Lost terminal identification method and system for metropolitan area network, and medium
By optically processing Internet access traffic at the metropolitan area network interface, combining the threat intelligence database and unit IP database, using secure DNS devices to collect internal domain name requests, the problem of difficult to identify and locate the lost terminal in the existing technology is solved, and the rapid and accurate identification and positioning of the lost terminal is achieved.
Patent Information
- Application Number
- CN202510647531.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art is difficult to identify and locate lost terminals in the metropolitan area network in a timely and accurate manner.
By spectroscopic processing of the uplink Internet access traffic of the operator's core network equipment at the metropolitan area network interface, filtering and formatting the domain name request resolution message flow, using the threat intelligence library to double-order iterative matching to identify risk domain names, combining the unit IP library and secure DNS equipment to collect internal domain name requests, and finally identifying the lost terminal through multi-level data matching.
It realizes fast and accurate identification and positioning of lost terminals, and improves the protection efficiency and accuracy of complex attack modes.
Smart Images

Figure CN120165990A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital communication technologies, and particularly to a method, system and medium for identifying compromised terminals for a metropolitan area network. Background Art
[0002] With the continuous upgrading of network attack means, traditional security protection methods have been difficult to effectively cope with complex attacks such as advanced persistent threats (APTs) and botworms. These attacks usually spread horizontally by controlling a large number of terminal devices (compromised terminals), greatly increasing the difficulty of detection and prevention. Existing security protection systems mainly rely on static rules and external threat intelligence for protection, but due to the lack of comprehensive monitoring of large-scale network traffic and real-time detection of internal terminals, it is difficult to quickly and accurately identify the devices that have been attacked and controlled. Therefore, there is an urgent need for a security protection technology that can monitor large-scale network traffic in real time and accurately identify compromised terminals. Summary of the Invention
[0003] This application provides a method, system and medium for identifying compromised terminals for a metropolitan area network, which is used to solve the technical problem that traditional network security protection in the prior art cannot timely and accurately identify and locate compromised terminals controlled by attackers in the metropolitan area network.
[0004] In the first aspect of this application, a method for identifying compromised terminals for a metropolitan area network is provided. The method includes: at the interface of the metropolitan area network, performing optical splitting processing on the upstream Internet access traffic of the core network devices of the operator, importing the optical signal into a splitter, filtering and forwarding the Internet access traffic in the splitter to obtain a set of domain name request resolution message streams; sending the set of domain name request resolution message streams to a collector for formatting processing to obtain a set of domain name request resolution records; performing two-stage iterative matching on the domain name fields in the set of domain name request resolution records with malicious domain names in a threat intelligence database, marking the threat level of the domain name request resolution records according to the matching results, and summarizing the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records; retrieving based on the set of risk domain name request resolution records and a unit IP library respectively to determine a set of compromised units, and deploying a secure DNS device at the Internet exit of each compromised unit to collect internal domain name request records to obtain a cluster of internal domain name request records; matching the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the cluster of internal domain name request records to obtain a set of compromised terminals.
[0005] The second aspect of the present application provides a compromised terminal identification system for a metropolitan area network. The system includes: a domain name request resolution module, which is used to perform optical splitting on the upstream Internet access traffic of the core network equipment of the operator at the interface of the metropolitan area network, import the optical signal into a splitter, filter and forward the Internet access traffic in the splitter to obtain a set of domain name request resolution message streams; a formatting processing module, which is used to send the set of domain name request resolution message streams to a collector for formatting processing to obtain a set of domain name request resolution records; a risk domain name identification module, which is used to perform two-stage iterative matching on the domain name fields in the set of domain name request resolution records and the malicious domain names in the threat intelligence library, identify the threat levels of the domain name request resolution records according to the matching results, and summarize the domain name request resolution records marked as serious threats into a set of risk domain name request resolution records; an internal domain name request record collection module, which is used to retrieve based on the set of risk domain name request resolution records and the unit IP library of each unit to determine a set of compromised units, and deploy a secure DNS device at the Internet exit of each compromised unit to collect internal domain name request records to obtain a cluster of internal domain name request records; a compromised terminal matching module, which is used to match the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the cluster of internal domain name request records to obtain a set of compromised terminals.
[0006] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages: The method, system and medium for identifying compromised terminals for a metropolitan area network provided by the present application relate to the field of digital communication technologies. By performing optical splitting on the metropolitan area network traffic, filtering and formatting the domain name request resolution message streams, using two-stage matching and the threat intelligence library to identify risk domain names, determining compromised units by matching with the unit IP library, and deploying a secure DNS device at their Internet exits to collect internal domain name requests, and finally identifying compromised terminal devices by matching internal records, it solves the technical problem in the prior art that traditional network security protection cannot timely and accurately identify and locate compromised terminals controlled by attackers in the metropolitan area network, and achieves the technical effect of quickly and accurately identifying and locating compromised terminals through multi-level data matching and internal and external network traffic analysis. Description of the Drawings
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0009] Figure 1 Schematic flow diagram of the compromised terminal identification method for the metropolitan area network provided by the embodiment of the present application; Figure 2 Schematic structural diagram of the compromised terminal identification system for the metropolitan area network provided by the embodiment of the present application.
[0010] Explanation of reference numerals: Domain name request resolution module 11, formatting processing module 12, risk domain name identification module 13, internal domain name request record collection module 14, compromised terminal matching module 15. Detailed implementation manners
[0011] The present application provides a compromised terminal identification method, system and medium for the metropolitan area network, which is used to solve the technical problem that traditional network security protection in the prior art cannot timely and accurately identify and locate compromised terminals controlled by attackers in the metropolitan area network.
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0013] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0014] Embodiment 1, as Figure 1 shown, the present application provides a compromised terminal identification method for the metropolitan area network, and the method includes: P10: At the interface of the metropolitan area network, perform optical splitting on the upstream Internet access traffic of the operator's core network devices, import the optical signal into the splitter, filter and forward the Internet access traffic in the splitter, and obtain a set of domain name request resolution message flows.
[0015] Among them, in the splitter, use the five-tuple to filter the traffic in the Internet access traffic that does not meet the preset requirements to obtain the set of domain name request resolution message flows. Among them, the preset requirement is that the domain name request traffic must conform to the UDP or TCP protocol, and the destination port is the corresponding UDP port or TCP port 53.
[0016] It should be understood that at the interface of the metropolitan area network, first perform optical splitting on the upstream Internet access traffic of the operator's core network devices. The purpose is to separate the full network traffic from different operators. Specifically, the optical signal is extracted from the operator's core network devices and imported into the splitter. Through this optical splitting process, all Internet access traffic from different operators is concentrated in the splitter for unified traffic screening and processing.
[0017] In the splitter, screen the received Internet traffic through the five-tuple filtering technology. The five-tuple filtering method classifies and identifies data streams based on the five-tuple (source IP address, source port, destination IP address, destination port, and protocol type). In this step, the preset five-tuple filtering requirement is that the traffic must conform to the UDP protocol or TCP protocol, and the destination port must be the UDP port 53 or TCP port 53. These two ports are specifically used for DNS (Domain Name System) request resolution.
[0018] During the filtering process, by checking the five-tuple, the splitter filters out all traffic that does not meet the above conditions. For example, if the destination port of a certain data stream is not 53, or its protocol is not UDP or TCP, then this traffic will be excluded, thus avoiding interference from irrelevant traffic to subsequent processing. After this screening, what is finally left is the traffic that conforms to the domain name request resolution message flow, and these message flows contain key data related to DNS requests.
[0019] Through the above screening, the splitter finally obtains a set of domain name request resolution message flows, which includes all qualified DNS requests and resolution messages. The key data included in these message flows includes fields such as domain name request time, source IP address, destination IP address, resolved domain name, and resolved IP address. These domain name request resolution records will be transmitted to the downstream collector for further processing and storage for risk analysis, entity identification, and terminal location in subsequent steps.
[0020] Through precise screening by the splitter, we ensure that only data flows related to DNS domain name request resolution are collected, and avoid wasting non-target traffic. This method not only improves the accuracy of data collection, but also enables subsequent security analysis to focus on high-risk domain name requests, thereby improving the efficiency and accuracy of the system.
[0021] P20: Send the domain name request resolution message flow set to the collector for formatting, and obtain a domain name request resolution record set.
[0022] The domain name request resolution record set includes at least the domain name request time, domain name request source IP address, source port, domain name server IP address, destination port, requested resolution domain name, domain name resolution response time and resolution IP address.
[0023] Optionally, in this step, the domain name request resolution message flow set is sent to the collector for formatting. After receiving these message flows, the collector will parse them, extract key information and convert them into a standardized domain name request resolution record set. These records contain multiple fields to ensure that the detailed information of each domain name request can be fully reflected, which is convenient for subsequent analysis and tracing.
[0024] First, the collector extracts the domain name request time and records the specific time when each request occurs. This is crucial for analyzing the time distribution of requests and whether there is abnormal traffic or attack mode. Next, the domain name request source IP address is extracted. This field can help trace back to the device or terminal that initiated the request and further identify potential attack sources. The source port is also necessary information. It identifies the port used when the request was initiated, which is crucial for locating network sessions and analyzing the source of requests.
[0025] At the same time, the domain name server IP address will be recorded, which points to the DNS server that performs domain name resolution, helping to determine which server the request is resolved through. This field helps analyze whether there is an abnormality in the DNS server itself or the possibility of being attacked. The destination port field is usually port 53, which is the standard port of the DNS protocol, ensuring that the request meets the basic requirements of DNS requests. Next, the collector will extract the domain name requested for resolution, that is, the target domain name that initiated the request. This field is the core information of the request and can directly reflect the attacker's goal or attack intention.
[0026] The collector also records the domain name resolution response time, which reflects the response speed of the DNS service and can help identify abnormal behaviors of the DNS service. Especially in an attack scenario, abnormal response time may indicate problems such as a Denial of Service (DoS) attack. Finally, the collector extracts the resolved IP address, which is the result returned by the DNS server and identifies the target server or device to which the domain name is ultimately resolved.
[0027] Through formatting, the collector converts these raw message streams into structured records and stores them in the server database. These records will include various key information of the domain name requests, ensuring that subsequent analysis can perform effective risk assessment, threat identification, and terminal location based on this data. At the same time, the records stored by the collector need to have a sufficient traceability duration to enable traceability analysis in case of subsequent security incidents. Table 1 shows the specific descriptions of each field in the domain name request resolution record set: Table 1: Domain Name Resolution Record Table (Resolution) Field Name Field Name Meaning Request_Time Domain Name Request Time Source_IP Domain Name Request Source IP Address Source_Port Source Port DNS_IP Domain Name Server IP Destination_Port Destination Port DomainName Domain Name to be Resolved Response_Time Domain Name Resolution Response Time Resolved_IP Resolved IP Address Threat_Type Threat Type Threat_Level Threat Level The formatted domain name request resolution record set not only improves the readability and operability of the data but also provides accurate data support for subsequent threat detection, entity identification, and terminal location in the system.
[0028] P30: Perform a two-stage iterative matching between the domain name field in the domain name request resolution record set and the malicious domain names in the threat intelligence library. According to the matching results, mark the threat levels of the domain name request resolution records, and summarize the domain name request resolution records marked as severe threats into a risk domain name request resolution record set.
[0029] Furthermore, step P30 of the embodiment of the present application further includes: P31: Extract two-stage keywords from the malicious domain names in the threat intelligence library to obtain a first-order keyword set and a second-order mapped keyword group set; P32: Use an encoder to perform a first-order matching on the domain name field in the domain name request resolution record set according to the first-order keyword set to obtain a first-order matching result; P33: Use the second-order mapped keyword group set to perform implicit iterative matching recognition on the domain name request resolution records corresponding to the first-order matching result to obtain a second-order matching information set, where each second-order matching information corresponds to a domain name request resolution record in the first-order matching result; P34: Call a threat level identifier to identify the threat levels of the second-order matching information set, and perform mapping marking on the domain name request resolution records in the first-order matching result according to the recognition results to obtain a marked first-order matching result; P35: Summarize the domain name request resolution records marked as severe threats in the marked first-order matching result to obtain a risk domain name request resolution record set.
[0030] Specifically, a two - stage iterative matching is performed between the domain name fields in the domain name request resolution record set and the malicious domain names in the threat intelligence library. The core of this process lies in accurately identifying domain name requests that may be related to malicious activities through a multi - level keyword extraction and matching mechanism, assigning corresponding threat levels to these domain name requests, and finally screening out domain name requests with serious threats and aggregating them into a risk domain name request resolution record set.
[0031] Specifically, first, two - stage keyword extraction is performed on the malicious domain names in the threat intelligence library. The purpose of this process is to extract a set of first - order keywords and a set of second - order mapped keyword groups from the malicious domain names. The set of first - order keywords includes common keyword parts in the malicious domain names, usually the explicit parts in the domain names (e.g., "malware", "phishing", etc.), while the set of second - order mapped keyword groups is a set of implicit features derived from the first - order keywords, which may be domain name suffixes, IP address ranges, etc. related to malicious activities. Exemplarily, collect intelligence on various malicious domain names, and the threat intelligence of each manufacturer must be aggregated into a threat intelligence record table, with fields at least as listed in Table 2: Table 2: Threat Intelligence Record Table (Threat) Field Name Field Name Meaning DomainName Domain Name CNAME Alias DomainName_Owner Domain Name Owner Registration_Time Domain Name Registration Time Expiration_Time Domain Name Expiration Time Service_Provider Domain Name Service Provider Resolved_IP Resolved IP IP_Location Geographical Location of Resolved IP Threat_Information Threat Intelligence Information Threat_Type Threat Type Threat_Level Threat Level Information_Update_Time Intelligence Update Time Information_Source Intelligence Source Next, use an encoder to process the set of first - order keywords and perform a first - order matching on the domain name fields in the domain name request resolution record set. The purpose of this step is to initially screen out records that contain first - order keywords in the domain name. The encoder generates a preliminary matching result through the matching of domain name fields and identifies domain name request records that may pose a threat.
[0032] Furthermore, use the set of second - order mapped keyword groups to perform implicit iterative matching and identification on the domain name request records in the first - order matching results. This process is a deeper analysis and identification of the first - order matching results, taking into account implicit features and more complex attack patterns. Through this implicit matching, more potential threat information can be identified, forming a set of second - order matching information, where each element in the set corresponds to a domain name request resolution record in the first - order matching results.
[0033] Then, call a threat level identifier to identify the threat level of the set of second - order matching information. The identifier determines the threat level of the domain name request based on the data characteristics in the second - order matching information and maps this level information back to the domain name request records in the first - order matching results. In this way, each domain name request record can be labeled with a threat level, providing a basis for subsequent risk assessment and defense.
[0034] Finally, extract the domain name request resolution records identified as serious threats from the first-order matching results and summarize them to form a set of risk domain name request resolution records. These records represent domain name requests that have been determined to be of high risk during the entire matching process, usually domain names related to cyberattacks or malicious behavior.
[0035] Exemplarily, for the risk analysis of domain name requests: match each domain name request resolution record one by one with the threat intelligence record table, and update fields such as the threat type and threat level of the domain name request resolution record.
[0036] The reference SQL statements for database operations are as follows (Domain Name Resolution Record Table: Table A, Threat Intelligence Record Table: Table B, Domain Name Field: Field D, Threat Type Field: Field X, Threat Level Field: Field Y): UPDATE Table A; SET Field X = (SELECT b.Field X FROM Table B b WHERE b.Field D = Table A.Field D); Field Y = (SELECT b.Field Y FROM Table B b WHERE b.Field A = Table A.Field A) WHERE EXISTS (SELECT 1 FROM Table B b WHERE b.Field D = Table A.Field D); The reference SQL statement for finding the source IP address of the domain name request with the threat level of "serious threat" is as follows (Domain Name Resolution Record Table: Table A, Source IP Address Field of Domain Name Request: Field S, Threat Level Field: Field Y): SELECT Field S FROM Table A; WHERE Field Y = "serious threat"; In this way, the IP addresses infected with malicious programs of the "serious threat" level are output.
[0037] Through the above steps, this method can not only achieve accurate threat detection through double-order keyword matching, but also ensure that the system can identify potential network security threats in a timely and accurate manner through multi-level matching and threat level identification. This method can provide strong data support for network security protection, especially providing an efficient detection means in the face of complex and changing malicious attacks.
[0038] Further, step P31 of the embodiment of the present application further includes: P31-1: Extract malicious domain names from the threat intelligence library to obtain a set of first-order keywords; P31-2: Retrieve the records of the first-order keyword set in the threat intelligence respectively according to the preset set of implicit malicious domain name keyword types to determine a set of second-order mapping keyword groups. The preset set of implicit malicious domain name keyword types includes at least DNS resolution, historical deformation, alias, and jump pointer.
[0039] It should be understood that the processing of malicious domain names in the threat intelligence library can be further refined to extract a more targeted and accurate set of keywords.
[0040] First, extract all malicious domain name information from the threat intelligence database and analyze these malicious domain names to generate a first-order keyword set. This set consists of characteristic words that appear explicitly in malicious domain names. These characteristics are usually iconic words in domain names or words related to attack types, such as "malware" and "phishing". These words reflect the characteristics of specific malicious activities or attack families. By extracting these keywords, we can lay the foundation for subsequent malicious domain name matching and identification.
[0041] Next, the system further retrieves and analyzes the records in the first-order keyword set according to the preset set of hidden malicious domain name keyword types. The set of hidden malicious domain name keyword types includes multiple features, which may not appear explicitly in the domain name, but can effectively help identify malicious behavior. First, the system analyzes the resolution behavior of domain name requests based on the feature of DNS resolution. Malicious domain names often use specific resolution methods, such as using non-standard DNS requests or adopting special DNS query modes. These behaviors can help the system identify potential attack domain names. Secondly, historical deformation is another important analysis point. Malicious domain names often perform character replacement or spelling deformation in order to evade static matching rules or increase the concealment of attacks. By looking back at historical data, the system can identify these deformed domain names and further increase the accuracy of matching.
[0042] In addition, malicious domain names usually have aliases, that is, multiple domain names are used to point to the same malicious control point. Attackers often control victim terminals by using multiple domain names or IP addresses. These aliases are usually highly correlated with the main domain name. By extracting this alias information, the system can identify more potential malicious activities. Finally, the characteristics of jump pointing also provide important clues for the identification of malicious domain names. Attackers may resolve a domain name to another domain name or IP address through DNS jump. This jump is usually used to evade detection and increase the concealment of the attack. By analyzing the jump pointing of the domain name, the system can track these jump paths and further reveal the potential threats of malicious domain names.
[0043] By searching and analyzing these hidden keyword types, the system can mine more hidden malicious features from the first-order keyword set, and finally form a second-order mapping keyword group set. These mapping keywords represent the malicious features of the domain name at a deeper level, and can help the system identify more complex and hidden malicious domain names.
[0044] After this process is completed, the system will perform implicit iterative matching on the domain name request resolution records in the first-order matching results based on the second-order mapping keyword group set, so as to obtain a more accurate malicious domain name recognition result. This iterative matching process can not only further improve the matching accuracy of malicious domain names, but also discover some malicious domain names that evade traditional matching rules through means such as deformation, aliasing, or redirection, providing a more detailed basis for subsequent threat level identification.
[0045] In summary, through the extraction of first-order keywords and the in-depth exploration of second-order implicit features, this embodiment can achieve the accurate identification of malicious domain names, ensuring that even malicious activities carried out through hidden means such as deformation and redirection can be discovered in a timely manner, providing strong technical support for network security protection.
[0046] Furthermore, step P33 of the embodiment of the present application further includes: P33-1: Extract the first domain name request resolution record from the first-order matching results; P33-2: Perform mapping matching on the second-order mapping keyword group set according to the first-order keywords corresponding to the first domain name request resolution record to obtain the matching second-order mapping keyword group; P33-3: Perform associated matching between the matching second-order mapping keyword group and the first field of the first domain name request resolution record. If the matching is successful, obtain the first matching mapping keyword; P33-4: Add the semantic annotation corresponding to the first matching mapping keyword, the first matching mapping keyword, and the matching second-order mapping keyword group into an initially empty vector to obtain the first memory vector; P33-5: Obtain the second field of the first domain name request resolution record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain the second memory vector, and so on. The memory vector obtained after the field matching is completed is used as the first second-order matching information of the first domain name request resolution record; P33-6: Perform implicit iterative matching recognition on the domain name request resolution records corresponding to the first-order matching results in combination with the second-order mapping keyword group set to obtain a second-order matching information set.
[0047] Specifically, the processing process of the domain name request resolution records in the first-order matching results can be further refined. This process uses a series of sub-steps to deeply explore potential threat information through implicit iterative matching recognition.
[0048] First, extract the first domain name request resolution record from the previously obtained first-order matching results. The purpose of this operation is to select the records in the first-order matching results that need to be further processed. These records have been preliminarily matched with malicious domain names through first-order keywords and are the basic data for subsequent in-depth analysis.
[0049] Next, based on the first-order keywords corresponding to the parsed records of the first domain name request extracted, the system will perform mapping matching on the second-order mapping keyword group set. The key to this process is that the system finds the relevant second-order mapping keywords in the threat intelligence library according to the first-order keywords, so as to discover more concealed and complex malicious features. The keywords in the second-order mapping keyword group set usually contain some features that are not easily detected directly, and these features are potential signs of malicious activities.
[0050] Next, the matched second-order mapping keyword group is associated and matched with the first field of the first domain name request parsing record. The first field here may be a domain name field, a request time, or other key data fields. In this way, the system can verify the relationship between the domain name request and the second-order keywords, and further confirm whether the domain name request is related to malicious behavior. If the match is successful, the first matching mapping keyword will be obtained, which is the key link between the domain name request and the malicious activity.
[0051] Furthermore, the obtained first matching mapping keyword and its corresponding semantic annotation are added to the initially empty vector together with the second-order mapping keyword group, thereby generating the first memory vector. The purpose of this step is to encode the matched keywords and their semantic information to form a vector representation for subsequent iterative matching and analysis. This memory vector not only contains the matching information but also helps the system understand the semantic meaning behind each keyword, providing richer context for subsequent matching.
[0052] Subsequently, the second field of the first domain name request parsing record is obtained, and the first memory vector is used to perform implicit iterative matching on this field. Implicit iterative matching means that without directly showing all the information, based on the memory vector obtained in the previous step, further matching is performed on the second field. This process continuously updates the information in the vector and finally obtains the second memory vector. In this way, the system can continuously optimize the matching result through step-by-step iteration and obtain a more accurate malicious domain name recognition result.
[0053] Finally, the system continues to perform implicit iterative matching and recognition on the domain name request parsing record corresponding to the first-order matching result in combination with the second-order mapping keyword group set, and finally obtains the second-order matching information set. This set contains all the relevant information obtained through the second-order keywords and implicit matching, marking the deep matching result of the domain name request in multiple dimensions. Through these second-order matching information, the system can more accurately identify potential malicious activities and provide stronger threat recognition capabilities.
[0054] Through the above steps, the matching accuracy of domain name request resolution records can be effectively improved, not limited to the matching of explicit features, but also delving into the implicit features behind domain name requests. Through iterative matching and vectorization techniques, this method provides the system with more powerful malicious domain name detection capabilities, enabling the discovery of more complex and concealed attack behaviors and providing stronger data support for network security protection.
[0055] Further, step P33-5 of the embodiment of the present application further includes: P33-51: Use the first memory vector to perform associative matching on the second field of the first domain name request resolution record. If the matching is successful, obtain the second matching mapping keyword; P33-52: Add the semantic annotation corresponding to the first matching mapping keyword and the second matching mapping keyword into the first memory vector to update it, and obtain the second memory vector.
[0056] In a possible embodiment of the present application, the implicit iterative matching process between the first memory vector and the second field can be further refined. This process mainly performs associative matching on the second field of the first domain name request resolution record, continuously optimizing and updating the memory vector to obtain a more accurate malicious domain name recognition result.
[0057] First, use the first memory vector to perform associative matching on the second field of the first domain name request resolution record. The second field may be other information related to the domain name request, such as the source port of the request, response time, resolved IP address, etc. Using the information in the first memory vector, the system matches the second field to find potential features related to malicious activities. When the second field matches the keyword in the memory vector successfully, the system will obtain the second matching mapping keyword. These mapping keywords represent the malicious features in the second field, further deepening the association between the domain name request and malicious activities.
[0058] Next, add the first matching mapping keyword and the second matching mapping keyword obtained through matching and their corresponding semantic annotations into the first memory vector and update it. The core of this process is to continuously integrate the matched keywords and their semantic information into the memory vector, making the memory vector gradually enriched and optimized during multiple iterative matching processes. This updated memory vector, that is, the second memory vector, contains more context information and malicious activity features, which helps to improve the accuracy of subsequent matching.
[0059] Through the operations of these two sub-steps, the system can gradually improve the matching accuracy of domain name request resolution records through iterative matching. Each matching and vector update will make the system's identification of malicious domain names more in-depth and accurate, thus enhancing the detection ability of potential threats. Finally, the second memory vector, as the updated vector, will provide more accurate data support for the subsequent identification of second-order matching information, helping the system accurately identify threatening domain name request records.
[0060] P40: Retrieve based on the set of risk domain name request resolution records and the unit IP library respectively to determine the set of compromised units, and deploy security DNS devices at the Internet egress of each compromised unit to collect internal domain name request records, obtaining a cluster of internal domain name request records.
[0061] Specifically, retrieve based on the set of risk domain name request resolution records obtained in the previous step and the unit IP library. The purpose is to determine the set of compromised units by comparing risk domain name requests with the public IP addresses of each unit. These units refer to those that may have been controlled by malware or are targets under cyber attacks.
[0062] First, by retrieving the unit IP library, the system can identify which units' public IP addresses match the IP addresses in the risk domain name requests. The unit IP library contains the public IP address records of each unit. Through these records, the system can compare the received risk domain name requests and further locate the specific victim units. Whenever the system finds a matching public IP address, it can determine that the unit has become a target of cyber attacks and thus include it in the set of compromised units. The network egress devices of these units may have been controlled by attackers to transmit malicious instructions or execute attack behaviors.
[0063] Exemplarily, through the unit IP table (see Table 3), the corresponding unit name can be queried by the IP address (the IP address infected with malware analyzed in the analysis step), so as to locate which unit is infected with what malware.
[0064] Table 3: Unit IP Table (IP_Table) Field Name Field Name Meaning Organization_Name Organization Name Address Organization Address Contacts Contacts Telephone Contact Phone Region Region Belonged to IP_address IP Address Operator Operator Belonged to The SQL statement for the database operation to find the unit name from the unit IP table by the specific IP address is as follows: SELECT Organization_Name FROM IP_Table; WHERE IP_address = "IP address".
[0065] Next, to further identify and prevent these affected units, the system will deploy security DNS devices at the Internet exits of each compromised unit. The function of these security DNS devices is to monitor and record all domain name request data within the unit, thereby obtaining a cluster of domain name request records within the unit. The internal domain name request record cluster refers to the set of domain name request data initiated by all terminals within the internal network of the unit. By deploying security DNS devices at the Internet exit, the system can capture all DNS requests sent from the unit and further determine whether there are compromised terminal devices within the unit through the analysis of these requests.
[0066] The purpose of deploying security DNS devices is as follows: on the one hand, it can collect domain name request data within the unit in real time, providing basic data for subsequent malicious domain name detection; on the other hand, the collected internal domain name request records can also be used to analyze the security status of the internal network of the unit, especially to determine whether there are compromised terminals. During the collection process, the DNS device will not only record the specific content of the domain name request but also further determine whether there are malicious activities based on the response information of the domain name resolution. Exemplarily, at the unit's exit, install a security DNS device and send all domain name request records of the unit to the security DNS device for identification, storage, and disposal. The stored unit domain name resolution record table is shown in Table 4: Table 4: Unit Domain Name Resolution Record Table (Resolution) Field Name Field Name Meaning Organization_Name Organization Name Request_Time Domain Name Request Time Source_IP Domain Name Request Source IP Address Source_Port Source Port DNS_IP Domain Name Server IP Destination_Port Destination Port DomainName Domain Name to be Resolved Response_Time Domain Name Resolution Response Time Resolved_IP Resolved IP Address Threat_Type Threat Type Threat_Level Threat Level Through this process, the system can achieve security monitoring from the external public network to the internal network, ensuring that compromised units that have been controlled can be accurately identified and located, and providing sufficient data support for subsequent terminal location, risk prevention, and repair.
[0067] P50: Match the set of risk domain name request resolution records with the corresponding set of internal domain name request records in the internal domain name request record cluster to obtain the set of compromised terminals.
[0068] Optionally, match the set of risk domain name request resolution records obtained in the previous step with the corresponding records in the internal domain name request record cluster to identify and locate the specific set of compromised terminals.
[0069] Specifically, first extract those domain name requests marked as having a serious threat from the set of risk domain name request resolution records. These domain name requests are usually related to known malicious activities and may include domain names from malicious control servers or domain names related to attack behaviors. Then, the system will compare these risk domain name requests with the domain name request record cluster in the internal network of the unit. The internal domain name request record cluster refers to all domain name request records initiated within the compromised unit, and these requests come from all terminal devices of the unit.
[0070] By comparing the corresponding records in the set of risk domain name request resolution records and the cluster of internal domain name request records, the system can identify which internal terminal devices have sent requests related to known malicious domain names. The key to this step is to associate external risk domain names with internal domain name request data, so as to achieve accurate positioning of compromised terminals. The successfully matched records will represent those terminals that may have been controlled by attackers, and the system will include these terminals in the set of compromised terminals.
[0071] Once the compromised terminals are matched, the system can clarify the association between these terminals and malicious domain names, and then conduct further security analysis and response. In this way, the system not only identifies the controlled terminals, but also can trace back to the attack sources involved by these terminals, helping the security team to quickly take measures to prevent further damage.
[0072] In summary, by matching risk domain names with the domain name request data within the organization, the specific terminals under attack control can be accurately identified, thus providing an important basis for subsequent security repair and threat disposal.
[0073] In summary, the embodiments of the present application at least have the following technical effects: In the present application, the upstream Internet traffic of the operator's core network equipment is split at the metropolitan area network interface, and the domain name request resolution packet stream set is filtered and obtained, and then formatted into a set of domain name request resolution records. Through double-stage matching with malicious domain names in the threat intelligence library, serious threats are identified and summarized into a set of risk domain name request resolution records. Based on this set, a search is performed with the organization's IP library to determine the compromised organization, and a secure DNS device is deployed to collect internal domain name requests. Finally, the risk domain names are matched with the internal records to obtain a set of compromised terminals and locate the controlled devices.
[0074] It achieves the technical effects of identifying and positioning compromised terminals in real time and accurately through double-stage iterative matching and real-time domain name request analysis, and improving the protection efficiency and accuracy against complex attack patterns.
[0075] Embodiment 2, based on the same inventive concept as the method for identifying compromised terminals in the metropolitan area network in the foregoing embodiment, as Figure 2 shown, the present application provides a system for identifying compromised terminals in the metropolitan area network. The system in the embodiments of the present application and the method embodiments are based on the same inventive concept. Among them, the system includes: A domain name request resolution module 11, which is used to perform splitting processing on the upstream Internet access traffic of the operator's core network equipment at the interface of the metropolitan area network, import the optical signal into a splitter, and filter and forward the Internet access traffic in the splitter to obtain a set of domain name request resolution packet streams.
[0076] The formatting module 12 is configured to send the domain name request parsing message stream set to a collector for formatting processing to obtain a domain name request parsing record set.
[0077] The risk domain name identification module 13 is configured to perform a two-stage iterative matching between the domain name fields in the domain name request parsing record set and the malicious domain names in the threat intelligence library, identify the threat levels of the domain name request parsing records according to the matching results, and summarize the domain name request parsing records marked as serious threats into a risk domain name request parsing record set.
[0078] The internal domain name request record collection module 14 is configured to retrieve based on the risk domain name request parsing record set and the unit IP library respectively to determine the set of compromised units, and deploy security DNS devices at the Internet egress of each compromised unit to collect internal domain name request records to obtain an internal domain name request record cluster.
[0079] The compromised terminal matching module 15 is configured to match the risk domain name request parsing record set with the corresponding internal domain name request record set in the internal domain name request record cluster to obtain a set of compromised terminals.
[0080] Furthermore, the domain name request parsing module 11 is further configured to perform the following steps: Filter the Internet access traffic that does not meet the preset requirements in the shunt using the five-tuple to obtain the domain name request parsing message stream set, where the preset requirement is that the domain name request traffic should conform to the UDP or TCP protocol, and the destination port is the corresponding UDP port or TCP port 53.
[0081] Furthermore, the formatting module 12 is further configured to perform the following steps: The domain name request parsing record set at least includes the domain name request time, the source IP address of the domain name request, the source port, the IP address of the domain name server, the destination port, the requested domain name to be resolved, the domain name resolution response time, and the resolved IP address.
[0082] Furthermore, the risk domain name identification module 13 is further configured to perform the following steps: Perform two - stage keyword extraction on the malicious domain names in the threat intelligence database to obtain a set of first - order keywords and a set of second - order mapped keyword groups; use an encoder to perform first - order matching on the domain name fields in the domain name request parsing record set according to the set of first - order keywords to obtain a first - order matching result; use the set of second - order mapped keyword groups to perform implicit iterative matching recognition on the domain name request parsing records corresponding to the first - order matching result to obtain a set of second - order matching information, where each piece of second - order matching information corresponds to a domain name request parsing record in the first - order matching result; call a threat level identifier to perform threat level recognition on the set of second - order matching information, and perform mapping identification on the domain name request parsing records in the first - order matching result according to the recognition result to obtain a completed first - order matching result with identification; summarize the domain name request parsing records marked as serious threats in the first - order matching result with identification to obtain a set of risk domain name request parsing records.
[0083] Further, the risk domain name recognition module 13 is further configured to perform the following steps: Extract the first domain name request parsing record from the first - order matching result; perform mapping matching on the set of second - order mapped keyword groups according to the first - order keywords corresponding to the first domain name request parsing record to obtain a matching second - order mapped keyword group; perform associated matching between the matching second - order mapped keyword group and the first field of the first domain name request parsing record. If the matching is successful, obtain a first matching mapped keyword; add the semantic annotation corresponding to the first matching mapped keyword, the first matching mapped keyword, and the matching second - order mapped keyword group into an initially empty vector to obtain a first memory vector; obtain the second field of the first domain name request parsing record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain a second memory vector, and so on. The memory vector obtained after the field matching is completed is used as the first second - order matching information of the first domain name request parsing record; perform implicit iterative matching recognition on the domain name request parsing records corresponding to the first - order matching result in combination with the set of second - order mapped keyword groups to obtain a set of second - order matching information.
[0084] Further, the risk domain name recognition module 13 is further configured to perform the following steps: Perform associated matching on the second field of the first domain name request parsing record using the first memory vector. If the matching is successful, obtain a second matching mapped keyword; add the semantic annotation corresponding to the first matching mapped keyword and the second matching mapped keyword into the first memory vector to update it, and obtain a second memory vector.
[0085] Further, the risk domain name recognition module 13 is further configured to perform the following steps: Extract malicious domain names from the threat intelligence library to obtain a first-order keyword set; retrieve the records of the first-order keyword set in the threat intelligence according to a preset set of implicit malicious domain name keyword types respectively to determine a second-order mapped keyword group set.
[0086] Further, the risk domain name recognition module 13 is further configured to perform the following steps: The preset set of implicit malicious domain name keyword types includes at least DNS resolution, historical deformation, aliases, and jump pointers.
[0087] Embodiment 3, based on the same inventive concept as the method for identifying compromised terminals for a metropolitan area network in the foregoing embodiments, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method in Embodiment 1 is implemented.
[0088] Through the foregoing detailed description of the method for identifying compromised terminals for a metropolitan area network in this specification, those skilled in the art can clearly know the method, system, and medium for identifying compromised terminals for a metropolitan area network in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method part.
[0089] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0090] It should be noted that the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. In addition, the above specific embodiments of this specification have been described. Further, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0091] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0092] This specification and the drawings are merely illustrative of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A method for identifying a lost terminal in a metropolitan area network, characterized in that: The method comprises: At the interface of the metropolitan area network, the uplink Internet access traffic of the operator's core network equipment is split and processed, and the optical signal is introduced into the splitter, where the Internet access traffic is filtered and forwarded to obtain a set of domain name request resolution message flows; Sending the domain name request resolution message flow set to a collector for formatting, to obtain a domain name request resolution record set; Performing a double-order iterative match between the domain name field in the domain name request resolution record set and the malicious domain name in the threat intelligence library, marking the threat level of the domain name request resolution record according to the matching result, and aggregating the domain name request resolution records marked as serious threats into a risky domain name request resolution record set; Based on the risk domain name request resolution record set and the IP database of each unit, the set of compromised units is determined, and a secure DNS device is deployed at the Internet exit of each compromised unit to collect internal domain name request records and obtain internal domain name request record clusters; The risky domain name request resolution record set is matched with the corresponding internal domain name request record set in the internal domain name request record cluster to obtain the compromised terminal set.
2. The method for identifying a lost terminal in a metropolitan area network according to claim 1, characterized in that: Performing a two-stage iterative match between the domain name field in the domain name request resolution record set and the malicious domain name in the threat intelligence library, identifying the threat level of the domain name request resolution record according to the matching result, and aggregating the domain name request resolution records identified as serious threats into a risky domain name request resolution record set, including: Performing two-order keyword extraction on malicious domain names in the threat intelligence database to obtain a first-order keyword set and a second-order mapping keyword group set; Using an encoder to perform a first-order match on the domain name field in the domain name request resolution record set according to the first-order keyword set to obtain a first-order matching result; Using the second-order mapping keyword group set, implicitly iterate matching and identifying the domain name request resolution record corresponding to the first-order matching result, to obtain a second-order matching information set, wherein each second-order matching information corresponds to a domain name request resolution record in the first-order matching result; Calling a threat level identifier to identify the threat level of the second-order matching information set, mapping and identifying the domain name request resolution record in the first-order matching result according to the identification result, and obtaining an identified first-order matching result with the identification completed; The domain name request resolution records identified as serious threats in the first-order matching results are summarized to obtain a risk domain name request resolution record set.
3. The method for identifying a lost terminal in a metropolitan area network according to claim 2, characterized in that: Using the second-order mapping keyword group set, implicit iterative matching and identification are performed on the domain name request resolution record corresponding to the first-order matching result to obtain a second-order matching information set, including: Extracting a first domain name request resolution record from the first-order matching result; According to the first-order keyword corresponding to the first domain name request resolution record, mapping and matching the second-order mapping keyword group set is performed to obtain a matching second-order mapping keyword group; Associating and matching the matching second-order mapping keyword group with the first field of the first domain name request resolution record, and if the match is successful, obtaining the first matching mapping keyword; Add the semantic annotation corresponding to the first matching mapping keyword, the first matching mapping keyword and the matching second-order mapping keyword group into the initially empty vector to obtain a first memory vector; Obtain the second field of the first domain name request resolution record, perform implicit iterative matching on the second field using the first memory vector, and update the first memory vector according to the matching result to obtain the second memory vector, and so on, and use the memory vector obtained after the field matching is completed as the first and second order matching information of the first domain name request resolution record; The domain name request resolution record corresponding to the first-order matching result is combined with the second-order mapping keyword group set to perform implicit iterative matching identification to obtain a second-order matching information set.
4. The method for identifying a lost terminal in a metropolitan area network according to claim 3, characterized in that: Obtaining a second field of the first domain name request resolution record, performing implicit iterative matching on the second field using the first memory vector, and updating the first memory vector according to the matching result to obtain a second memory vector, including: Using the first memory vector to perform an association match on the second field of the first domain name request resolution record, if the match is successful, obtaining a second matching mapping keyword; The semantic annotation corresponding to the first matching mapping keyword and the second matching mapping keyword are added into the first memory vector to update it, so as to obtain the second memory vector.
5. The method for identifying a lost terminal in a metropolitan area network according to claim 4, characterized in that: include: Extract malicious domain names from the threat intelligence database to obtain a first-order keyword set; According to the preset implicit malicious domain name keyword type set, the records of the first-order keyword set in the threat intelligence are searched respectively to determine the second-order mapping keyword group set.
6. The method for identifying a lost terminal in a metropolitan area network according to claim 5, characterized in that: The preset hidden malicious domain name keyword type set includes at least DNS resolution, historical deformation, alias and jump pointing.
7. The method for identifying a lost terminal in a metropolitan area network according to claim 1, characterized in that: In the diverter, the quintuple is used to filter the Internet access traffic that does not meet the preset requirements to obtain the domain name request resolution message flow set, wherein the preset requirement is that the domain name request traffic must comply with the UDP or TCP protocol, and the target port is the corresponding UDP port or TCP53 port.
8. The method for identifying a lost terminal in a metropolitan area network according to claim 1, characterized in that: The domain name request resolution record set includes at least the domain name request time, the domain name request source IP address, the source port, the domain name server IP address, the destination port, the domain name requested for resolution, the domain name resolution response time and the resolution IP address.
9. A lost terminal identification system for metropolitan area networks, characterized in that: The system comprises: A domain name request resolution module, which is used to perform optical splitting processing on the uplink Internet access traffic of the operator's core network equipment at the interface of the metropolitan area network, guide the optical signal into the splitter, filter and forward the Internet access traffic in the splitter, and obtain a domain name request resolution message flow set; A formatting processing module, the formatting processing module is used to send the domain name request resolution message flow set to the collector for formatting processing to obtain a domain name request resolution record set; A risky domain name identification module, the risky domain name identification module is used to perform a two-stage iterative match between the domain name field in the domain name request resolution record set and the malicious domain name in the threat intelligence library, identify the threat level of the domain name request resolution record according to the matching result, and aggregate the domain name request resolution records identified as serious threats into a risky domain name request resolution record set; An internal domain name request record collection module, which is used to search based on the risk domain name request resolution record set and the IP library of each unit, determine the set of compromised units, and deploy a secure DNS device at the Internet exit of each compromised unit to collect internal domain name request records and obtain an internal domain name request record cluster; A compromised terminal matching module is used to match the risky domain name request resolution record set with the corresponding internal domain name request record set in the internal domain name request record cluster to obtain a compromised terminal set.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the method for identifying a lost terminal for a metropolitan area network according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
APT detection method based on matching of flow fingerprint and communication features
CN108833437A
Detection method and device of lost terminal, electronic equipment and storage medium
CN117978472A
Apparatus for detecting and filtering ddos attack based on request URI type
US20110107412A1
Hardware acceleration device for denial-of-service attack identification and mitigation
US20210306373A1
Attack detection method, and apparatus
WO2023216792A1
Cited By
Multi-source threat intelligence-driven DNS (Domain Name Server) security protection method and system
CN120979750A