Method, device and electronic equipment for determining a botnet master

By performing feature preprocessing and semantic reasoning using a large language model on the network flow data of the botnet master control terminal, combined with attack correlation analysis, the problem of inaccurate botnet master control terminal location was solved, and more efficient botnet tracing was achieved.

CN122339736APending Publication Date: 2026-07-03BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-03-23
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing botnet detection technologies struggle to accurately identify botnet master control units, primarily due to a lack of deep temporal logic and interactive semantic understanding of network communication behavior, as well as a lack of systematic correlation analysis of multiple controlled hosts, leading to inaccurate location.

Method used

By preprocessing the network flow data of the controlled host to identify botnet characteristics, host behavior description information is generated. Semantic reasoning is then performed using a large language model in the security domain to output a list of suspected master control endpoints. Finally, attack correlation analysis is used to determine the master control endpoint of the target botnet.

Benefits of technology

It achieves precise location of the botnet master control terminal, breaks through the limitations of traditional single host detection, improves the targeting and semantic understanding capabilities of attack tracing, and reduces false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122339736A_ABST
    Figure CN122339736A_ABST
Patent Text Reader

Abstract

This disclosure provides a method for identifying the master controller of a botnet, relating to the field of network security technology, particularly attack attribution, botnets, and deep learning. The specific implementation scheme is as follows: In response to the detection of attack traffic targeting external communication addresses, a set of controlled hosts corresponding to the attack traffic is determined, and network flow data of each controlled host within a preset attack attribution time window is extracted; for any controlled host, botnet feature preprocessing is performed on the network flow data to generate host behavior description information; the host behavior description information of each controlled host is input into a large language model in the security field, and through the prompt information configured for botnet feature analysis in the large language model, a list of suspected master controllers corresponding to each controlled host is output; the lists of suspected master controllers of each controlled host are aggregated, attack correlation analysis is performed, and the target botnet master controller is determined based on the analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of cybersecurity technology, particularly to attack attribution, botnets, and deep learning technologies. Background Technology

[0002] Current botnet detection technologies have the following limitations in identifying and tracing their command and control servers (C2Servers): Traditional methods rely heavily on feature rule matching or statistical model analysis. While they can identify some known or statistically anomalous attack traffic, they generally lack the ability to understand the deep temporal logic and interactive semantics of network communication behavior. The processing of network flow data is mostly limited to shallow feature extraction or simple filtering, failing to transform it into behavioral description information that can be deeply semantically analyzed. Furthermore, existing technologies rely heavily on the independent detection results of a single host, lacking systematic correlation analysis of suspected master control information from multiple controlled hosts. This makes it difficult to effectively integrate multi-source clues, penetrate the disguise of the botnet master control server, and result in inaccurate location of the target master control server in botnets. Summary of the Invention

[0003] This disclosure provides a method, apparatus, and electronic device for fixing a botnet master control terminal to solve at least one of the above-mentioned technical problems.

[0004] According to a first aspect of this disclosure, a method for determining the master control terminal of a botnet is provided, wherein the method includes: In response to the detection of attack traffic targeting external communication addresses, the set of controlled hosts corresponding to the attack traffic is determined, and network flow data of each controlled host in the set of controlled hosts is extracted within a preset attack tracing time window. For any one of the controlled hosts in the set of controlled hosts, the network flow data is preprocessed with botnet features to generate host behavior description information of the controlled host; The host behavior description information of each of the controlled hosts is input into a large language model in the security field, and semantic reasoning is guided by the prompt information configured for the feature analysis of botnets configured for the large language model, and a list of suspected master control terminals corresponding to each controlled host is output. The list of suspected master controllers of each of the controlled hosts is aggregated, and attack correlation analysis is performed to determine the master controller of the target botnet based on the analysis results.

[0005] According to a second aspect of this disclosure, an apparatus for determining the master control terminal of a botnet is provided, wherein the apparatus includes: The host determination module is used to determine the set of controlled hosts corresponding to the attack traffic in response to the detection of attack traffic targeting external communication addresses, and to extract the network flow data of each controlled host in the set of controlled hosts within a preset attack tracing time window. The host description module is used to perform botnet feature preprocessing on the network flow data for any one of the controlled hosts in the set of controlled hosts, and generate host behavior description information of the controlled host. The master control reasoning module is used to input the host behavior description information of each of the controlled hosts into a large language model in the security field, and guide it to perform semantic reasoning by using the prompt information configured for the feature analysis of botnets in the large language model, and output a list of suspected master control terminals corresponding to each controlled host. The master control determination module is used to aggregate the list of suspected master control terminals of each of the controlled hosts, perform attack correlation analysis, and determine the target botnet master control terminal based on the analysis results.

[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described method.

[0007] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described above.

[0008] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described above.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating a method for determining the master control terminal of a botnet according to the first embodiment of this disclosure; Figure 2 This is a schematic diagram of the structured template 01; Figure 3This is a flowchart illustrating another method for determining the master control terminal of a botnet, provided in the first embodiment of this disclosure. Figure 4 This is a flowchart illustrating an automated security verification process; Figure 5 This is a schematic diagram of the structure of a device for determining the master control terminal of a botnet, provided in the second embodiment of this disclosure; Figure 6 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation

[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0012] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0013] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0014] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0015] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0016] The method for determining the master controller of a botnet according to this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method for determining the master controller of a botnet provided in this disclosure can be executed by a server.

[0017] First, the key technical terms involved in each step are explained.

[0018] External communication address: refers to the target Internet Protocol (IP) address that is attacked in a network attack incident. In this disclosure, it is also referred to as "network address", "communication address", etc.

[0019] Attack traffic: refers to malicious network data streams initiated by a controlled host with the aim of compromising the availability, integrity, or confidentiality of a target system.

[0020] Controlled host set: refers to a set of multiple hosts within the same security domain (such as a cloud tenant environment or enterprise intranet) that are controlled by the same botnet and participate in the attack traffic.

[0021] Network flow data refers to information that records network communication metadata, which may include at least one of the following: timestamp, source and destination Internet Protocol addresses, port, protocol type, number of bytes transmitted, number of data packets, etc., but does not include the specific payload content of the communication.

[0022] Suspected master control server list: refers to the potential botnet command and control server output by the large language model for each controlled host, that is, the master control server of the botnet. It can be a structured result list, such as JSON format, which can contain the inferred address of the potential botnet command and control server and related supporting reasons descriptions, also known as "reason information for the suspected judgment".

[0023] Target botnet master control terminal: refers to the network address of the botnet command and control server or its cluster that is ultimately determined by this method and controls the set of controlled hosts participating in this attack.

[0024] In the first disclosed embodiment, see Figure 1 , Figure 1 This diagram illustrates a flowchart of a method for determining the master control terminal of a botnet according to a first embodiment of this disclosure. The method includes: S101. In response to detecting attack traffic targeting external communication addresses, determine the set of controlled hosts corresponding to the attack traffic, and extract network flow data of each controlled host in the set of controlled hosts within a preset attack tracing time window.

[0025] The attack attribution time window refers to a preset time range for accurately capturing communication data between the botnet's master control unit and the controlled host. Specifically, it can include a first attack attribution time window consisting of a preset first duration before the attack begins, and / or a second attack attribution time window consisting of a preset second duration after the attack begins. The attack begins when attack traffic targeting an external communication address is detected.

[0026] This step narrows the scope of analysis from massive amounts of data across the entire network to a small number of hosts and a limited time window that are highly relevant to specific attack events, providing precise input for subsequent in-depth analysis and reducing the amount of data input.

[0027] Specifically, when the security system detects attack traffic targeting a specific external communication address, such as a distributed denial-of-service attack exceeding a threshold, the system responds immediately. First, it resolves the source IP address of the attack from the attack traffic and filters out controlled hosts within its security jurisdiction, forming a set of controlled hosts for this event. Next, based on the attack's start time, an attack tracing time window is defined, for example, 30 minutes before the attack begins and 10 minutes after. The system automatically extracts all network flow data from each controlled host in this set within this attack tracing time window, using it as the raw data for analysis.

[0028] S102. For any controlled host in the set of controlled hosts, perform botnet feature preprocessing on the network flow data to generate host behavior description information of the controlled host.

[0029] Among them, botnet feature preprocessing of network stream data refers to data cleaning, filtering and feature enhancement operations specifically targeting the characteristics of botnet control channel behavior in order to make network stream data more suitable for subsequent semantic analysis.

[0030] For example, host behavior description information refers to natural language text paragraphs that are converted from preprocessed network flow data belonging to a single controlled host and can describe its outbound connection behavior in chronological order.

[0031] This step filters and semantically transforms the raw network flow data, filtering out noise from a large amount of normal business traffic, and converting structured log data into textual descriptions that large language models can understand and that contain security semantics, thus preprocessing the model for effective inference.

[0032] Specifically, for each controlled host in the controlled host set, botnet characteristic preprocessing is performed on the extracted network flow data. For example, the following preprocessing can be performed: filtering based on historical baselines, utilizing the baseline of compliant external connections established by each host during non-attack periods (e.g., compliant update servers, business databases, cloud service application programming interfaces, etc.) to filter out normal business connections matching the baseline, significantly reducing data volume; and heuristic filtering based on control channel characteristics. Since botnet attack commands typically involve small data streams (small packets), network flow data with excessively high average bandwidth can be filtered out. Furthermore, normal business connections are usually not limited to a single connection, thus filtering out connections established throughout the entire attack tracing process. Isolated connections appearing only once within the source time window; feature enhancement and semantic transformation are performed to calculate traffic statistics for each filtered network flow data, such as the payload-to-packet ratio (the ratio of total bytes to the size of the corresponding data packet), and / or connection duration category, including long connections and short connections. Finally, each network flow data and its traffic statistics are converted into a descriptive natural language sentence, for example: "[25 minutes before the attack] initiated a short connection to port 8080 of an overseas IP (104.xxx), sending a small payload of 160 bytes." Finally, all descriptive statements for a single host within the attack tracing time window are concatenated chronologically to generate a coherent host behavior description, i.e., a natural language paragraph. This step transforms messy data into a behavioral story description with temporal context.

[0033] S103. Input the host behavior description information of each controlled host into the large language model of the security domain, and guide it to perform semantic reasoning by using the prompt information configured for the feature analysis of botnets, and output a list of suspected master control terminals corresponding to each controlled host.

[0034] In the security field, large language models refer to large language models that have been trained or fine-tuned with knowledge related to cybersecurity (such as attack patterns, malware behavior, protocol analysis, etc.) and have the ability to understand cybersecurity-related texts and perform reasoning.

[0035] Prompts for botnet feature analysis refer to specially structured instructions or contextual information input to a large language model. Their purpose is to guide and constrain the model to analyze the input text from dimensions specific to the botnet control channel, such as periodicity, payload characteristics, and temporal logic.

[0036] Semantic reasoning refers to the process by which large language models, based on their understanding of the semantics of the input text and combined with domain knowledge, make logical inferences to identify hidden patterns, causal relationships, or intentions.

[0037] This step leverages the deep semantic understanding and logical reasoning capabilities of large language models in the security domain to identify hidden botnet control channel characteristics from host behavior description information that are difficult to detect using manual rules or traditional models.

[0038] The host behavior descriptions generated in step S102 are input into a large language model for the security domain. Simultaneously, the model is provided with hints constructed based on botnet characteristics. This hint is a structured template that explicitly requires the model to analyze from multiple botnet-specific dimensions, thus adapting to the semantic reasoning specific to the botnet scenario. Guided by this external knowledge, the large language model performs semantic reasoning on the input host behavior descriptions to understand the underlying intent behind the behavioral sequences. Finally, the model outputs a structured list of suspected master controllers for each controlled host, such as in JOSN format. The list includes the address of the control server it considers suspicious, the reasoning for the suspicion, and a confidence score. The reasoning for the suspicion is the reason for determining that the host is a legitimate master controller, expressed in natural language.

[0039] S104. Gather a list of suspected master controllers from each controlled host, perform attack correlation analysis, and determine the master controller of the target botnet based on the analysis results.

[0040] Among them, attack correlation analysis refers to the process of combining the analysis results of multiple controlled hosts, comparing and correlating them to find commonalities or behavioral similarities, thereby improving the accuracy of judgment.

[0041] The technical objective of this step is to overcome the limitations of analyzing a single host. By integrating the analysis results of multiple controlled hosts and analyzing the correlation between them, the real or clustered botnet master control terminals or clusters can be located with higher confidence.

[0042] The method disclosed herein utilizes a large language model in the security domain to perform semantic reasoning guided by prompts from botnet feature analysis. This enables a deep understanding of the complex interaction logic and temporal relationships between the controlled host and the master control unit from the network flow data of the controlled host, thereby achieving effective identification of botnet master control units that are disguised or use dynamic technologies. Furthermore, by aggregating the analysis results of multiple controlled hosts and performing correlation analysis, it can discover hidden control sources from group behavior, overcoming the problems of false detection and missed detection in single-point detection. This achieves accurate positioning of botnet master control units, effectively breaking through the limitations of traditional single-host detection, improving the targeting and semantic understanding capabilities of attack tracing, and reducing misjudgments caused by irrelevant data interference.

[0043] In some examples, S101 includes the attack attribution time window: A first attack tracing time window, from a preset first duration before the attack begins to the first attack tracing time window at the start of the attack; and / or, a second attack tracing time window, from the start of the attack to a preset second duration after the start of the attack.

[0044] The first attack tracing time window refers to the time range extending backward from the start of the attack by a preset first duration. The preset first duration refers to the pre-attack analysis duration based on the communication patterns of the botnet, and can be set as needed, for example, from 30 minutes to 2 hours. The second attack tracing time window refers to the time range extending backward from the start of the attack by a preset second duration. The preset second duration refers to the preset post-attack analysis duration, and can be set as needed, for example, from 10 minutes to 1 hour. Specifically, the first and second durations can be preset based on prior knowledge of botnet activity patterns or historical data analysis.

[0045] The attack attribution time window is designed to accurately capture critical communication data between the botnet's master control unit and the controlled host. Its implementation can include three scenarios: using only the first attack attribution time window, using only the second attack attribution time window, and using both the first and second attack attribution time windows simultaneously. The first attack attribution time window focuses on communication data prior to the attack, as the botnet's master control unit typically sends attack commands or confirms the liveness status of the controlled host some time before initiating the attack. The second attack attribution time window covers communication data after the attack, capturing any subsequent commands or status feedback data that the master control unit may send after the attack. By flexibly configuring the attack attribution time window, it is possible to ensure that no critical communication clues are missed, based on the characteristics of different attack scenarios.

[0046] In S102, botnet feature preprocessing refers to a series of processing operations performed on network flow data to filter out network flow data with analytical value. It includes two methods: data flow filtering preprocessing and network flow data feature enhancement preprocessing. Each method has multiple specific implementation forms, and these methods can be used individually or in combination. The execution order can also be flexibly adjusted according to the actual scenario without any mandatory restrictions. Data flow filtering preprocessing refers to filtering network flow data through specific rules to remove normal or invalid data without analytical value; network flow data feature enhancement preprocessing refers to improving the semantic expressive power of network flow data by calculating specific statistical features. The specific implementation forms of each preprocessing method are described in detail below.

[0047] In some examples, as an implementation of data stream filtering preprocessing, S102 includes: Step A1: Obtain the historical behavior baseline of each controlled host in the controlled host set, and match the outbound connections in the network flow data of the controlled host with their corresponding historical behavior baselines. The historical behavior baselines include the compliant external connection objects of the controlled hosts.

[0048] The historical behavior baseline refers to the behavior benchmark built based on the normal business traffic of the controlled host over a preset historical period. It includes the host's compliant external connection objects, which are the connection targets required for normal business operations, such as the server address of business partners, the service address of cloud vendors, the system update server address, etc. The preset historical period can be set as needed, such as 10 days, 20 days, etc., and is not limited here.

[0049] Step A2: Filter out the network flow data corresponding to the successfully matched connection records from the network flow data to obtain the first network flow data.

[0050] Among them, "first network flow data" refers to network flow data suspected of communicating with the botnet master control terminal that is retained after filtering by historical behavior baseline matching.

[0051] Step A3: Generate host behavior description information based on the first network flow data of the controlled host.

[0052] In the above method, step A1 involves pre-obtaining the historical behavior baseline for each controlled host in the controlled host set, and matching all outbound connections in the network flow data of that host within the attack tracing time window with the compliant outbound connections in the historical behavior baseline one by one. Step A2 involves removing the network flow data corresponding to the successfully matched connection records from the network flow data, i.e., normal business connection data, and the remaining unmatched network flow data is the first network flow data. Step A3 involves generating host behavior description information for the controlled host based on the filtered first network flow data. The purpose of this method is to filter normal business traffic through the historical behavior baseline, accurately filter out unfamiliar and suspicious outbound connection data, reduce the interference of normal business data on subsequent analysis, and reduce the amount of data processing.

[0053] In some examples, as an implementation of data stream filtering preprocessing, S102 includes: Step B1: Filter out network flow data whose average data transmission rate exceeds a preset bandwidth threshold, and / or filter out network flow data that appears less than a preset number of times within the attack tracing time window and has no reconnection behavior.

[0054] The preset bandwidth threshold refers to the average data transmission rate threshold preset based on the characteristic that botnet control commands are usually low-volume data. For example, it can be configured to be between 0.5Mbps and 5Mbps. The preset number of attempts refers to the threshold used to determine whether a connection is an isolated and invalid connection, which can be configured according to the actual scenario, for example, 1 or 2 times. No reconnection behavior means that within the attack tracing time window, a certain connection only appears once and does not reconnect.

[0055] Step B2: Obtain second network flow data based on the filtered network flow data, and generate host behavior description information based on the second network flow data of the controlled host.

[0056] The second network flow data refers to the network flow data obtained after bandwidth filtering and isolated connection filtering.

[0057] In the above method, step B1 involves dual filtering of the network flow data of the controlled host: first, filtering out network flow data with an average data transmission rate exceeding a preset bandwidth threshold, because botnet control commands are mostly low-bandwidth data, while high-bandwidth data is usually normal business data; second, filtering out network flow data that appears less than a preset number of times within the attack tracing time window and shows no reconnection behavior, because such isolated connections are mostly invalid interference data and do not possess the persistent characteristics of botnet communication. These two filtering operations can be performed individually or simultaneously. Step B2 involves determining the remaining network flow data after filtering as the second network flow data, and generating host behavior description information for the controlled host based on this second network flow data. The purpose of this method is to further filter out suspicious data that conforms to the characteristics of botnet communication through bandwidth and connection persistence characteristics, thereby improving the accuracy of data filtering and reducing the amount of data processing.

[0058] In some examples, as an implementation of network stream data feature enhancement preprocessing, S102 includes: Step C1: For any network flow data, calculate its traffic statistics characteristics, wherein the traffic statistics characteristics include at least one of the following: the ratio of the total number of bytes to the size of the corresponding data packet of the network flow data, and the connection category characteristics based on the connection duration of the network flow data.

[0059] Among them, traffic statistics features refer to statistical indicators that can reflect the characteristics of network flow data communication, which are used to enhance the semantic expression of host behavior description information; the ratio of total bytes to the size of the corresponding data packet of network flow data (i.e., payload ratio feature) refers to the ratio of the total number of bytes of a single network flow data to the size of the data packet in that flow, which is used to determine whether the data packet is a small payload instruction packet, that is, an instruction packet that conforms to the attack instructions issued by the botnet; the connection category feature based on the connection duration of network flow data refers to the feature of classifying network flow into long connections or short connections according to the connection duration. For example, a long connection refers to a connection whose duration exceeds the preset connection duration, and a short connection refers to a connection whose duration does not exceed the preset connection duration. The preset connection duration can be set as needed, for example, 30 seconds.

[0060] In this step, the network stream data being processed can be either the first network stream data or the second network stream data; or the second network stream data can be data that has been processed by the first network stream data and then further processed. The specific combination of processing the second network stream data in this step can be flexibly combined according to the actual application scenario, and is not limited here.

[0061] Step C2: Based on the traffic statistics characteristics of the network flow data of each controlled host, generate host behavior description information of the controlled host.

[0062] In the above method, step C1 involves calculating at least one of the aforementioned traffic statistical features for the network flow data to be processed. The network flow data to be processed can be flexibly selected; it can be the raw network flow data without data flow filtering preprocessing, the first network flow data obtained after processing in steps A1-A3, or the second network flow data obtained after processing in steps B1-B2. If both data flow filtering preprocessing and feature enhancement preprocessing of this form are used simultaneously, the second network flow data can be obtained by further processing the first network flow data through steps B1-B2. This application does not limit the preprocessing steps for data processing. Step C2 involves generating host behavior description information for the controlled host based on the calculated traffic statistical features and other core information of the network flow data. The purpose of this method is to supplement the dimensions of host behavior description information by extracting key statistical features, enabling subsequent large language models to more accurately identify the specific characteristics of botnet communication.

[0063] It should be noted that the above three preprocessing methods can be flexibly combined according to the actual application scenario. For example, first execute steps A1-A3 to obtain the first network stream data, then execute steps B1-B2 to obtain the second network stream data, and finally execute steps C1-C2 to perform feature enhancement; or directly execute steps C1-C2 to perform feature enhancement on the original network stream data. The execution order of each method can also be freely interchanged, and no limitation is made here.

[0064] In some examples, the host behavior description information in S102 is in natural language form. Natural language form is better suited to the semantic understanding characteristics of large language models in the security domain, facilitating accurate inference by subsequent models. Therefore, the host behavior description information of the controlled host generated in S102 includes: Sub-step 1: Convert each preprocessed network stream data into a natural language sentence describing the network stream data.

[0065] Among them, natural language form refers to a text format that is human-understandable and conforms to the logic of language expression, which is different from the structured log format of network stream data; preprocessed network stream data refers to network stream data obtained after the aforementioned filtering preprocessing and / or network stream data feature enhancement preprocessing, such as various types of suspicious network stream data, including first network stream data, second network stream data, etc.; natural language sentence refers to an independent text statement formed by refining and transforming the core information of a single preprocessed network stream data.

[0066] Sub-step two: For any controlled host, concatenate all natural language sentences of the controlled host within the attack tracing time window according to the timestamp order to form a host behavior log segment in natural language form.

[0067] Among them, the timestamp refers to the time mark that records the time when the network stream data is generated, which can reflect the chronological order of the network communication behavior of the controlled host; the host behavior log segment refers to the host behavior description information in natural language form, which is a complete text segment formed by splicing multiple natural language sentences in chronological order.

[0068] The natural language sentence may specifically include at least one of the following: time information, referring to the specific time when the network connection occurred; destination address, referring to the destination IP address of the network connection; port information, referring to the destination port number used by the network connection; packet size description, referring to a quantitative description in natural language based on the total number of bytes or average packet size of a network flow, such as "sent 160 bytes of small payload data" or "performed a large volume transmission"; and connection status description, referring to a general description of the basic characteristics of the network connection, such as "established a long connection" or "connection was reset".

[0069] For example, the data record of the network stream is: {Time:10:05:01,Dst:104.xxx,Port:8080,Bytes:160,Pkts:3}; the converted natural language sentence is: [Time:-25m] Initiated a connection to port 8080 of an overseas IP (104.xxx), sent a small payload of data (160B), and the connection was quickly disconnected.

[0070] The above steps, for any one of the controlled hosts, collect all the aforementioned natural language sentences generated by the controlled host within the attack tracing time window. Then, based on the timestamps corresponding to each natural language sentence, and in chronological order from earliest to latest, all the natural language sentences are sequentially concatenated to form a complete text segment. This text segment is the host behavior description information in natural language form, i.e., the host behavior log segment. This chronological arrangement preserves the sequence and interval information of events, forming a behavior description with a continuous temporal context. It reconstructs the overall network communication behavior trajectory of the controlled host within the attack tracing time window, thus enabling subsequent large language models in the security field to capture long-term temporal dependencies.

[0071] Through the above transformation and splicing process, the original non-semantic log data is constructed into natural language text that includes time, action, object, and state information. This adapts to the semantic understanding capabilities of large language models, enabling them to be directly read and understood by large language models in the security field. This allows them to use natural language processing capabilities to analyze the logical relationships between behaviors, identify hidden patterns, and thus perform subsequent semantic reasoning tasks.

[0072] See in some examples Figure 2 , Figure 2 The diagram shows the composition of the structured template 01. The prompt information for the feature analysis of botnets in S103 is a structured prompt template. The structured prompt template is a pre-set text frame or instruction set with a specific organizational structure. Its core function is not only to transmit data, but also to explicitly guide and constrain the reasoning direction of the large language model, guiding it to input host behavior description information from the feature analysis of botnet control behavior.

[0073] Among them, the feature dimension of the botnet refers to the core judgment angle used to identify the communication behavior of the botnet, and is the core component of the structured prompt template; structured data refers to data organized according to a preset field format, which is convenient for subsequent correlation analysis and machine reading, such as JSON format.

[0074] For example, the structured prompt template is configured to guide the large language model to analyze at least one of the following feature dimensions of the botnet in sequence: Heartbeat jitter feature dimension 001 is configured to identify whether the connection interval of the controlled host fluctuates around a fixed period value, and the fluctuation is within a preset fluctuation range.

[0075] In this context, the heartbeat jitter feature dimension 001 is configured to guide the model in identifying whether the outward connections of the controlled host exhibit pseudo-periodicity. Specifically, the model is instructed to check whether the time intervals of consecutive connections exhibit a pattern of fluctuation around a fixed periodic value, such as 60 seconds, and whether this fluctuation is controlled within a preset range, such as ±15%. This "regular randomness" is a typical means for botnets to circumvent detection rules based on strict periods.

[0076] The payload fingerprint feature dimension 002 is configured to determine whether the similarity of the packet sizes of multiple connections of the controlled host is higher than a preset similarity threshold, wherein the multiple connections are for the same external communication address or different external communication addresses.

[0077] Among them, the payload fingerprint feature dimension 002 indicates whether the data packet sizes transmitted by the controlled host initiating multiple connections at different times show a high degree of similarity, that is, the similarity is higher than a preset similarity threshold. The preset similarity threshold refers to a pre-set threshold value for data packet size matching, which can be set as needed, for example, 90%. It should be noted that the analysis of this dimension requires the model to ignore changes in the target address. That is, regardless of whether multiple connections are for the same external communication address or different addresses, their payload sizes should be compared. Stable payload sizes are usually fingerprints of specific instruction codes or protocol data units.

[0078] The behavioral mutation dimension 003 before and after the attack is configured to analyze the characteristics of the controlled host's connections within a preset time before the attack begins, and to identify whether there is any behavior under the attack command.

[0079] The dimension 003, which describes behavioral mutations before and after an attack, is configured to guide the model to focus on key behavioral changes in the controlled host before and after the attack begins. It instructs the model to analyze whether the characteristics of the last one or more connections of the controlled host undergo abrupt changes within a preset time period before the attack begins. The preset time period refers to the key observation period before the attack begins and can be set as needed, such as 1 minute. Connection characteristics include sudden changes in packet size, a surge in connection frequency, etc., such as changing from a stable small heartbeat packet to a slightly larger or structurally different packet. Such mutations are often associated with the issuance of attack commands. The botnet master usually issues attack trigger commands to the controlled host before the attack, and these commands cause abnormal mutations in the controlled host's connection behavior. The model can use this dimension to identify such key communications.

[0080] Port anomaly dimension 004 determines whether the target port connected by the controlled host deviates from the normal port range corresponding to the preset business role of the controlled host.

[0081] Among them, port anomaly dimension 004 is configured to guide the model to make reasonable judgments based on the host's business role, instructing the model to evaluate whether the target port connected by the controlled host deviates significantly from the normal port range typically used by the host's established business role. For example, if a host identified as a database server frequently and actively connects to non-standard external ports, it is considered abnormal behavior.

[0082] In some examples, see further. Figure 2 The structured template 01 may also include: an attack event context field 005, used to inject information about the start time of the attack and the business role information of the controlled host in the attack event corresponding to the attack traffic into the large language model. In some cases, the attack event context field 005 can be omitted, and this is not a limitation here.

[0083] The attack event context field 005 is used to inject information about the attack start time and the business role information of the controlled host in the attack event corresponding to the attack traffic into the large language model. Injecting the attack start time information helps the model identify key time nodes and accurately analyze the correlation between communication behaviors before and after those nodes; injecting the business role information of the controlled host allows the model to judge the rationality of connection behavior based on role attributes. Business roles include web servers, database servers, etc., further improving the accuracy of inference.

[0084] In some examples, the list of suspected masterminds output by the large language model is structured data, including the network address of the suspected mastermind and suspected descriptive information. The network address of the suspected mastermind is the IP address or domain name of the potential botnet command and control server inferred by the model. The suspected descriptive information includes at least one of the following fields: controlled host identifier, indicating which controlled host the judgment is for; reasoning information for the suspected judgment, in natural language, indicating which feature dimensions or dimensions the model used to arrive at the suspected conclusion; and confidence score, a numerical rating used to represent the model's degree of certainty regarding the judgment. This approach provides standardized input data for subsequent multi-source association analysis, effectively improving the accuracy and efficiency of botnet mastermind identification.

[0085] In some examples, the attack correlation analysis in S104 can include multiple methods, such as target overlap correlation analysis and / or behavioral similarity correlation analysis. Of course, other analysis methods can also be used, which are not limited here. Among them, target overlap correlation analysis refers to determining the analysis method of the target master terminal by statistically analyzing the frequency of occurrence of the suspected master terminal's network address in multiple controlled hosts; behavioral similarity correlation analysis refers to determining the analysis method of the target master terminal by converting the relevant information of the suspected master terminal into feature vectors, calculating the vector similarity corresponding to different controlled hosts, and then determining the analysis method of the target master terminal.

[0086] In one feasible approach, attack correlation analysis includes: target overlap correlation analysis, which includes: Step 1: Calculate the intersection of the suspected master control network address sets of different controlled hosts.

[0087] For each controlled host in the controlled host set, extract all suspected master control network addresses from its suspected master control list to form a set of suspected master control network addresses for a single controlled host; then, calculate the intersection of the address sets of all controlled hosts, that is, find the suspected master control network addresses that appear together in the address sets of multiple controlled hosts.

[0088] Step 2: If any suspected master control network address appears in the list of suspected master control hosts of more than a preset threshold number of controlled hosts, it is marked as the master control host of the target botnet.

[0089] The frequency of occurrence of each commonly appearing suspected master control network address (i.e., the number of controlled hosts containing that address) is compared with a preset threshold. If the frequency of any suspected master control network address exceeds the preset threshold, that network address is directly marked as the target botnet master control address. The underlying logic is that the probability of multiple unrelated hosts coincidentally connecting to the same suspicious external address before and after an attack is extremely low; this address is highly likely to be a common source of command and control.

[0090] The preset threshold number refers to the minimum frequency of occurrence of the target master control terminal that is pre-set and used to determine the target master control terminal. It can be configured according to the actual scenario, for example, 2 to 5 units.

[0091] In one feasible approach, attack correlation analysis includes behavioral similarity correlation analysis, which is better suited to scenarios where attackers use dynamically changing addresses (such as Fast-Flux techniques). Behavioral similarity correlation analysis includes: Step 1: For the list of suspected master control terminals, use a semantic embedding model to convert them into behavioral feature vectors; Among them, semantic embedding models refer to models that can transform text information into high-dimensional numerical vectors, and the output vectors can accurately represent the semantic features of the text.

[0092] Behavioral feature vectors refer to high-dimensional numerical vectors obtained after processing the list of suspected master control terminals through a semantic embedding model. They are used to quantitatively characterize the communication behavior features between the controlled host and the suspected master control terminal.

[0093] Step 2: Calculate the behavioral similarity between the behavioral feature vectors corresponding to different controlled hosts; One approach is to use the cosine similarity algorithm to calculate the behavioral similarity between the behavioral feature vectors of different controlled hosts.

[0094] Step 3: Determine the suspected master control terminal connected to each controlled host as the master control terminal of the target botnet based on the similarity of their behaviors.

[0095] Behavioral similarity correlation analysis targets scenarios where attackers evade detection by switching suspected master control network addresses, such as dynamic address switching and cluster control. It determines correlations based on behavioral feature consistency. Specifically: for each controlled host's suspected master control list, information is extracted from the list and input into a semantic embedding model. Through model processing, the textual information is transformed into high-dimensional behavioral feature vectors that quantify behavioral characteristics. The behavioral similarity between the behavioral feature vectors corresponding to different controlled hosts is calculated; this similarity quantifies the semantic closeness of suspicious behavioral patterns across different hosts. The system makes a judgment based on the calculated behavioral similarity results. The judgment criteria are: if two controlled hosts are connected to different suspected master control network addresses, but the behavioral similarity between their corresponding behavioral feature vectors exceeds a preset similarity threshold (e.g., 0.95), and these suspicious connection behaviors occur just before the attack, then the system will determine that these two different network addresses are highly likely to belong to the same target botnet master control cluster, for example, different entry nodes under the same control infrastructure or different resolved IPs under the same Fast-Flux domain name.

[0096] This method enables intelligent association of different IPs with the same behavioral patterns (i.e., originating from the same source). This approach can accurately identify the master control terminal of a botnet with dynamic addresses and clustered control.

[0097] It should be noted that target overlap correlation analysis and behavior similarity correlation analysis can be used alone or in combination. This method can flexibly and effectively deal with different strategies of static or dynamic addresses used by the botnet master control end, thereby ensuring accurate source tracing under various conditions.

[0098] See in some examples Figure 3 , Figure 3 This diagram illustrates another method for determining the master control terminal of a botnet according to the first embodiment of this disclosure. After S104, the method provided by this disclosure further includes: S105. Perform automated security verification on the target botnet's master control terminal.

[0099] Among them, automated security verification refers to the process of automatically verifying the malicious attributes of the target botnet master control terminal through preset verification methods, which can improve verification efficiency and accuracy.

[0100] S106. Execute corresponding cybersecurity measures based on the verification results.

[0101] Among them, network security response actions refer to network protection operations performed based on automated security verification results in order to block attacks and eliminate threats.

[0102] See in some examples Figure 4 , Figure 4 The flowchart illustrating automated security verification is shown, with S105 including: S1051. Perform a passive Domain Name System (DNS) query on the network address of the target botnet master control terminal to obtain its historically associated domain name information and verify whether the historically associated domain name information includes associated malicious domain names. And / or, S1052. Perform a port scan on the network address of the target botnet master control terminal to determine whether it has opened malicious ports related to the known botnet master control terminal.

[0103] Passive DNS lookup refers to obtaining information about domain names that an Internet Protocol address has been associated with by querying a third-party database of historical DNS records, rather than by directly requesting the target. Historically associated domain name information refers to all domain name records that the target botnet's control address has been associated with through DNS resolution over a period of time. Associated malicious domain names refer to domain names that have been marked by security agencies or that are associated with known botnets or malicious attack activities. Port scanning refers to probing the ports corresponding to the target botnet's control address to determine whether the ports are open. Malicious ports refer to ports that are commonly used by the botnet's control address for issuing control commands.

[0104] Regarding S1051, in specific implementation, the system initiates a passive Domain Name System (DNS) query on the network address of the target botnet's master control terminal to retrieve which domain names have historically resolved to that address. Subsequently, the system verifies whether these historically associated domain name information contains known or highly suspicious associated malicious domain names. If such associations exist, it corroborates the malicious nature of the target botnet's master control terminal's network address.

[0105] For S1052, a port scan is performed on the network address of the target botnet master control terminal to detect whether the various ports corresponding to that address are open. The open ports are then matched with a list of commonly used malicious ports by known botnet master control terminals to determine if any open malicious ports exist. The technical purpose of this verification method is to further confirm the malicious attributes of the target through port characteristics. Botnet master control terminals typically keep specific ports open for communication with controlled hosts, which is an important characteristic that distinguishes them from normal network addresses.

[0106] In some examples, S106 includes: At the network gateway of the security domain, access control list rules are issued to block bidirectional network communication between the controlled host set and the target botnet master control terminal.

[0107] Among them, a security domain refers to a logical area in the network divided according to security requirements, used to achieve isolation and protection of different security levels; an access control list rule refers to a set of rules configured on the network gateway to allow or prohibit communication between specific network addresses; bidirectional network communication refers to the communication between the controlled host set and the target botnet master control terminal, including both connections initiated by the controlled host to the master control terminal and connections in which the master control terminal issues instructions to the controlled host.

[0108] In practice, the system automatically generates and issues one or more access control list rules at the network gateway of the security domain, i.e., a key node at the network boundary. The content of the access control list rules may include: blocking bidirectional network communication between all hosts in the controlled host set and the target botnet's master control network address, thereby immediately cutting off the botnet's control channel and data leakage channel.

[0109] In some examples, S106 may also include: At least a portion of the controlled hosts in the controlled host set are migrated to an isolated network environment for security testing and cleanup.

[0110] An isolated network environment refers to a dedicated network area that is independent of the normal business network and has security detection capabilities. It is used to clean up threats from infected hosts and prevent the threats from spreading to the normal network.

[0111] The system can migrate all or at least some of the controlled hosts in a controlled host set to an isolated network environment, either entirely or in batches. Within this isolated network environment, in-depth security testing and cleanup can be performed without disruption to business operations, completely removing malware.

[0112] The method disclosed herein filters suspicious traffic through multi-dimensional preprocessing, including historical behavior baselines, bandwidth, and reconnection behavior. It guides large-scale models to make accurate inferences using structured prompt templates containing botnet-specific feature dimensions. Combining target overlap and behavior vector similarity dual-dimensional correlation analysis, it can not only identify botnet master control terminals with fixed IPs but also identify master control terminals or master control terminal clusters corresponding to different IPs. It can effectively deal with various countermeasures such as C2 servers using fixed or dynamic IPs, and accurately determine botnet master control terminals and their clusters. Subsequent automated verification through passive DNS queries and port scanning, as well as handling actions such as gateway blocking and host isolation, significantly improve the efficiency, accuracy, and automation level of botnet attack tracing.

[0113] In the disclosed second embodiment, see Figure 5 ,like Figure 1 The principle shown Figure 5 This disclosure shows a device 50 for determining the master control terminal of a botnet according to a second embodiment, wherein the device includes: The host identification module 501 is used to identify the set of controlled hosts corresponding to the attack traffic in response to the detection of attack traffic targeting external communication addresses, and to extract the network flow data of each controlled host in the set of controlled hosts within a preset attack tracing time window. The host description module 502 is used to perform botnet feature preprocessing on network flow data for any controlled host in the controlled host set, and generate host behavior description information of the controlled host. The master control reasoning module 503 is used to input the host behavior description information of each controlled host into the large language model of the security domain, and guide it to perform semantic reasoning through the prompt information configured for the feature analysis of botnets, and output a list of suspected master control terminals corresponding to each controlled host. The master control determination module 504 is used to aggregate a list of suspected master control terminals from various controlled hosts, perform attack correlation analysis, and determine the target botnet master control terminal based on the analysis results.

[0114] In some examples, the host description module 502 is specifically used for: Obtain the historical behavior baseline of each controlled host in the controlled host set, and match the outbound connections in the network flow data of the controlled host with their corresponding historical behavior baselines. The historical behavior baselines include the compliant outbound connections of the controlled hosts. Filter out the network flow data corresponding to the successfully matched connection records from the network flow data to obtain the first network flow data; Host behavior description information is generated based on the first network flow data of the controlled host.

[0115] In some examples, the host description module 502 is specifically used for: Filter out network flow data whose average data transmission rate exceeds a preset bandwidth threshold, and / or filter out network flow data that appears less than a preset number of times within the attack tracing time window and has no reconnection behavior. The second network flow data is obtained from the filtered network flow data, and host behavior description information is generated based on the second network flow data of the controlled host.

[0116] In some examples, the host description module 502 is specifically used for: For any network flow data, calculate its traffic statistics characteristics, wherein the traffic statistics characteristics include at least one of the following: the ratio of the total number of bytes to the size of the corresponding data packet of the network flow data, and the connection category characteristics based on the connection duration of the network flow data; Based on the traffic statistics characteristics of the network flow data of each controlled host, host behavior description information of the controlled host is generated.

[0117] In some examples, the host behavior description information is in natural language form; the host description module 502, when generating host behavior description information for the controlled host, is specifically used for: Each preprocessed network stream data is converted into a natural language sentence describing the network stream data; For any controlled host, all natural language sentences of the controlled host within the attack tracing time window are concatenated in timestamp order to form a host behavior log segment in natural language form.

[0118] In some examples, the prompts for botnet feature analysis in the master inference module 503 are structured prompt templates. These structured prompt templates are configured to guide the large language model to analyze at least one of the following botnet feature dimensions in sequence: The heartbeat jitter feature dimension is configured to identify whether the connection interval of the controlled host fluctuates around a fixed period value, and the fluctuation is within a preset fluctuation range. The payload fingerprint feature dimension is configured to determine whether the similarity of the packet sizes of multiple connections of the controlled host is higher than a preset similarity threshold, wherein multiple connections are for the same external communication address or different external communication addresses. The dimension of behavioral change before and after the attack is configured to analyze the characteristics of the connection of the controlled host within a preset time before the start of the attack, and to identify whether there is behavior under the attack command. The port anomaly dimension determines whether the target port connected by the controlled host deviates from the normal port range corresponding to the preset business role of the controlled host.

[0119] In some examples, the structured hint template also includes an attack event context field, used to inject information about the start time of the attack into the large language model and the business role information of the controlled host in the attack event corresponding to the attack traffic.

[0120] In some examples, the list of suspected master control devices is structured data, including the network address of the suspected master control device and suspected description information. The suspected description information includes at least one of the following fields: controlled host identifier, reason for the suspected determination, and confidence score.

[0121] In some examples, the attack correlation analysis in the master control determination module 504 includes: target overlap correlation analysis, which includes: Analyze the intersection of the suspected master control network address sets of different controlled hosts; If any suspected master control network address appears in the list of suspected master control hosts of more than a preset threshold number of controlled hosts, it is marked as the master control host of the target botnet.

[0122] In some examples, the attack correlation analysis in the master control determination module 504 includes: behavioral similarity correlation analysis, which includes: For the list of suspected master control terminals, a semantic embedding model is used to convert them into behavioral feature vectors; Calculate the behavioral similarity between behavioral feature vectors corresponding to different controlled hosts; The suspected master control terminal connected to each controlled host is identified as the master control terminal of the target botnet based on the similarity of the behaviors of each controlled host.

[0123] In some examples, when the master control determination module 504 determines, based on the behavioral similarity of each controlled host, that the suspected master control terminal it connects to is the master control terminal of the target botnet, it is used to: If two controlled hosts connected to different suspected master control network addresses have behavioral feature vector similarity exceeding a preset similarity threshold and their communication time is just before the attack, then it is determined that the different suspected master control network addresses belong to the same target botnet master control cluster.

[0124] In some examples, when the master control determination module 504 calculates the behavioral similarity between behavioral feature vectors corresponding to different controlled hosts, it is used for: The cosine similarity algorithm is used to calculate the behavioral similarity between behavioral feature vectors corresponding to different controlled hosts.

[0125] In some examples, the attack attribution time window includes: A first attack tracing time window, from a preset first duration before the attack begins to the first attack tracing time window at the start of the attack; and / or, a second attack tracing time window, from the start of the attack to a preset second duration after the start of the attack.

[0126] In some examples, the device also includes: The security verification module is used to perform automated security verification on the target botnet's master control terminal; The security handling module is used to perform corresponding cybersecurity handling actions based on the verification results.

[0127] In some examples, the security verification module is specifically used to: perform a passive Domain Name System query on the network address of the target botnet master control terminal, obtain its historically associated domain name information, and verify whether the historically associated domain name information includes associated malicious domain names; And / or, perform a port scan on the network address of the target botnet master control terminal to determine whether it has opened malicious ports related to the known botnet master control terminal.

[0128] In some examples, the security handling module is specifically used to: issue access control list rules at the network gateway of the security domain to block bidirectional network communication between the controlled host set and the target botnet master control terminal.

[0129] In some examples, the security handling module is specifically used to migrate at least a portion of the controlled hosts in the controlled host set to an isolated network environment for security testing and cleanup.

[0130] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0131] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0132] Computer instructions stored on a non-transitory computer-readable storage medium are used to cause a computer to perform the above-described method.

[0133] Computer program products include computer programs that, when executed by a processor, implement the methods described above.

[0134] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0135] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from a storage unit into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0136] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0137] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the methods described above. For example, in some embodiments, the methods described above can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).

[0138] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0139] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0142] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0143] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0144] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0145] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for determining the master control terminal of a botnet, wherein, The method includes: In response to the detection of attack traffic targeting external communication addresses, the set of controlled hosts corresponding to the attack traffic is determined, and network flow data of each controlled host in the set of controlled hosts is extracted within a preset attack tracing time window. For any one of the controlled hosts in the set of controlled hosts, the network flow data is preprocessed with botnet features to generate host behavior description information of the controlled host; The host behavior description information of each of the controlled hosts is input into a large language model in the security field, and semantic reasoning is guided by the prompt information configured for the feature analysis of botnets configured for the large language model, and a list of suspected master control terminals corresponding to each controlled host is output. The list of suspected master controllers of each of the controlled hosts is aggregated, and attack correlation analysis is performed to determine the master controller of the target botnet based on the analysis results.

2. The method according to claim 1, wherein, The step of performing botnet feature preprocessing on the network flow data for any one of the controlled hosts in the set of controlled hosts to generate host behavior description information for the controlled host includes: Obtain the historical behavior baseline of each of the controlled hosts in the controlled host set, and match the outbound connections in the network flow data of the controlled hosts with their corresponding historical behavior baselines, wherein the historical behavior baselines include the compliant outbound connection objects of the controlled hosts; Filter out the network flow data corresponding to the successfully matched connection records from the network flow data to obtain the first network flow data; The host behavior description information is generated based on the first network flow data of the controlled host.

3. The method according to claim 1 or 2, wherein, The step of performing botnet feature preprocessing on the network flow data for any one of the controlled hosts in the set of controlled hosts to generate host behavior description information for the controlled host includes: Filter out network flow data whose average data transmission rate exceeds a preset bandwidth threshold from the network flow data, and / or filter out network flow data that appears less than a preset number of times within the attack tracing time window and has no reconnection behavior from the network flow data; The second network flow data is obtained based on the filtered network flow data, and the host behavior description information is generated based on the second network flow data of the controlled host.

4. The method according to any one of claims 1-3, wherein, The step of performing botnet feature preprocessing on the network flow data for any one of the controlled hosts in the set of controlled hosts to generate host behavior description information for the controlled host includes: For any of the aforementioned network stream data, its traffic statistics characteristics are calculated, wherein the traffic statistics characteristics include at least one of the following: the ratio of the total number of bytes to the size of the corresponding data packet of the network stream data, and the connection category characteristics based on the connection duration of the network stream data; Based on the traffic statistics characteristics of the network flow data of each of the controlled hosts, host behavior description information of the controlled hosts is generated.

5. The method according to any one of claims 1-4, wherein, The host behavior description information is in natural language form; The generation of host behavior description information for the controlled host includes: Each preprocessed network stream data is converted into a natural language sentence describing the network stream data; For any of the controlled hosts, all the natural language sentences of the controlled host within the attack tracing time window are concatenated in timestamp order to form the host behavior log segment in natural language form.

6. The method according to any one of claims 1-5, wherein, The prompts for feature analysis of botnets are structured prompt templates, which are configured to guide the large language model to analyze at least one of the following feature dimensions of botnets in sequence: The heartbeat oscillation feature dimension is configured to identify whether the connection interval of the controlled host fluctuates around a fixed period value, and the fluctuation is within a preset fluctuation range. The payload fingerprint feature dimension is configured to determine whether the similarity of the packet sizes of multiple connections of the controlled host is higher than a preset similarity threshold, wherein the multiple connections are for the same external communication address or different external communication addresses. The dimension of behavioral change before and after the attack is configured to analyze the characteristics of the connection of the controlled host within a preset time before the start of the attack, and to identify whether there is behavior under the attack command. The port anomaly dimension determines whether the target port connected to the controlled host deviates from the normal port range corresponding to the preset service role of the controlled host.

7. The method according to claim 6, wherein, The structured prompt template also includes an attack event context field, used to inject information about the start time of the attack and the business role information of the controlled host in the attack event corresponding to the attack traffic into the large language model.

8. The method according to any one of claims 1-7, wherein, The list of suspected master control terminals is structured data, including the network address of the suspected master control terminal and suspected description information. The suspected description information includes at least one of the following fields: controlled host identifier, reason information for suspected judgment, and confidence score.

9. The method according to any one of claims 1-8, wherein, The attack correlation analysis includes: target overlap correlation analysis, which includes: Calculate the intersection of the suspected master control network address sets of different controlled hosts; If any suspected master control network address appears in the list of suspected master control hosts of more than a preset threshold number of controlled hosts, it is marked as the master control host of the target botnet.

10. The method according to any one of claims 1-9, wherein, The attack correlation analysis includes: behavioral similarity correlation analysis, which includes: For the list of suspected master control terminals, a semantic embedding model is used to convert them into behavioral feature vectors; Calculate the behavioral similarity between the behavioral feature vectors corresponding to different controlled hosts; The suspected master control terminal connected to each of the controlled hosts is determined to be the master control terminal of the target botnet based on the similarity of their behaviors.

11. The method according to claim 10, wherein, The step of determining the suspected master control terminal connected to each of the controlled hosts as the master control terminal of the target botnet based on the behavioral similarity of each of the controlled hosts includes: If two controlled hosts connected to different suspected master control network addresses have a behavioral feature vector similarity exceeding a preset similarity threshold and their communication time is just before the attack, then it is determined that the different suspected master control network addresses belong to the same target botnet master control cluster.

12. The method according to claim 10 or 11, wherein, The calculation of behavioral similarity between the behavioral feature vectors corresponding to different controlled hosts includes: The cosine similarity algorithm is used to calculate the behavioral similarity between the behavioral feature vectors corresponding to different controlled hosts.

13. The method according to any one of claims 1-12, wherein, The attack attribution time window includes: The first attack tracing time window, which is a preset first duration before the attack starts and extends to the attack starts; and / or, the second attack tracing time window, which is a preset second duration after the attack starts.

14. The method according to any one of claims 1-13, wherein, After aggregating the suspected master control lists of each of the controlled hosts, performing attack correlation analysis, and determining the target botnet master control based on the analysis results, the method further includes: Perform automated security verification on the target botnet's master control terminal; Based on the verification results, appropriate cybersecurity actions will be taken.

15. The method according to claim 14, wherein, The automated security verification of the target botnet's master control terminal includes: A passive domain name system query is performed on the network address of the target botnet master control terminal to obtain its historically associated domain name information, and to verify whether the historically associated domain name information includes associated malicious domain names; And / or, perform a port scan on the network address of the target botnet master control terminal to determine whether it has opened malicious ports related to known botnet master control terminals.

16. The method according to claim 14 or 15, wherein, Based on the verification results, corresponding cybersecurity actions will be taken, including: At the network gateway of the security domain, an access control list rule is issued to block bidirectional network communication between the controlled host set and the target botnet master control terminal.

17. The method according to any one of claims 14-16, wherein, Based on the verification results, corresponding cybersecurity actions will be taken, including: At least a portion of the controlled hosts in the controlled host set are migrated to an isolated network environment for security testing and cleanup.

18. A device for identifying the master control terminal of a botnet, wherein, The device includes: The host determination module is used to determine the set of controlled hosts corresponding to the attack traffic in response to the detection of attack traffic targeting external communication addresses, and to extract the network flow data of each controlled host in the set of controlled hosts within a preset attack tracing time window. The host description module is used to perform botnet feature preprocessing on the network flow data for any one of the controlled hosts in the set of controlled hosts, and generate host behavior description information of the controlled host. The master control reasoning module is used to input the host behavior description information of each of the controlled hosts into a large language model in the security field, and guide it to perform semantic reasoning by using the prompt information configured for the feature analysis of botnets in the large language model, and output a list of suspected master control terminals corresponding to each controlled host. The master control determination module is used to aggregate the list of suspected master control terminals of each of the controlled hosts, perform attack correlation analysis, and determine the target botnet master control terminal based on the analysis results.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-17.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-17.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-17.