A network attack trace evidence method, system and storage medium

By acquiring network attack data and utilizing knowledge bases and large models to determine information collection strategies and tools, the accuracy and reliability issues of traditional tracing and evidence collection technologies in complex network attack scenarios have been resolved. This has enabled comprehensive collection and analysis of device information, thereby improving the accuracy and reliability of tracing and evidence collection.

CN121644248BActive Publication Date: 2026-04-17HANGZHOU DPTECH TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DPTECH TECH
Filing Date
2026-02-04
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional source tracing and evidence collection techniques are difficult to adapt to complex cyberattack scenarios, cannot accurately determine the actual extent of damage to attacked devices, and are prone to misjudgment of proactive probing behaviors, leading to interruptions in the evidence collection process. Furthermore, the lack of correlation analysis between scan results and contextual data such as network traffic makes it difficult to form a complete chain of evidence from isolated clues, resulting in insufficient accuracy and credibility of source tracing conclusions.

Method used

By acquiring network attack data, using a pre-set knowledge base and large model to determine the dimensions and tools for information collection, pushing the information to the attacked devices to collect device information, retrieving relevant knowledge for source tracing and evidence collection from the knowledge base, and combining the large model to output source tracing and evidence collection results, including attack verification, potential attack determination, and path tracing.

Benefits of technology

It enables precise collection and comprehensive analysis of attack evidence, improves the accuracy, comprehensiveness and reliability of source tracing and evidence collection, and ensures the credibility and integrity of the evidence collection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644248B_ABST
    Figure CN121644248B_ABST
Patent Text Reader

Abstract

The application provides a network attack traceability evidence method, system and storage medium, the method comprising: obtaining attack data with attack characteristics and attack attributes; searching the knowledge base for information collection knowledge matching the attack type of the data, including the information collection dimensions and tools of the attacked device in historical cases; determining the information collection dimensions and tools of the attacked device based on the attack data and the knowledge using a large model; pushing the information collection tools to the attacked device to collect device information according to the information collection dimensions; searching the knowledge base for traceability evidence knowledge matching the device information, including the device information, attack data and traceability evidence results in historical cases; outputting the traceability evidence results based on the device information, knowledge and attack data using a large model, including the results of verifying whether the attack characteristics exist, the results of determining whether the device is subject to potential attacks and the corresponding types, the traceability results of the attack path, and the evidence for each result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to a method, system and storage medium for tracing and obtaining evidence of network attacks. Background Technology

[0002] With the acceleration of digitalization, cyberspace has become a core carrier of socio-economic activities. Various cyberattack methods are becoming increasingly complex and diverse, posing serious challenges to personal privacy and business operations. Source tracing and evidence collection, as a crucial link in the cybersecurity defense system, is typically initiated after cyberattack detection equipment detects cyberattack data exhibiting cyberattack characteristics. It uses technical means to trace the source of the attack, reconstruct the attack path, and secure key evidence, providing a solid foundation for subsequent security protection optimization, emergency response, and legal accountability. Its accuracy and comprehensiveness directly affect the overall effectiveness of cybersecurity defense.

[0003] However, traditional attribution and forensic techniques are ill-suited to the complexities of current cyberattack scenarios and exhibit significant limitations. One common technique involves capturing and storing all network traffic packets, reconstructing the attack session based on source and destination addresses, port numbers, and transmitted content, and then extracting attack-related evidence for network attack analysis. However, this technique relies solely on passive analysis of network-level data, neglecting the collection of information about the attacked device. It can only reconstruct the attack process but cannot accurately determine the actual extent of damage to the attacked device or verify critical statuses such as whether device vulnerabilities have been patched. Another technique employs fixed scanning tools like vulnerability scanners and port scanners to actively probe the attacked device, detecting open ports, known vulnerabilities, and abnormal processes to obtain attack clues. However, such active probing is easily misjudged as malicious attacks by network security devices, interrupting the forensic process. Furthermore, fixed scanning tools cannot dynamically adjust the information collection dimensions based on the attack type, failing to cover comprehensive device information such as temporary files and operation logs. Moreover, the scan results lack correlation analysis with contextual data such as network traffic, making it difficult to form a complete chain of evidence from isolated clues, resulting in insufficient accuracy and credibility of attribution conclusions. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this application provides a method, system and storage medium for tracing and obtaining evidence of network attacks.

[0005] According to a first aspect of the embodiments of this application, a method for tracing and obtaining evidence of network attacks is provided, the method comprising:

[0006] Acquire network attack data, wherein the network attack data is data that has been pre-determined to have network attack characteristics and can characterize the attributes of the current network attack;

[0007] Retrieve information collection association knowledge that matches the attack type corresponding to the network attack data from a preset knowledge base. The information collection association knowledge includes: historical information collection dimensions and historical information collection tools for historically attacked devices in historical source tracing and forensics cases that match the attack type.

[0008] Using a pre-defined large model, based on the network attack data and the information collection knowledge, determine the information collection dimensions and information collection tools of the currently attacked device pointed to by the network attack data;

[0009] The information collection tool is pushed to the attacked device so that the attacked device runs the information collection tool according to the information collection dimension to collect the device information of the attacked device. The device information can characterize the operating status of the attacked device under the network attack.

[0010] Retrieve source tracing and forensics-related knowledge that matches the device information from the knowledge base. The source tracing and forensics-related knowledge includes: historical device information in historical source tracing and forensics cases that match the device information, as well as historical network attack data and historical source tracing and forensics results corresponding to the historical device information.

[0011] Using the large model, based on the device information, the source tracing and evidence collection association knowledge, and the network attack data, source tracing and evidence collection results are output. The source tracing and evidence collection results include: verification results to confirm whether the network attack features of the network attack data actually exist; determination results to determine whether there are potential network attack features on the attacked device and the attack type corresponding to the potential network attack features; source tracing results to trace the attack path of the network attack; and evidence for the verification results, the determination results, and the source tracing results, respectively.

[0012] According to a second aspect of the embodiments of this application, a network attack tracing and forensics system is provided, the system comprising:

[0013] Several network devices, wherein the network devices are devices in the network that may be subject to network attacks;

[0014] Several network attack detection devices are communicatively connected to the network device to detect whether the network device is under network attack, and when a network attack is detected, to identify and extract network attack data with network attack characteristics.

[0015] The source tracing and evidence collection device is communicatively connected to the network device and the network attack detection device, respectively, and is used to execute the source tracing and evidence collection method for network attacks described in the first aspect.

[0016] According to a third aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the method described in the first aspect.

[0017] The technical solutions provided in this application embodiment may include the following beneficial effects:

[0018] In this embodiment, network attack data representing network attack attributes is acquired, and then information collection-related knowledge matching the attack type corresponding to the data is retrieved from a preset knowledge base. A large model is used to combine the network attack data with the retrieved knowledge to determine the information collection dimensions and tools corresponding to the attacked device and push them to the attacked device, enabling it to collect device information that represents the device's operational status. Subsequently, source tracing and evidence collection-related knowledge matching the device information is retrieved from the knowledge base. The large model integrates the device information, source tracing and evidence collection-related knowledge, and network attack data to output source tracing and evidence collection results that include attack verification results, potential attack judgment results, attack path tracing results, and corresponding evidence. This effectively achieves accurate collection and comprehensive analysis of attack evidence, improving the accuracy, comprehensiveness, and reliability of source tracing and evidence collection.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this application, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] Figure 1 This is a schematic diagram of the structure of a traceability and evidence collection system according to an exemplary embodiment of this application.

[0022] Figure 2 This is a flowchart illustrating a source tracing and evidence collection method according to an exemplary embodiment of this application.

[0023] Figure 3 This is a flowchart illustrating another source tracing and evidence collection method according to an exemplary embodiment of this application.

[0024] Figure 4 This is a flowchart illustrating a network attack detection method according to an exemplary embodiment of this application.

[0025] Figure 5 This is a schematic diagram of the structure of a traceability and evidence collection device according to an exemplary embodiment of this application. Detailed Implementation

[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0027] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0029] With the acceleration of digitalization, cyberspace has become a core carrier of socio-economic activities. Various cyberattack methods are becoming increasingly complex and diverse, posing serious challenges to personal privacy and business operations. Source tracing and evidence collection, as a crucial link in the cybersecurity defense system, is typically initiated after cyberattack detection equipment detects cyberattack data exhibiting cyberattack characteristics. It uses technical means to trace the source of the attack, reconstruct the attack path, and secure key evidence, providing a solid foundation for subsequent security protection optimization, emergency response, and legal accountability. Its accuracy and comprehensiveness directly affect the overall effectiveness of cybersecurity defense.

[0030] However, traditional attribution and forensic techniques are ill-suited to the complexities of current cyberattack scenarios and exhibit significant limitations. One common technique involves capturing and storing all network traffic packets, reconstructing the attack session based on source and destination addresses, port numbers, and transmitted content, and then extracting attack-related evidence for network attack analysis. However, this technique relies solely on passive analysis of network-level data, neglecting the collection of information about the attacked device. It can only reconstruct the attack process but cannot accurately determine the actual extent of damage to the attacked device or verify critical statuses such as whether device vulnerabilities have been patched. Another technique employs fixed scanning tools like vulnerability scanners and port scanners to actively probe the attacked device, detecting open ports, known vulnerabilities, and abnormal processes to obtain attack clues. However, such active probing is easily misjudged as malicious attacks by network security devices, interrupting the forensic process. Furthermore, fixed scanning tools cannot dynamically adjust the information collection dimensions based on the attack type, failing to cover comprehensive device information such as temporary files and operation logs. Moreover, the scan results lack correlation analysis with contextual data such as network traffic, making it difficult to form a complete chain of evidence from isolated clues, resulting in insufficient accuracy and credibility of attribution conclusions.

[0031] Based on this, in order to solve the problems existing in related technologies, this application provides a method for tracing and obtaining evidence of network attacks. The method first obtains network attack data that characterizes the attributes of network attacks, then retrieves information collection-related knowledge that matches the attack type corresponding to the data from a preset knowledge base, and uses a large model to combine network attack data and retrieved knowledge to determine the information collection dimensions and information collection tools corresponding to the attacked device and push them to the attacked device so that it can collect device information that can characterize the device's operating status. Subsequently, it retrieves tracing and obtaining evidence-related knowledge that matches the device information from the knowledge base, and integrates device information, tracing and obtaining evidence-related knowledge, and network attack data through a large model to output tracing and obtaining evidence results that include attack verification results, potential attack judgment results, attack path tracing results, and corresponding evidence. This effectively realizes the accurate collection and comprehensive analysis of attack evidence and improves the accuracy, comprehensiveness, and reliability of tracing and obtaining evidence.

[0032] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0033] Before introducing the source tracing and evidence collection method provided in the embodiments of this application, the source tracing and evidence collection system applied by the method is first described in order to better understand the application scenario of the method.

[0034] Figure 1 This is a schematic diagram illustrating the structure of a network attack tracing and forensics system according to an exemplary embodiment of this application. Figure 1As shown, the source tracing and evidence collection system includes several network devices 101, several network attack detection devices 102, and a source tracing and evidence collection device 103.

[0035] Network devices 101 represent various devices within the network that may be vulnerable to cyberattacks. Their specific types can be flexibly configured according to actual application scenarios, including but not limited to servers, switches, routers, industrial control equipment, IoT terminals, personal computers, and mobile terminals. These network devices 101 generate a large amount of network data during daily operation, covering various information such as data transmission between devices, command interaction, and status feedback. This data not only comprehensively reflects the real-time operating status of the network but also serves as a potential carrier for cyberattacks, making it a key analytical object for subsequent network attack detection and tracing. Specifically, network data may include, but is not limited to, network interaction messages, device security logs, and network traffic statistics. Network interaction messages can record in detail the specific content, protocol type, and interaction sequence of data transmission between devices; security logs can completely retain device operating events, user access records, permission change operations, and abnormal behavior trigger information; and traffic statistics can intuitively reflect core network operating characteristics such as network bandwidth usage, data transmission rate, and number of connection sessions.

[0036] The network attack detection device 102 is communicatively connected to the network device 101 and can collect various types of network data generated by each network device 101 in real time. It can perform network attack detection on the collected network data based on preset attack detection methods, accurately determining whether each network device 101 has been subjected to a network attack. When a clear network attack is detected, the network attack detection device 102 will further identify and extract network attack data from the network data that exhibits network attack characteristics and can characterize the current network attack attributes, providing basic data support for subsequent tracing and evidence collection.

[0037] The source tracing and evidence collection device 103 is communicatively connected to the network device 101 and the network attack detection device 102, respectively. Its deployment method can be flexibly selected according to actual application needs. It can be directly integrated into the network attack detection device 102 to achieve integrated detection and evidence collection, or it can be independently deployed and integrated into other network devices such as preset servers, cloud computing devices, or dedicated security devices to meet the application needs of distributed network environments. This source tracing and evidence collection device 103 can acquire network attack data determined by the network attack detection device 102 in real time and perform source tracing and evidence collection based on the source tracing and evidence collection method provided in this application embodiment. The final output includes source tracing and evidence collection results, including verification of whether network attack characteristics truly exist, determination of whether the attacked device has potential network attack characteristics and corresponding attack types, tracing the attack path of the network attack, and all evidence supporting the above results. This provides a strong basis for network security emergency response and responsibility determination.

[0038] Figure 2 This is a flowchart illustrating a method for tracing and obtaining evidence of network attacks according to an exemplary embodiment of this application. Figure 2 As shown, the method includes steps S201 to S206.

[0039] Step S201: Obtain network attack data, which is data that has been pre-determined to have network attack characteristics and can characterize the attributes of the current network attack.

[0040] Network attack data is the fundamental input for attribution investigation and forensics. It comprehensively covers various attributes of current network attacks, ensuring sufficient data support for subsequent analysis. Specific types of data include, but are not limited to, the IP address of the attacked device, the name of the network attack, the type of the network attack, the corresponding payload information, and the network context messages at the time of the attack. Specifically, the IP address of the attacked device is used to accurately locate the target of the attack, providing a clear direction for subsequent tool deployment and information collection; the attack name and attack type can quickly identify the core characteristics of the attack (such as "SQL injection attack," "ransomware attack," etc.), providing a basis for matching historical forensic experience and determining information collection strategies; the payload information corresponding to the attack contains key content such as malicious code fragments and command data carried by the attack behavior, serving as important clues for identifying attack intent; and the network context messages record the network interaction environment at the time of the attack, covering information such as the communication sequence, protocol type, and data transmission content between the attacker and the attacked device, enabling the reconstruction of the network scenario at the time of the attack.

[0041] Specifically, the method of acquiring network attack data can be flexibly configured according to the actual application scenario. Either the network attack detection devices 102 can automatically transmit the extracted network attack data to the source tracing and evidence collection device 103 through a preset communication interface after detecting a network attack, or in special scenarios (such as when the network attack detection devices 102 are not fully adapted to the automatic transmission function), staff can enter the confirmed network attack data through the manual input interface of the source tracing and evidence collection device 103 to ensure that the source tracing and evidence collection device 103 can obtain valid data in a timely manner and start the subsequent source tracing and evidence collection process.

[0042] Step S202: Retrieve information collection association knowledge that matches the attack type corresponding to the network attack data from the preset knowledge base. The information collection association knowledge includes: historical information collection dimensions and historical information collection tools for historically attacked devices in historical tracing and forensic cases that match the attack type.

[0043] Step S203: Using a pre-defined large model, based on network attack data and information gathering knowledge, determine the information gathering dimensions and tools of the currently attacked device pointed to by the network attack data.

[0044] Step S204: Push the information collection tool to the attacked device so that the attacked device runs the information collection tool according to the information collection dimension to collect the device information of the attacked device. The device information can characterize the operating status of the attacked device under network attack.

[0045] After obtaining network attack data, directly conducting subsequent source tracing and forensics based on this data can only reconstruct the transmission path and interaction process of the network attack at the network level. It cannot delve into the internal damage state of the attacked device, accurately determine the actual extent of damage caused, or verify whether the exploited vulnerabilities have been patched, ultimately leading to significant limitations in the source tracing and forensics results. Therefore, targeted device-side information collection is necessary to supplement the operational data and status information of the attacked device to achieve comprehensive and in-depth source tracing and forensics. To efficiently collect device-side information, it is first necessary to determine the information collection dimensions and tools suitable for the current attack type. Ignoring the differences in attack types and applying a fixed information collection strategy to all attack scenarios will not only result in incomplete collection of key information due to insufficient strategy targeting but may also lead to the acquisition of a large amount of redundant data unrelated to source tracing and forensics, wasting resources and reducing efficiency. Therefore, this embodiment proposes a scheme to formulate differentiated information collection strategies for attacked devices based on different attack types.

[0046] To quickly determine an accurate information collection strategy, the most direct and effective approach is to draw on successful experiences in historical source tracing and forensics. Therefore, after acquiring network attack data, the source tracing and forensics device 103 can first search a pre-built knowledge base for information collection related knowledge that matches the current attack type. This knowledge base is a pre-built knowledge reserve for network attack source tracing and forensics, containing a massive amount of categorized and organized historical source tracing and forensics case data. It covers various common attack types such as SQL injection attacks, ransomware attacks, DDoS attacks, and cross-site scripting attacks, as well as handling cases for some new and rare attacks. Each case is associated with a complete handling plan for the corresponding attack type, including historical information collection dimensions and tools for different historically attacked devices. At the same time, this knowledge base supports dynamic updates based on new source tracing and forensics cases and technical experience, ensuring the timeliness and comprehensiveness of its knowledge reserves. In a specific search, the source tracing and evidence collection device 103 can first extract key feature fields from the acquired network attack data, namely the attack type of the current network attack, and then use the attack type as the search keyword to perform a precise matching search in the knowledge base, filter out historical source tracing and evidence collection cases that are consistent with or highly similar to the current attack type, and then extract the corresponding information collection association knowledge in these cases to provide experience support for determining the information collection strategy of the current attacked device.

[0047] Based on the information collected and related knowledge obtained through retrieval, the source tracing and evidence collection device 103 can determine the information collection dimensions and tools suitable for the currently attacked device through a pre-set large model. This large model is an intelligent decision-making model pre-trained based on deep learning algorithms, which has the ability to perform attack type matching analysis, historical experience transfer, and tool dimension adaptation. Specifically, if the knowledge base contains information collection association knowledge that precisely matches the attack type of the current network attack data, the big model directly transfers and determines the historical information collection dimensions and tools for historically attacked devices from this association knowledge as the information collection dimensions and tools corresponding to the current attacked device. This enables the reuse of past successful experiences and improves the efficiency and accuracy of information collection strategy formulation. If the knowledge base does not contain information collection association knowledge that matches the current attack type, i.e., the current attack type is a new or rare attack, the big model automatically calls the preset tool library and determines all information collection tools stored in the tool library as the information collection tools required for this forensics. At the same time, it determines all information collection dimensions that these tools can cover as the information collection dimensions for this forensics, thereby ensuring that there are no blind spots in the information collection of new attacks and avoiding the omission of key information due to lack of experience. The information collection tools include mainstream information collection tools such as osquery, Fenrir, and LinPEAS, which can respectively realize information collection functions of different dimensions such as system configuration query, malicious code scanning, and privilege escalation vulnerability detection.

[0048] After determining the information collection dimensions and tools, the forensic tracing device 103, based on the determined information collection tools, can push the corresponding toolkit to the currently attacked device through a preset secure communication channel, and simultaneously issue information collection dimension instructions matching the tools. Compared to the active detection method in related technologies, this embodiment allows the attacked device to autonomously run and collect data after receiving the toolkit, effectively avoiding the problem of evidence collection interruption caused by external active scanning being mistakenly judged as malicious activity by network security devices. After receiving the toolkit and instructions, the attacked device automatically runs the information collection tools according to the preset information collection dimensions to carry out comprehensive device information collection. The collected device information can comprehensively characterize the real-time operating status of the attacked device during the network attack, specifically including but not limited to the attacked device's device configuration information, process running information, operation logs, and temporary files generated during the network attack. Among them, device configuration information can reflect the device's system version, software and hardware environment, and security policy configuration, providing a basis for analyzing attack vulnerabilities; process running information can present the device's process startup and running status when the attack occurs, making it easier to identify malicious processes; operation logs can completely record user operation behaviors and system events before and after the attack; and temporary files generated during the attack may contain core forensic clues such as malicious code residues and attack command fragments, providing key support for subsequent source tracing analysis.

[0049] Step S205: Retrieve source tracing and forensic knowledge that matches the device information from the knowledge base. Source tracing and forensic knowledge includes: historical device information in historical source tracing and forensic cases that match the device information, as well as historical network attack data and historical source tracing and forensic results corresponding to the historical device information.

[0050] Step S206: Using the large model, based on device information, source tracing and evidence collection related knowledge, and network attack data, output source tracing and evidence collection results. The source tracing and evidence collection results include: verification results to confirm whether the network attack characteristics of the network attack data actually exist; judgment results to determine whether there are potential network attack characteristics on the attacked device and the attack type corresponding to the potential network attack characteristics; source tracing results to trace the attack path of the network attack; and evidence for the verification results, judgment results, and source tracing results respectively.

[0051] After collecting device information of the attacked device, directly relying on a large model to analyze scattered network attack data and device information is prone to analytical bias due to a lack of historical experience for reference, and may even lead to a "large model illusion," resulting in a lack of credibility and rigor in the output forensic results. Therefore, similar to the information collection strategy determined earlier, this embodiment can still choose to reuse historical source tracing and forensic experience to provide a reliable reference for in-depth analysis of the large model, thereby improving the accuracy and effectiveness of the forensic results. That is, the source tracing and forensic device 103 searches the preset knowledge base for source tracing and forensic related knowledge that matches the currently collected device information. In the specific search process, the source tracing and forensic device 103 can first extract core features from the collected device information, such as device configuration features, process abnormal features, log key event features, and malicious code features in temporary files, and then use these features as search conditions to perform multi-dimensional matching searches in the knowledge base, and filter out historical source tracing and forensic cases that highly match the features of the current device information. The extracted knowledge related to source tracing and evidence collection not only includes historical device information that matches the current device information in the case, but also covers historical network attack data corresponding to the historical device information, as well as complete historical source tracing and evidence collection results derived from the analysis of both, forming a complete experience chain from data input to result output, providing comprehensive reference support for subsequent large-scale model analysis.

[0052] Based on the acquired current device information, network attack data, and retrieved source tracing and forensic knowledge, the source tracing and forensic device 103 can utilize a pre-set large model to comprehensively analyze multi-dimensional data and ultimately output comprehensive and accurate source tracing and forensic results. During the analysis, the large model first compares the current network attack data with historical network attack data, combining this with the device operating status recorded in the current device information during the attack period to verify whether the network attack features in the network attack data truly exist and act on the attacked device, forming corresponding verification results. Secondly, relying on the experience in identifying potential attack features summarized from historical cases, the large model deeply mines the current device information to determine whether there are potential network attack features in the attacked device that were not captured by the initial network attack data, and further determines the attack type corresponding to these potential features, forming a judgment result. Simultaneously, the large model combines the attack transmission path information in the network attack data, the process and log records in the device information, and the attack path tracing experience from historical cases to reconstruct the complete attack path from the initiator to the attacked device, forming a source tracing result. To ensure the traceability and legal validity of the evidence, the large model generates corresponding evidence chains simultaneously while outputting the above three types of results. Evidence for verification results can include feature comparison reports of network attack data and device information; evidence for judgment results can include records of extraction and analysis of potential attack features; and evidence for source tracing results can include interactive data and log evidence from each node in the attack path, thus achieving closed-loop management of source tracing and evidence collection.

[0053] However, in actual forensic investigations, the device information collected from attacked devices often contains a large amount of scattered and independent raw data. Directly inputting this data into a large model for analysis not only increases the computational burden on the model and reduces analysis efficiency, but also may prevent the model from accurately capturing key features due to a lack of clear logical connections between the data, thus affecting the accuracy of the forensic investigation results. Therefore, in some embodiments, before using a large model for forensic investigation analysis, the collected device information can be structured to build logical connections between the data, improving the efficiency and accuracy of subsequent analysis.

[0054] Specifically, the source tracing and evidence collection device 103 can perform structured processing of device information through the relationships between various historical device information pre-stored in the knowledge base. These relationships are the inherent logic of device information extracted and summarized from massive historical source tracing and evidence collection cases, covering correspondences of multiple data dimensions, including but not limited to the relationship between process running information and operation logs, and the relationship between file information and network transmission traffic. Among them, the relationship between process running information and operation logs can reflect the matching of the time nodes of a specific process's start, run, and termination with the corresponding event records in the system operation log. For example, the start of a malicious process is often accompanied by the generation of operation logs such as system permission changes and file read / write operations. The relationship between file information and network transmission traffic can establish a corresponding link between newly added, modified, and deleted file data within the device and the transmission traffic of a specific IP and port during a network attack. For example, malicious temporary files generated during an attack usually have temporal and content correlations with the file transmission traffic of external attack sources.

[0055] During structured processing, the source tracing and evidence collection device 103 can first retrieve the aforementioned preset association rules from the knowledge base, and then use these rules as a basis to classify, match, and integrate the collected raw device information. For example, the process ID, startup time, and running path in the process running information are associated and bound with the permission change records and file access records within the same time interval in the operation log; the generation time and file hash value of temporary files are matched and associated with the transmission traffic characteristics of the corresponding time period in the network attack data. Finally, the originally scattered device configuration information, process running information, operation logs, temporary files, and other data are transformed into structured data with clear logical relationships. The device information after structured processing can clearly present the inherent connections between data, greatly reducing the analysis difficulty of large models, and at the same time providing a more solid data foundation for subsequent in-depth analysis based on device information, source tracing and evidence collection association knowledge, and network attack data.

[0056] Furthermore, to ensure the reliability and practicality of the source tracing and evidence collection results, avoid invalid evidence conclusions due to analytical bias, and improve the closed-loop management level of source tracing and evidence collection work, some embodiments, such as Figure 3As shown, after outputting the source tracing and evidence collection results through the large model, the source tracing and evidence collection device 103 can also execute step S207 to further assess the security of the source tracing and evidence collection results using a preset security assessment model. Specifically, the security assessment model can be trained based on a massive number of historical valid evidence collection cases and has built-in multi-dimensional preset security assessment standards, such as: the completeness of the evidence chain of the verification results, the accuracy of identifying potential attack features of the judgment results, the logical coherence of the attack path of the source tracing results, and the correlation between the evidence and data corresponding to various results. During the assessment process, if the source tracing and evidence collection results meet the preset security assessment standards, that is, all assessment dimensions meet the standards, it indicates that the results have sufficient rigor and credibility. At this time, the source tracing and evidence collection device 103 executes step S208 to generate a formal source tracing and evidence collection report based on the source tracing and evidence collection results. The report not only comprehensively covers the verification results, judgment results, source tracing results, and corresponding complete evidence, but also deeply analyzes the severity of the cyberattack based on the source tracing results. For example, it classifies the severity level according to the importance of the attacked device, the scope of system function impact caused by the attack, and the risk of data leakage. It also provides targeted security protection recommendations, such as patch solutions for the vulnerabilities exploited by the attack, optimization of network access control policies for the attack path, and strengthening of real-time device monitoring rules for potential attack characteristics. If the source tracing results do not meet the preset security assessment standards, such as broken evidence chains, logical contradictions in the attack path, or doubts about the identification of potential attack characteristics, it indicates that the current source tracing has deficiencies. The entire source tracing method described above needs to be re-executed, from acquiring cyberattack data and collecting device information to retrieving source-related knowledge and conducting large-scale model analysis, to correct any deviations. The re-execution process will continue until the output traceability and evidence collection results meet the preset security assessment standards. If the number of re-executions reaches the preset maximum number of executions and no results that meet the standards are obtained, an alarm mechanism will be triggered to prompt staff to intervene for manual verification and handling, so as to ensure that the traceability and evidence collection work is not interrupted and no key risks are overlooked.

[0057] The above content fully presents the complete source tracing and evidence collection method of the source tracing and evidence collection device 103. However, the conduct of source tracing and evidence collection work requires accurate network attack detection as a prerequisite. Only by accurately identifying network attacks and locking relevant attack data can the source tracing and evidence collection device 103 conduct targeted subsequent source tracing and evidence collection analysis. Therefore, the following section will also detail the attack detection process of the network attack detection device 102 in the pre-source tracing and evidence collection stage, based on which accurate detection of various network attacks can be achieved, providing a reliable data foundation for subsequent source tracing and evidence collection.

[0058] Figure 4 This is a flowchart illustrating a network attack detection method according to an exemplary embodiment of this application. Figure 4As shown, this method can be applied to the network attack detection device 102 mentioned above to achieve accurate detection of various network attacks, specifically including the following steps S401 to S407.

[0059] Step S401: Obtain network data, which is used to characterize the real-time network operating status.

[0060] During network operation, the network attack detection device 102 can collect various network data generated by each network device 101 in real time through communication connections. This network data can effectively reflect the network operating status. Specifically, the network data may include, but is not limited to, network interaction messages, device security logs, and network traffic statistics. Among them, network interaction messages can record in detail the specific content, protocol type, and interaction sequence of data transmission between devices; security logs can completely retain device operation events, user access records, permission change operations, and abnormal behavior trigger information; and traffic statistics can intuitively reflect core network operating characteristics such as network bandwidth usage, data transmission rate, and number of connection sessions.

[0061] To ensure the efficiency and accuracy of subsequent detection processes, in some embodiments, the network attack detection device 102 can perform standardized preprocessing on the network data after acquiring it. For example, it can first parse the network data to extract core metadata such as source IP address, target IP address, communication port number, and protocol type, while extracting key segments from the packet payload and mining temporal features such as connection duration, packet size distribution, and request frequency. Subsequently, the parsed information is structured according to the key-value pair format to remove duplicate fields, invalid and redundant information, and data with incorrect format, unify the data encoding standard, and form standardized network data to provide high-quality data support for the input of subsequent detection models.

[0062] Step S402: Based on network data, construct the first prompt word of the first detection model. The first prompt word includes the role definition of the first detection model, network data, recognition task instructions of the first detection model, and the result output format of the first detection model.

[0063] Step S403: Input the first prompt word into the first detection model to perform preliminary network attack detection on the network data using the first detection model, and output the preliminary detection results; the preliminary detection results are used to preliminarily determine whether the network data has network attack characteristics and the corresponding attack type.

[0064] Step S404: Mark network data that is determined to have network attack characteristics by the preliminary detection results as target network data.

[0065] In network attack detection scenarios, network data volumes are typically large and contain a significant amount of normal data without network attacks. Directly using a single detection model to perform network attack detection on the entire dataset would generate extremely high computational pressure, leading to low detection efficiency and accuracy, making it difficult to meet real-time detection requirements. Therefore, this embodiment adopts a two-tier detection architecture. Two interconnected detection models process network data step-by-step. The first detection model performs preliminary attack detection on the entire dataset, quickly filtering out target network data that may exhibit network attack characteristics. Then, the second detection model performs deep attack detection on the filtered target network data. This achieves a rapid screening and precise verification detection effect, ensuring overall detection efficiency while effectively improving attack detection accuracy.

[0066] To conduct initial attack detection using the first detection model, the primary task is to construct a suitable first prompt word to ensure that the first detection model can clearly identify the detection target, accurately acquire data information, and complete the detection task as required. Specifically, the network attack detection device 102 can construct the first prompt word corresponding to the first detection model based on network data. The first prompt word can include four key components: First, the role definition of the first detection model, which clarifies the model's identity and positioning, such as "You are a professional preliminary network attack detection analyst, responsible for conducting preliminary attack detection on network data and outputting preliminary detection results," enabling the model to conduct detection work from a corresponding professional perspective; Second, the network data, i.e., the network data obtained in step S402, which serves as the core data basis for model detection and must be completely and standardizedly embedded in the prompt word to ensure that the model can fully obtain the basic information required for detection; Third, the detection task instruction of the first detection model, which clearly informs the model of the specific detection task, such as "Please conduct preliminary attack detection on the provided network data, determine whether there are network attack characteristics, and if so, identify the corresponding attack type (such as DDoS attack, SQL injection attack, port scanning attack, etc.)"; Fourth, the result output format of the first detection model, which constrains the model's output content through a standardized format, such as "Output format requirements: 1. Whether network attack characteristics exist: Yes / No; 2. Attack type (if present): XXX," so as to facilitate rapid analysis and processing of the detection results later.

[0067] After constructing the first prompt word, the network attack detection device 102 can input the constructed first prompt word into the first detection model. The first detection model then performs preliminary attack detection on the input network data based on the role positioning and task instructions in the prompt word. Specifically, in this embodiment, the first detection model can be a lightweight model adapted to the preliminary detection scenario. A lightweight model is a model with a relatively small number of parameters (e.g., in the range of millions to tens of millions of parameters). Compared to large models with hundreds of millions or more parameters, it does not require excessive computing and storage resources, and has the advantages of low deployment cost and fast response speed, which can meet the needs of rapid screening of massive network data in the preliminary detection stage. At the same time, through model transfer learning technology, the rich network attack detection knowledge learned by the large model can be transferred to the lightweight model, which can effectively make up for the limited learning ability of the lightweight model itself, significantly improve the accuracy of preliminary detection, and achieve the dual goals of efficient screening and accurate preliminary judgment. During the detection process, the first detection model will combine its own understanding of the characteristics of various network attacks to analyze key information such as metadata and temporal characteristics in the network data, determine whether there are characteristic clues of attack behavior, preliminarily identify possible attack types, and finally output preliminary detection results to determine whether the network data has network attack characteristics and the corresponding attack types.

[0068] Subsequently, the network attack detection device 102 analyzes the preliminary detection results output by the first detection model, marking network data that is determined to "exhibit network attack characteristics" as target network data. This type of target network data is suspected attack data and will enter the subsequent deep detection stage for further verification to ensure the accuracy of attack detection; while network data that is initially determined to "not exhibit network attack characteristics" will not be included in the subsequent deep detection scope, thereby achieving the purpose of rapid screening and focusing on key detection targets.

[0069] However, in real-world detection scenarios, the complexity of network data and the stealth of attack behaviors can lead to misjudgments by the first detection model. Therefore, to improve the reliability of preliminary detection results, in some embodiments, the first detection model outputs a preliminary confidence level along with the preliminary detection result. This confidence level characterizes the model's trust in the preliminary detection result; a higher confidence level indicates stronger reliability. Based on this confidence level, the preliminary detection process can be further optimized: when the preliminary detection result determines that the network data does not exhibit network attack characteristics, and the corresponding preliminary confidence level is less than a first preset threshold (e.g., set to 80%, with the specific value adjustable according to actual detection needs), it indicates that the first detection model has low trust in the detection result and there is a risk of misjudgment. At this point, the network attack detection device 102 can supplement the acquisition with historical network data within the corresponding historical period, such as the same IP interaction records related to the network data within the past 5 minutes, and traffic statistics characteristics within the historical period (such as packet size variance, data request frequency, etc.). Subsequently, the acquired historical network data is fused with the current network data, and the first prompt word is reconstructed based on the fused complete data to ensure that the prompt word contains more comprehensive detection basis. Finally, the reconstructed first prompt word is input again into the first detection model, and the first detection model re-performs preliminary network attack detection on the network data, thereby reducing the probability of false positives and improving the reliability of the preliminary detection results.

[0070] Step S405: Retrieve relevant knowledge matching the attack type corresponding to the target network data from the preset knowledge base.

[0071] After labeling the target network data, this embodiment does not directly perform deep detection on the target network data. Instead, it first retrieves relevant knowledge matching the attack type corresponding to the target network data from a pre-set knowledge base, and then uses this as a basis to advance the subsequent deep detection process. This is done to compensate for the inherent shortcomings of deep detection in related technologies, and to solve the problem of detection limitations caused by relying solely on sample features learned by the model itself for judgment, lacking professional knowledge support. In related technologies, when performing deep detection on suspected attack data, the data is only input into the detection model to complete feature comparison and judgment. The model's analysis logic is completely limited to the sample features learned during the training phase. This not only makes it difficult to accurately identify new and mutated attacks with insufficient sample coverage, but also fails to combine the attack's tactical logic, technical principles, historical cases, and other in-depth information for analysis. It is very easy to make misjudgments due to feature similarity, and at the same time, the judgment of attack behavior lacks professional basis support, which greatly reduces the accuracy and generalization ability of deep detection. This embodiment retrieves and matches related knowledge in advance, which can supplement the subsequent deep detection stage with professional knowledge such as the technical details, judgment rules, and historical cases corresponding to the attack. This provides a more comprehensive and professional basis for judgment in deep detection, which can effectively solve the problem of insufficient detection generalization ability in related technologies. It can also make the judgment logic of deep detection more in line with the professional understanding in the field of network security, further improve the accuracy of identifying complex and new attacks, and reduce the detection bias caused by feature mismatch.

[0072] Specifically, in some embodiments, the pre-defined knowledge base can adopt a hybrid storage mode of "structured data + unstructured data" to comprehensively cover various types of knowledge information related to network attacks, ensuring that the retrieved related knowledge is complete and practical. The structured data mainly covers the knowledge content corresponding to three core databases: the CVE vulnerability database (containing standardized information such as the number, description, scope of impact, and exploitation conditions of various known vulnerabilities), the ATT&CK tactical technology database (systematically summarizing the tactical targets commonly used by attackers, corresponding technical means, and implementation steps), and the attack signature database (storing typical characteristics of various attack behaviors, such as signature codes, abnormal traffic patterns, and specific interactive commands). The unstructured data includes APT attack reports released by security vendors (detailed records of the background, attack path, technical details, and source tracing results of APT attack events), vulnerability exploitation technology white papers (in-depth analysis of the exploitation principles, technical implementation methods, and defense points of various vulnerabilities), and attack and defense exercise case analyses (summarizing the manifestations of attack behaviors, response strategies, and lessons learned in actual attack and defense scenarios).

[0073] Furthermore, to improve the efficiency and accuracy of knowledge retrieval, a knowledge classification index can be established during the construction of the knowledge base. All structured and unstructured data are classified and stored according to a three-level classification dimension: "attack type - attack tactic - attack technique". For example, attack type (such as DDoS attack, SQL injection attack, ransomware attack, etc.) can be used as the primary classification standard to initially categorize various types of knowledge; then, under the same attack type, secondary subdivision can be carried out according to attack tactics (such as initial access, persistence, privilege escalation, data leakage, etc.); finally, under the same tactic dimension, it can be further refined to specific attack techniques (such as remote code execution by exploiting vulnerabilities, brute-force cracking of account passwords, obtaining initial access rights through phishing emails, etc.), forming a hierarchical knowledge classification system.

[0074] Based on the above knowledge base architecture, the specific implementation process of step S205 can be as follows: The network attack detection device 102 first extracts the preliminary attack type corresponding to the target network data (i.e., the attack type output by the first detection model), uses the attack type as the core retrieval condition, and combines it with the three-level classification index to conduct a precise retrieval in the preset knowledge base; during the retrieval process, it matches all structured and unstructured data related to the attack tactics and attack techniques corresponding to the attack type, and finally filters out the related knowledge that can provide support for subsequent in-depth detection, such as the typical technical characteristics, common implementation paths, known vulnerability exploitation conditions, and behavioral performance in historical attack cases corresponding to the attack type.

[0075] Step S406: Based on the associated knowledge, construct the second prompt word of the second detection model. The second prompt word includes the set of historical attack cases in the associated knowledge, the location information of historical network attack features in the associated knowledge, the historical network attack detection rules in the associated knowledge, and the result output format of the second detection model.

[0076] After retrieving relevant knowledge matching the attack type corresponding to the target network data from the knowledge base, a second prompt word adapted to the second detection model can be constructed based on this relevant knowledge, providing targeted input guidance for the subsequent deep detection of the second detection model. It is worth noting that the second prompt word constructed in this embodiment is not a uniform, fixed template style, but is dynamically generated based on the attack type of the target network data obtained from the initial attack detection. That is, different attack types of target network data will generate corresponding exclusive second prompt words adapted to their characteristics. This dynamic generation of prompt words effectively avoids the drawbacks of using fixed prompt word templates to drive model detection in related technologies. In related technologies, fixed prompt word templates cannot adjust the model detection guidance logic according to the differences in attack types. Faced with differences in the core features and judgment points of different attack types, using a uniform expression to guide model detection can easily lead to the model failing to focus on key detection dimensions, thereby exacerbating detection bias and reducing detection accuracy. This embodiment dynamically generates second prompt words adapted to specific attack types, allowing the prompt word content to closely match the characteristics of the target network data to be detected. This accurately guides the second detection model to focus on the core features and key judgment rules of the attack type for in-depth analysis, which not only improves detection efficiency but also significantly enhances the accuracy of in-depth detection, ensuring that the verification results of suspected attack data are more reliable.

[0077] Specifically, the second prompt word constructed based on association knowledge in this embodiment mainly consists of four core elements:

[0078] The first is a collection of historical attack cases within the associated knowledge. This collection consists of typical cases selected from the search results that perfectly match the current attack type. It covers core information such as the attack scenario, implementation steps, key behavioral characteristics, attack chain, and ultimate impact of each case. Incorporating these cases into the second set of prompts helps the second detection model quickly establish an understanding of the current attack type. By comparing the target network data with the features in historical cases, it can more accurately capture attack clues and provide a reference for in-depth judgment.

[0079] Secondly, there is the location information of historical network attack features within the associated knowledge. This information is extracted from the associated knowledge and shows the specific distribution of historical attack features corresponding to the current attack type within the network data. For example, features of a certain type of attack often appear in the payload field of network packets, the "operation type" field of security logs, or the time series of "packet size distribution" in traffic statistics. Clearly defining this location information directly guides the second detection model to focus on core data areas for in-depth analysis, avoiding the inefficient detection caused by indiscriminately traversing the entire network data, while significantly improving the accuracy of feature recognition.

[0080] Thirdly, there are historical network attack detection rules within the associated knowledge base. These rules can originate from standardized detection logic corresponding to the current attack type in structured knowledge bases (such as the ATT&CK tactical and technical database and attack signature database), or they can include empirical detection criteria summarized in unstructured knowledge bases (such as attack and defense exercise case analyses and vulnerability exploitation technology white papers). For example, "When a specific vulnerability exploitation code fragment appears in network data, and the data request frequency exceeds a preset threshold, it can be determined that this type of attack exists." Integrating these rules into the second prompt word can provide clear judgment criteria for the second detection model, standardize its detection logic, and reduce bias caused by subjective judgment of the model.

[0081] Fourthly, the output format of the second detection model. To facilitate the rapid analysis, storage, and application of the deep detection results, a unified output format can be clearly specified in the second prompt. For example, the model may be required to output "Confirm whether network attack characteristics exist: Yes / No", "Attack type (if present): XXX", "Core attack characteristics and corresponding locations: XXX", "Detection basis (related cases / rules): XXX", etc., to ensure the standardization, completeness, and readability of the output results.

[0082] The network attack detection device 102 integrates and sorts out the above four core elements to form a second prompt word that is precisely matched with the current target network data attack type, laying a solid foundation for the second detection model to carry out efficient and accurate in-depth detection.

[0083] Step S407: Input the second prompt word and the target network data into the second detection model to perform deep network attack detection on the target network data using the second detection model, and output the deep detection result; the deep detection result is used to determine whether the target network data has network attack characteristics and the corresponding attack type.

[0084] After constructing the second prompt word, the network attack detection device 102 can input the constructed second prompt word and the marked target network data into the second detection model, which will then perform targeted deep network attack detection. Specifically, in this embodiment, the second detection model can be a medium-sized model. A medium-sized model is one with a parameter count between a lightweight model and a large model (e.g., a parameter count in the tens of millions to hundreds of millions). Compared to a lightweight model, it has stronger feature learning and deep analysis capabilities, enabling it to uncover hidden deep attack features in the target network data and analyze the logical links of attack behaviors. Compared to a large model, it does not require excessive computing and storage resources, making deployment costs more controllable. It also has a faster response speed, ensuring both deep detection accuracy and detection efficiency, thus meeting the core requirement of accurate verification in a two-level detection architecture.

[0085] During the detection process, the second detection model first establishes a deep cognitive framework for the attack type corresponding to the current target network data based on the historical attack case set, attack feature location information, and attack detection rules in the second prompt. Then, it focuses on the core areas of the target network data, performing a dimension-by-dimensional comparative analysis of the target network data's features and the associated knowledge in the prompt. For example, based on the explicitly stated attack feature location information in the prompt, the model accurately locates the corresponding fields in the target network data and checks for content matching historical attack features. Simultaneously, it combines the detection rules in the prompt to determine whether the characteristics of the target network data meet the attack judgment conditions and, referring to the behavioral patterns of historical attack cases, verifies whether the behavioral logic in the target network data matches the attack behavior. Finally, the second detection model outputs a deep detection result, which accurately determines whether the target network data indeed contains network attack features and the corresponding precise attack type, providing a core basis for the final handling of network attacks.

[0086] Similarly, to further ensure the reliability of depth detection results, in some embodiments, the second detection model outputs the depth confidence level corresponding to the depth detection result at the same time as outputting the depth detection result. The confidence level is used to characterize the degree of trust that the model has in the preliminary detection result. The higher the confidence level, the stronger the reliability of the detection result.

[0087] Based on this depth confidence level, this embodiment also sets up a complete supplementary detection and result confirmation process: when the depth confidence level output by the second detection model is greater than or equal to the second preset threshold, it indicates that the model has a high enough level of confidence in the current depth detection result, and the reliability of the detection result can be guaranteed. At this time, the depth detection result is directly confirmed as the final detection result, and there is no need to start the subsequent supplementary detection process. The network attack detection device 102 can synchronize the final detection result (including whether there is a network attack, the corresponding attack type, etc.) to the network security management platform, and trigger the corresponding handling response mechanism. For example, if it is determined that there is a network attack, it will automatically match the defense strategy (such as blocking attack traffic, blocking attack IPs, isolating affected devices, etc.) according to the attack type to achieve rapid response and handling of the attack; if it is determined that there is no network attack, the "suspected attack" mark of the target network data will be removed, and it will be included in the normal network data management scope. When the depth confidence level output by the second detection model is less than the second preset threshold (for example, it can be set to 90%, and the specific value can be adjusted according to the actual detection needs), it indicates that the model has a low level of confidence in the current depth detection result, and there is a risk of judgment deviation. At this point, the network attack detection device 102 can initiate a multi-model collaborative verification process, calling at least two different recognition models of the same scale as the second detection model. These different recognition models can have different model structures and feature learning focuses (for example, some models focus on temporal feature analysis, while others focus on text feature extraction), which can perform a deeper detection of the target network data from multiple perspectives, avoiding misjudgments caused by the cognitive limitations of a single model. After multi-model detection is completed, the network attack detection device 102 will perform a consistency comparison of the deep detection results output by each model: if the deep detection results output by all different identification models are consistent with the deep detection results output by the second detection model (i.e., the judgment conclusion and attack type are exactly the same), it means that even if the confidence of a single model is low, the multi-model collaborative verification supports the result, and the deep detection result can be confirmed as the final detection result. The corresponding result synchronization and processing procedures mentioned above can then be executed. If the deep detection result output by any different identification model differs from the deep detection result output by the second detection model (for example, the second detection model determines that "a certain type of attack exists," while a different identification model determines that "no attack exists," or the attack types determined by the two are different), it means that the current detection result is controversial and needs to be further verified. At this time, the network attack detection device 102 can automatically trigger the manual review process, and fully synchronize the target network data, the detection results of the second detection model, the detection results of different identification models, the corresponding confidence levels, the related knowledge retrieval results and other relevant information to the operation terminal of the security operation and maintenance personnel. The professional operation and maintenance personnel will then conduct manual analysis based on their own experience and professional knowledge to finally determine whether the target network data is subject to a network attack and the corresponding attack type.The conclusions drawn from manual review will serve as the final test results, which will then be synchronized to the management platform and appropriate actions will be taken to ensure the accuracy and authority of the test results and minimize the risk of misjudgment or omission.

[0088] Building upon the above embodiments, and considering the dynamic evolution of network attack methods and the continuous emergence of new attack characteristics, a fixed static knowledge base cannot adapt to these changes in real time, potentially leading to a lag in the identification of new attacks. Furthermore, if errors are found in the deep detection results during manual review, the knowledge base needs to be corrected promptly to avoid subsequent misjudgments. Therefore, in some embodiments, a closed-loop knowledge update mechanism can be constructed to achieve dynamic iterative optimization of the knowledge base.

[0089] Specifically, when the deep detection result determines that the target network data contains network attack characteristics, but the matching degree between the network attack characteristics of the target network data and the historical network attack characteristics stored in the knowledge base is less than a preset matching threshold (e.g., 60%, which is considered a discovery of a new attack characteristic), or when the network attack detection device 102 receives feedback from human intervention that the deep detection result is incorrect, the knowledge update process will be automatically triggered. First, the network attack detection device 102 extracts the core network attack characteristics from the target network data. Then, it standardizes the extracted attack characteristics to generate standardized network attack characteristics that include attack type, attack characteristic description, attack detection logic, and attack detection basis, and assigns a unique ID to each to ensure the uniqueness of the identifier. After standardization, the network attack detection device 102 can push the standardized network attack characteristics to the knowledge management platform for manual review. After the review is approved, it is synchronously updated to the structured feature library and vector database of the knowledge base, and a corresponding update log is generated, recording information such as the feature entry time, reviewer, and related cases. Meanwhile, in order to allow the detection model to adapt to the new knowledge in a synchronous manner, the first and second detection models can be incrementally learned and fine-tuned using newly added feature samples within a certain period (such as monthly) to improve the model's ability to identify new attack features.

[0090] At the prompt word construction level, to avoid excessive imperative descriptions interfering with the model's professional judgment and to improve the professionalism and reliability of the model's reasoning, some embodiments may employ a minimum instruction strategy to optimize the prompt word structure. Specifically, when constructing the second prompt word, the proportion of imperative descriptions can be strictly controlled (e.g., its proportion ≤30%), while the proportion of knowledge base references can be significantly increased. This allows the second detection model to rely more on retrieved professional knowledge rather than general instructions during the reasoning process, thereby weakening the interference of subjective imperative descriptions on the model's judgment and making the deep detection results output by the model more professional and accurate.

[0091] Furthermore, in actual network attack detection processes, when retrieving related knowledge for target network data, knowledge corresponding to certain attack types is repeatedly retrieved and used. Multiple retrievals from the knowledge base consume significant computing resources, prolong retrieval time, and negatively impact overall detection efficiency. Therefore, in some embodiments, a caching mechanism can be implemented to optimize the knowledge retrieval process and improve retrieval efficiency. Specifically, the network attack detection device 102 can statistically analyze the historical access frequency of each piece of knowledge in the knowledge base, identifying knowledge with a historical access frequency exceeding a preset frequency threshold as high-frequency knowledge. This high-frequency knowledge is then stored in a preset cache module. When performing subsequent network attack detection and retrieving related knowledge from the knowledge base, high-frequency knowledge is retrieved first from this cache module, eliminating the need to repeatedly access the original knowledge base. This effectively reduces resource consumption from repeated retrievals and significantly shortens the knowledge retrieval time.

[0092] Meanwhile, in some embodiments, in order to adapt to the detection needs of massive network data and ensure the stable progress of the detection process, dynamic resource scheduling can also be performed. Specifically, the network attack detection device 102 monitors the current network traffic load in real time and dynamically adjusts the concurrency of the first detection model and the second detection model according to the real-time load. When the traffic peak exceeds the preset threshold, some preliminary attack detection tasks are automatically diverted to the regional layer server to avoid overload and lag of the detection device on a single node, ensuring that the entire detection process can proceed efficiently and stably.

[0093] Corresponding to the aforementioned embodiments of the source tracing and evidence collection method, this application also provides a source tracing and evidence collection device 103, including a memory, a processor, and a computer program stored in the memory and executable on the processor; wherein, when the processor executes the computer program, it implements the steps of the source tracing and evidence collection method described in any of the above embodiments.

[0094] For example, processors include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs).

[0095] For example, the memory may include at least one type of storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc.

[0096] Figure 5 This is a schematic diagram illustrating the structure of a traceability and evidence collection device 103 according to an exemplary embodiment of this application. Figure 5 As shown, at the hardware level, the computer device includes a processor 501, an internal bus 502, a network interface 503, memory 504, and non-volatile memory 505, and may also include other hardware required for business operations. One or more embodiments of this application can be implemented in software, for example, the processor 501 reads the corresponding computer program from the non-volatile memory 505 into memory 504 and then runs it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0097] Corresponding to the embodiments of the aforementioned source tracing and evidence collection methods, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the source tracing and evidence collection methods described in any of the above embodiments.

[0098] Corresponding to the embodiments of the aforementioned source tracing and evidence collection methods, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the source tracing and evidence collection methods described in any of the above embodiments.

[0099] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention filed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the foregoing claims.

[0101] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0102] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for tracing forensics of a network attack, characterized in that, include: Acquire network attack data, wherein the network attack data is data that has been pre-determined to have network attack characteristics and can characterize the attributes of the current network attack; Retrieve information collection association knowledge that matches the attack type corresponding to the network attack data from a preset knowledge base. The information collection association knowledge includes: historical information collection dimensions and historical information collection tools for historically attacked devices in historical source tracing and forensics cases that match the attack type. Using a pre-defined large model, based on the network attack data and the information collection association knowledge, the information collection dimensions and information collection tools of the currently attacked device pointed to by the network attack data are determined; The information collection tool is pushed to the attacked device so that the attacked device runs the information collection tool according to the information collection dimension to collect the device information of the attacked device. The device information can characterize the operating status of the attacked device under the network attack. Retrieve source tracing and forensics-related knowledge that matches the device information from the knowledge base. The source tracing and forensics-related knowledge includes: historical device information in historical source tracing and forensics cases that match the device information, as well as historical network attack data and historical source tracing and forensics results corresponding to the historical device information. Using the large model, based on the device information, the source tracing and evidence collection association knowledge, and the network attack data, source tracing and evidence collection results are output. The source tracing and evidence collection results include: verification results to confirm whether the network attack features of the network attack data actually exist; determination results to determine whether there are potential network attack features on the attacked device and the attack type corresponding to the potential network attack features; source tracing results to trace the attack path of the network attack; and evidence for the verification results, the determination results, and the source tracing results, respectively.

2. The method of claim 1, wherein, The method further includes: The security assessment of the source tracing and evidence collection results is performed using a preset security assessment model. If the source tracing and evidence collection results meet the preset security assessment criteria, a source tracing and evidence collection report is generated based on the source tracing and evidence collection results. The source tracing and evidence collection report includes the source tracing and evidence collection results, as well as the degree of harm of the network attack determined based on the source tracing and evidence collection results and the corresponding security protection recommendations. If the source tracing and evidence collection results do not meet the preset security assessment criteria, the source tracing and evidence collection method will be re-executed until the source tracing and evidence collection results meet the preset security assessment criteria, or the number of times the source tracing and evidence collection method is re-executed reaches the preset maximum number of executions.

3. The method of claim 1, wherein, Using a pre-defined large model, based on the network attack data and the information collection correlation knowledge, the information collection dimensions and tools of the currently attacked device pointed to by the network attack data are determined, specifically including: If the knowledge base contains information collection association knowledge that matches the attack type of the network attack data, then the big model will determine the historical information collection dimension and historical information collection tool for the historically attacked device in the information collection association knowledge as the information collection dimension and the information collection tool. If the knowledge base does not contain any information collection-related knowledge that matches the attack type of the network attack data, then the large model will determine all information collection tools in the preset tool library as the information collection tools, and determine all information collection dimensions that can be covered by the information collection tools as the information collection dimensions.

4. The method of claim 1, wherein, The device information is structured data; Before outputting the source tracing and forensics results using the large model based on the device information, the source tracing and forensics association knowledge, and the network attack data, the process also includes: Based on the relationships between various historical device information stored in the knowledge base, the device information is processed into the structured data.

5. The method of claim 4, wherein, The associations include one or more of the following: the association between process running information and operation logs, and the association between file information and network transmission traffic.

6. The method according to claim 1, characterized in that, The network attack data includes one or more of the following: the IP address of the attacked device, the name of the network attack, the type of the network attack, the load information corresponding to the network attack, and the network context message when the network attack occurred.

7. The method of claim 1, wherein, The device information includes one or more of the following: device configuration information, process running information, operation logs, and temporary files generated during the network attack of the attacked device.

8. A cyber attack forensics system, characterized by, include: Several network devices, wherein the network devices are devices in the network that may be subject to network attacks; Several network attack detection devices are communicatively connected to the network device to detect whether the network device is under network attack, and when a network attack is detected, to identify and extract network attack data with network attack characteristics. The source tracing and evidence collection device is communicatively connected to the network device and the network attack detection device, respectively, and is used to execute the source tracing and evidence collection method for network attacks as described in any one of claims 1 to 7.

9. The system of claim 8, wherein, The source tracing and evidence collection device is integrated into the network attack detection device, or into other preset network devices.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the network attack tracing and forensics method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network evidence obtaining and attack chain reconstruction method based on threat alarm

    CN120915528A

  • Network attack tracing method, device, equipment and medium

    CN121173491A