Safety traceability method and system, storage medium and terminal equipment

By automating the processing of alarm event chains using large language models, the high cost and security risks caused by relying on manual analysis in existing technologies are solved, achieving efficient and accurate security tracing analysis.

CN120974481APending Publication Date: 2025-11-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410605586.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing security tracing methods rely on manual analysis by professional security experts, which increases labor costs and cybersecurity risks. Furthermore, insufficient expertise may lead to reduced system security.

Method used

The system employs a large language model to automate the processing of alarm event chains. Through scenario analysis, investigation steps, and function call code, it achieves automated source tracing analysis of alarm events.

Benefits of technology

It improves the efficiency and accuracy of security traceability, reduces reliance on professional security experts, lowers labor costs, and enhances system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974481A_ABST
    Figure CN120974481A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a safety traceability method and system, a storage medium and terminal equipment, and is applied to the technical field of information processing. When the alarm event chain is obtained, scene analysis is directly carried out according to at least one event and alarm information in the alarm event chain to obtain a candidate scene set, and then an investigation step comprising a plurality of investigation operations is generated according to the candidate scene set. And generating candidate reasoning information for selecting a final scene from the candidate scene set based on investigation results of the multiple investigation operations, generating a function call code according to the investigation steps and the candidate reasoning information, and finally determining a candidate scene in the candidate scene set according to an operation result of the function call code. And taking as a real scene of the alarm event chain. In the whole process, safety traceability of each alarm event can be automatically realized, and the safety traceability efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, in particular to a security traceability method and system, a storage medium and a terminal device. BACKGROUND

[0002] In the field of network security, potential security threats or abnormal behaviors can generally be detected by network security products, and security alarms can be given. The content of the alarm is different according to the different protection objects, such as Endpoint Detection and Response (EDR). The EDR scheme mainly monitors and responds to threats to computer endpoints (such as workstations, servers, etc.). Specifically, endpoint data is collected in real time, and certain analysis methods are used to detect, investigate and prevent potential threats. That is, through security traceability analysis, the defense capability of an organization against complex security threats can be improved.

[0003] An existing security traceability method mainly involves professional security experts first interpreting multiple event contents in alarm information, analyzing attack behaviors occurring in a real network environment, and investigating and collecting evidence in combination with system log information, etc. to ultimately determine the source of a network attack. This requires a large degree of manual input from professional security experts, which puts high demands on the technical ability of security experts and increases the difficulty and labor cost of security traceability. Moreover, if the professional ability of a security expert is insufficient in some aspects, it may increase the network security risks faced by the system. SUMMARY

[0004] The embodiments of the present application provide a security traceability method, system, storage medium and terminal device, which realize automatic security traceability analysis of alarm events.

[0005] In one aspect, the embodiments of the present application provide a security traceability method, comprising:

[0006] obtaining an alarm event chain, the alarm event chain comprising at least one event and alarm information arranged in order of occurrence;

[0007] performing scenario analysis according to the at least one event and alarm information in the alarm event chain to obtain a candidate scenario set, the candidate scenario set comprising multiple candidate scenarios in which the alarm event chain is located;

[0008] generating an investigation step and candidate reasoning information according to the candidate scenario set, the investigation step being used to describe multiple investigation operations, and the candidate reasoning information being used to select information of a final scenario from the candidate scenario set according to investigation results of the multiple investigation operations;

[0009] According to the investigation operation in the investigation step and the candidate reasoning information generation function call code;

[0010] According to the running result of the function call code, a candidate scene in the candidate scene set is determined as the real scene of the alarm event chain.

[0011] Another aspect of the embodiment of the application provides a safety traceability system, comprising:

[0012] An event chain acquisition unit is configured to acquire an alarm event chain, wherein the alarm event chain comprises at least one event and alarm information arranged in an occurrence order;

[0013] A scene analysis unit is configured to perform scene analysis according to the at least one event and the alarm information in the alarm event chain, and obtain a candidate scene set, wherein the candidate scene set comprises a plurality of candidate scenes in which the alarm event chain is located;

[0014] A call generation unit is configured to generate an investigation step and candidate reasoning information according to the candidate scene set, wherein the investigation step is used to describe a plurality of investigation operations, and the candidate reasoning information is used to select a final scene from the candidate scene set according to an investigation result of the plurality of investigation operations;

[0015] A code generation unit is configured to generate a function call code according to the investigation operation in the investigation step and the candidate reasoning information;

[0016] A real determination unit is configured to determine a candidate scene in the candidate scene set as a real scene of the alarm event chain according to a running result of the function call code.

[0017] Another aspect of the embodiment of the application further provides a computer readable storage medium, wherein the computer readable storage medium stores a plurality of computer programs, and the computer programs are suitable for being loaded and executed by a processor to implement the safety traceability method according to the aspect of the embodiment of the application.

[0018] Another aspect of the embodiment of the application further provides a terminal device, comprising a processor and a memory;

[0019] The memory is configured to store a plurality of computer programs, and the computer programs are used to load and execute the safety traceability method according to the aspect of the embodiment of the application by the processor; and the processor is configured to implement each computer program in the plurality of computer programs.

[0020] It can be seen that in the method of the embodiment, when the alarm event chain is obtained, the following operations are directly performed according to the corresponding large language model: performing scene analysis according to at least one event in the alarm event chain and the alarm information to obtain a candidate scene set, generating an investigation step including a plurality of investigation operations and candidate reasoning information according to the candidate scene set, and generating a function call code according to the investigation step and the candidate reasoning information, and finally determining a candidate scene in the candidate scene set as a real scene of the alarm event chain according to a running result of the function call code. In the whole process, the safety traceability of each alarm event can be automatically realized through the large language model, and the efficiency and accuracy of safety traceability are improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0022] Figure 1 is a schematic diagram of a safety traceability method provided by an embodiment of the present application;

[0023] Figure 2 is a flowchart of a safety traceability method provided by an embodiment of the present application;

[0024] Figure 3 is a flowchart of a safety traceability method provided by another embodiment of the present application;

[0025] Figure 4 is a flowchart of a safety traceability method provided by an application embodiment of the present application;

[0026] Figure 5a is a schematic diagram of an alarm event chain in an application embodiment of the present application;

[0027] Figure 5b is a schematic diagram of another alarm event chain in an application embodiment of the present application;

[0028] Figure 6 is a schematic diagram of an event and an alarm in an application embodiment of the present application;

[0029] Figure 7a is a schematic diagram of a scene analysis prompt in an application embodiment of the present application;

[0030] Figure 7b is a schematic diagram of an investigation prompt in an application embodiment of the present application;

[0031] Figure 8 is a schematic diagram of a scenario analysis interface in one application embodiment of the present application;

[0032] Figure 9 is a schematic diagram of a function investigation code in one application embodiment of the present application;

[0033] Figure 10 is a schematic diagram of a logical structure of a secure traceability system provided by an embodiment of the present application;

[0034] Figure 11 is a schematic diagram of a logical structure of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0036] The terms "first", "second", "third", "fourth" and the like (if any) in the description, claims and above drawings of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to include those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0037] The embodiments of the present application provide a secure traceability method, which is mainly applied to traceability analysis of an alarm event in any network system, so as to find out an actual scenario where the alarm event is located, such as Figure 1 As shown in the figure, when the network security system 11 is applied to any network device 10 (such as a terminal device or a server), the network security system 11 will alarm a network security problem according to certain alarm rules during the running of the network device 10, and the alarm information can be generally stored in the log of the network device 10.

[0038] In the embodiment, a secure traceability system 12 can be deployed, which can be applied to the network device 10 or other devices independent of the network device 10 Figure 1With the security tracing system 12 being independent of the network device 10 as an example, the security tracing system 12 can perform security tracing analysis through the following steps:

[0039] obtain an alarm event chain, the alarm event chain including at least one event and alarm information arranged in an occurrence order; perform scenario analysis according to the at least one event and the alarm information in the alarm event chain to obtain a candidate scenario set, the candidate scenario set including a plurality of candidate scenarios in which the alarm event chain is located; generate an investigation step and candidate reasoning information according to the candidate scenario set, the investigation step being used to describe a plurality of investigation operations, and the candidate reasoning information being used to select a final scenario from the candidate scenario set according to investigation results of the plurality of investigation operations; generate function call code according to the plurality of investigation operations in the investigation step and the candidate reasoning information; and determine a candidate scenario in the candidate scenario set as a real scenario of the alarm event chain according to a running result of the function call code.

[0040] It can be seen that the security tracing of each alarm event can be automatically implemented in the whole process, and the efficiency and accuracy of security tracing are improved.

[0041] The security tracing method provided in an embodiment of the present application is mainly applied to the above-mentioned scenario, and the method performed by the security tracing system 12 includes the following steps. Figure 1 Figure 2

[0042] In step 101, an alarm event chain is obtained, and the alarm event chain includes at least one event and alarm information arranged in an occurrence order.

[0043] It can be understood that when the network security system 11 is applied to any network device 10 (such as a terminal device or a server), the network security system 11 can perform alarm on network security problems according to certain alarm rules during the running of the network device 10, and the alarm information can be generally stored in the log of the network device.

[0044] For example, for the alarm rule “persistence: registry and startup folder”, when a process in the network device attempts to modify the registry or the startup folder, it may be that malicious software in the network device is performing persistence, and an alarm of persistence can be performed; for the alarm rule “disguise means: abnormal file name, common behavior of phishing attacks”, when a file with an abnormal file name appears in the running environment of the network device, such as “resume.pdf.exe” and the like, it may be that a deceptive file name is used in a phishing attack, and an alarm of disguise means can be performed.

[0045] ​​In one implementation, the network security system 11 can generate an alert based on log information of the network device 10, where the log information of the network device 10 can include process behavior information, traffic information, etc. In this way, the network security system 11 can determine whether a condition for triggering an alert is met through a preset algorithm, and the algorithm can determine that an alert needs to be generated through only a single piece of information (information of a single operation step), such as a case where a process is found to create a startup item in an EDR device, indicating that a malware infection is possible, and an alert is generated. In this case, the alert event chain obtained includes information of one event and one alert, which can also be referred to as an original alert.

[0046] For another example, the network security system 11 can determine that an alert is needed through a sequence of information (such as a sequence of single alerts or a sequence of process behaviors), such as a case where process behavior information is found in an EDR device, including the following operations: a user opens a mail application, generating an alert; the mail application downloads a file, generating another alert; the user double-clicks to run the downloaded file; the downloaded file attempts to scan the intranet, indicating that the user is likely to be subjected to a phishing attack, generating another alert. In this case, multiple alerts are generated, and the alert event chain obtained includes information of multiple events (sorted according to the order of occurrence) and multiple alerts involved in this process.

[0047] For a specific example: a user opens API Fox software (Apifox.exe) through the resource manager (explorer.exe) in a Windows system, and because Apifox.exe attempts to add itself to the system startup item (so-called boot self-starting), and boot self-starting is a feature of malware, the network security system 11 suspects that Apifox.exe can be a malware, and creates an alert. In the process of creating the alert, only one event is passed, that is, opening the API Fox software, and therefore the alert event chain includes only one event, and the event involves two processes, explorer.exe and Apifox.exe. The corresponding alert rule is “Persistence: registry and startup folder”.

[0048] Specifically, when the alert event chain is obtained, information of each event in the alert event chain can be obtained, such as a timestamp of the event, a process tree (including a process name, a startup path of the process, a hash value of an executable file, signature information, etc.) involved in the event, asset information (a host name, an IP address, responsible person information, etc.), etc. Alert information corresponding to the alert event chain is also obtained, such as a file path of a suspected malware that can be additionally obtained for an alert of a malware infection; an IP address and a port of a remote device that can be additionally obtained for an alert of remote task addition, but no process tree information is obtained.

[0049] Step 102, performing scenario analysis according to at least one event in the alarm event chain and the alarm information to obtain a candidate scenario set, the candidate scenario set including multiple candidate scenarios in which the alarm event chain is located.

[0050] Here, the candidate scenario refers to an actual scenario in which each event in the alarm event chain is located, such as a scenario of actual operation by a user or a scenario of malicious software infection, and any analysis granularity (such as process granularity of an event containing a process, event granularity, or IP or port calling granularity) can be used to analyze the candidate scenario. Specifically, any attribute involved in the alarm event chain can be used to divide the alarm event chain into analysis units, and each analysis unit is analyzed to realize overall scenario analysis. Thus, before step 102 is performed, the analysis granularity of scenario analysis can be determined according to the attribute information of the alarm event chain, and then scenario analysis is performed according to at least one event in the alarm event chain, the alarm information, and the analysis granularity.

[0051] The attribute information of the alarm event chain can include, but is not limited to, the following attributes: the number of processes contained in the events in the alarm event chain, the process tree, the text length of the text corresponding to the alarm event chain, and the IP or port involved in the events in the alarm event chain.

[0052] When the analysis granularity is determined, if the number of processes contained in the events in the alarm event chain or the number of process trees is less than a preset value, the analysis granularity of scenario analysis is determined to be process granularity, and if the number of processes contained in the events or the number of process trees is greater than or equal to the preset value, the analysis granularity of scenario analysis is determined to be event granularity. If the events do not contain processes but involve interface (such as IP or port) calling, the analysis granularity of scenario analysis is determined to be interface granularity. Alternatively, the analysis granularity of scenario analysis is determined to be text granularity according to the text length of the text corresponding to the alarm event chain, and the text length of the analysis unit based on the text granularity does not exceed a preset length.

[0053] Further, after the analysis granularity is determined, the analysis unit for scenario analysis can be obtained according to the analysis granularity. For example, if the analysis granularity is process granularity, the analysis unit is the information of each process contained in each event in the alarm event chain, and then each analysis unit is analyzed to obtain the candidate scenario of the entire alarm event chain.

[0054] When the process granularity is used for scenario analysis, specifically:

[0055] The description information of at least one process contained in each event in the alarm event chain can be determined, and the multiple candidate scenarios in which the alarm event chain is located can be determined according to the description information of the at least one process corresponding to each event and the alarm information.

[0056] Wherein, if at least one attribute description information is included in the description information of any process contained in each event, and the first process involved in the alarm information is included in the alarm description information of the first process, which is any process contained in each event, in determining each candidate scenario, specifically, at least one attribute description of any process involved in the alarm event chain can be selected to obtain the selected attribute description of any process, and then the selected attribute description of any process involved in the alarm event chain and the alarm description based on the first process are combined to obtain the combined attribute description of the alarm event chain, and then a corresponding candidate scenario is determined according to the combined attribute description. Wherein, the attribute description information refers to the information describing the attribute of the process (such as the function attribute, etc.).

[0057] For example: an alarm event chain includes n events T1, T2, …, Tn, any process contained in each event is Tn-Pi (i is any natural number from 1 to x, x is the number of processes contained in the event), and the selected attribute description Qj (j is any natural number from 1 to y, y is the number of attributes involved in the process) is obtained for each process Tn-Pi. The first process involved in the alarm information is Tn-Px1. The selected attribute description Qj of all processes involved in the alarm event chain and the alarm description of the first process Tn-Px1 are combined according to the occurrence order to obtain the combined attribute description, and then a candidate scenario can be determined.

[0058] Specifically, for example, an alarm event chain includes an event involving two processes explorer.exe and ApifoxAppAgent.exe, and the first process involved in the alarm is ApifoxAppAgent.exe.

[0059] If one attribute description "is a Windows file manager, usually directly or indirectly started by the user" of the process explorer.exe and one attribute description "the process may be generated by the user's API Fox tool, used to manage and protect the developer's tool" of the process ApifoxAppAgent.exe are selected, and the alarm description "ApifoxAppAgent.exe is found in the alarm directory, indicating that it may be trying to persist or obtaining unauthorized privileges on the system" based on the first process is combined, the corresponding candidate scenario can be determined as "user active operation".

[0060] If a property description "is Windows Explorer, usually launched by user directly or indirectly" in explorer.exe is selected, a property description "this process may be part of malware, which may try to deceive users by disguising as a legitimate application" in ApifoxAppAgent.exe is selected, and the alarm description based on the first process is "ApifoxAppAgent.exe is located in the alarm directory, which is an obvious signal that it attempts to achieve persistence or obtain unauthorized privileges on the system, in which case the process may have caused different degrees of damage to the system, including but not limited to data theft and remote control", a corresponding candidate scenario "malware infection" can be determined by combining the three.

[0061] In step 103, investigation steps and candidate reasoning information are generated according to the candidate scenario set. The investigation steps are used to describe a plurality of investigation operations, and the candidate reasoning information is used to select the final scenario from the candidate scenario set according to the investigation results of the plurality of investigation operations.

[0062] Here, the investigation steps are mainly used to determine whether a candidate scenario in the candidate scenario set is the actual scenario of the alarm event chain, and the steps required to obtain other specific information. Thus, when the investigation steps are generated, for each candidate scenario in the candidate scenario set, the corresponding investigation operation is generated as: an operation of obtaining specific information used to determine whether the candidate scenario is the final scenario, and these generated investigation operations are combined to form the investigation steps. In the process of obtaining the specific information, the specific information can be obtained based on the alarm information obtained in step 101.

[0063] The candidate reasoning information is mainly used to describe information that a candidate scenario in the candidate scenario set is the final scenario according to the investigation results of the respective investigation operations.

[0064] For example, the candidate scenarios included in the candidate scenario set are: scenario 1 harmless behavior of the host's responsible person, scenario 2 harmless behavior of software or system, and scenario 2 real attack behavior.

[0065] For scenario 1, the malicious process is created by the host responsibility person himself, although the alarm is triggered, but the behavior itself is harmless, which belongs to the behavior of the host responsibility person, for example, the host responsibility person is performing security testing, learning command line functions, etc., which is a false alarm. In this case, it is necessary to confirm whether the alarm behavior is the host responsibility person himself, and then generate the investigation operation of "inquiring the responsibility person", specifically, the operation of obtaining the confirmation information of the responsibility person for the bcdedit.exe operation, to determine whether the bcdedit.exe operation in the alarm is the behavior of the host responsibility person himself, such as prompting the user "is the bcdedit.exe operation in the alarm your own behavior?". If the user answers "yes", it can be confirmed that scenario 1 is the actual scenario, and if the answer is "no", this scenario 1 can be excluded.

[0066] For scenario 2, it is necessary to determine whether the alarm behavior is an expected system maintenance behavior, and then generate the investigation operation of "inquiring the responsibility person", specifically, in the case where it is confirmed that the alarm behavior is not the host responsibility person himself, prompting the user "is the bcdedit.exe operation in the alarm your expected system maintenance behavior".

[0067] In this case, the candidate reasoning information includes: if the investigation operation for scenario 1 is "yes", it is judged that the final scenario is scenario 1, i.e. the harmless behavior of the host responsibility person; if the investigation operation for scenario 1 is "no" and the investigation operation for scenario 2 is "yes", it is judged that the final scenario is scenario 2, i.e. the harmless behavior of software or system; if the investigation operation for scenario 1 is "no" and the investigation operation for scenario 2 is "no", it is judged that the final scenario is scenario 3, i.e. the real attack behavior.

[0068] It should be noted that in addition to the above "inquiring the responsibility person", the investigation operation can also include but is not limited to the following operations: querying the process list, querying the malware list, inquiring the operation and maintenance personnel, etc.

[0069] Step 104, generating function call code according to the multiple investigation operations in the investigation step and the candidate reasoning information.

[0070] Here, the function call code refers to the code that can be directly run by the security traceability system 12, and the function call code is used to obtain the investigation result through multiple investigation operations, and select the final scenario from the candidate scenario set according to the investigation result. When the function call code is generated, the function call code can be directly run, and the running result is obtained.

[0071] Step 105, determining a candidate scenario in the candidate scenario set as the real scenario of the alarm event chain according to the running result of the function call code.

[0072] As can be seen, in the method of the embodiment, when the security traceability system 12 obtains an alarm event chain, the scenario analysis is directly performed according to at least one event in the alarm event chain and the alarm information, the candidate scenario set is obtained, the investigation steps including a plurality of investigation operations are generated according to the candidate scenario set, the candidate reasoning information for selecting the final scenario from the candidate scenario set based on the investigation results of the plurality of investigation operations is generated, the function call code is generated according to the investigation steps and the candidate reasoning information, and finally, according to the running result of the function call code, a candidate scenario in the candidate scenario set is determined as the real scenario of the alarm event chain. The security traceability of each alarm event can be automatically realized in the whole process, and the efficiency and accuracy of the security traceability are improved.

[0073] Another embodiment of the present application provides a security traceability method, which is mainly a method performed by the security traceability system 12 described above, and is similar to the security traceability method shown in Figure 2 The difference is that in the method of the embodiment, a large language model is used for scenario analysis, generation of investigation steps and candidate reasoning information, and generation of function call code. The flow chart is shown in Figure 3 The specific steps include:

[0074] Step 201, obtaining an alarm event chain, the alarm event chain including at least one event and alarm information arranged in order of occurrence.

[0075] Step 202, obtaining a scenario analysis prompt, calling a preset first large language model, obtaining a candidate scenario set from the first large language model according to the scenario analysis prompt, and the candidate scenario set including a plurality of candidate scenarios.

[0076] It should be noted that in the embodiment, the scenario analysis is performed by the first large language model, which is a machine learning model based on artificial intelligence, and can be trained by a certain training method and pre-installed in the security traceability system 12.

[0077] Artificial intelligence (AI) is a theory, method, technology and application system for using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to design and implement principles and methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0078] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, machine learning and deep learning and other major directions.

[0079] Machine learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0080] The first large language model in this embodiment can be reinforced based on an open source large language model. The existing open source large language model is reinforced through fine-tuning and other methods to enhance its ability to analyze alarm event chains. This solution requires less computing power resources and supports private deployment, which can avoid privacy leakage.

[0081] The selection of the open source large language model does not necessarily choose a model with a large number of parameters (such as a model with a parameter quantity of 7B, 13B and above). Using a model with a parameter size of 1-2B can still achieve good analysis results. In this embodiment, a smaller parameter size model can be used, but the analysis effect can already approach a model with a parameter quantity of 7B, or even better, and because of the small parameter quantity, the inference speed is fast, and the throughput is greatly improved, making the alarm analysis efficiency greatly improved compared to large models with 7B and above, and the requirement for a graphics card is also very low, the pressure of private deployment is small, and the product cost is also greatly reduced.

[0082] It should be noted that if the first large language model is used for scenario analysis, the scenario analysis prompt needs to be obtained first, and then the scenario analysis prompt is input into the first large language model to obtain the candidate scenario set. The scenario analysis prompt (prompt) is a natural language suitable for the operation of the first language model.

[0083] Specifically, when the scenario analysis prompt is obtained, the analysis granularity of the scenario analysis can be determined according to the attribute information of the alarm event chain, and a prompt template corresponding to the analysis granularity can be determined according to the analysis granularity. Then, at least one event in the alarm event chain and the alarm information are formatted according to the prompt template to a format corresponding to the prompt template, and the scenario analysis prompt is obtained.

[0084] The information of the alarm event chain can be included in the scenario analysis prompt, and indication information can also be included, which is used to indicate specific details of the scenario analysis performed by the first large language model, such as analysis steps of the scenario analysis performed by the first large language model, or a preset candidate scenario.

[0085] In step 203, the investigation prompt is obtained according to the candidate scenario set and the alarm information, a preset second large language model is called, and the investigation steps and candidate reasoning information are obtained by the second large language model according to the investigation prompt.

[0086] In the investigation prompt, the candidate scenario set and the alarm information can be spliced to obtain the investigation prompt. The indication information can also be included in the investigation prompt, which is used to indicate specific details of the generation of the investigation steps and the candidate reasoning information by the second large language model, such as tools involved in each investigation operation in the investigation steps.

[0087] In step 204, the investigation code prompt is obtained according to the multiple investigation operations in the investigation steps and the candidate reasoning information, a preset third large language model is called, and the function call code is obtained by the third large language model according to the investigation code prompt.

[0088] In the investigation code prompt, the investigation steps and the candidate reasoning information can be spliced to obtain the investigation code prompt. The indication information can also be included in the investigation code prompt, which is used to indicate specific details of the generation of the function call code by the third large language model, such as tool information.

[0089] It should be noted that the second large language model and the third large language model are similar to the first large language model, and are all machine learning models based on artificial intelligence. The difference is that the functions of the models are different. The three large language models can be deployed separately or integrated into the same large language model.

[0090] In step 205, the function call code is run to obtain a running result.

[0091] It should be noted that in the process of executing the above steps 202 to 205, the security tracing system 12 can provide the network device 10 with a scenario analysis interface, which displays the above candidate scenario set, the investigation steps and the candidate reasoning information, and the running result of the function call code in sequence. Some information in the running process of the function call code can also be displayed in the scenario analysis result, such as the investigation result of some investigation operations, specific information obtained from the network device 10 based on some investigation operations, etc.

[0092] And in the process of running the function call code, in some cases, information needs to be obtained through interaction with the user of the network device 10, such as when the investigation operation is "ask the operation and maintenance personnel", etc., the user needs to make a selection according to the prompt information, and then obtain the selection information of the user.

[0093] Step 206, according to the running result of the function call code, a candidate scenario in the candidate scenario set is determined as the real scenario of the alarm event chain, thereby completing the security tracing analysis of the alarm event chain.

[0094] It can be seen that in the method of the embodiment, when the security tracing system 12 obtains the alarm event chain, it will directly perform the following operations according to the corresponding large language model: performing scenario analysis according to at least one event in the alarm event chain and the alarm information to obtain a candidate scenario set, generating investigation steps including a plurality of investigation operations and candidate reasoning information according to the candidate scenario set, and generating function call code according to the investigation steps and the candidate reasoning information, and finally determining a candidate scenario in the candidate scenario set as the real scenario of the alarm event chain according to the running result of the function call code. The security tracing of each alarm event can be automatically realized through the large language model in the whole process, which improves the efficiency and accuracy of security tracing.

[0095] The security tracing method of the present application will be described below with another specific application example. The security tracing method of the embodiment can be applied to the above-mentioned Figure 1 As shown in the figure, in this embodiment, the network device 10 is specifically a device applying an EDR network security solution, referred to as an EDR device, such as Figure 4 As shown in the figure, the security tracing method in this embodiment can include:

[0096] Step 301, obtaining the alarm event chain of the EDR device.

[0097] Specifically, it can be assumed in this embodiment that the alarm event chain L is an event information list E and triggered alarm information R, i.e. L=(E, R), wherein:

[0098] The event information list E is a set containing information about N events arranged in chronological order of occurrence, where N is greater than or equal to 1, meaning at least one event.

[0099] E = {e1, e2, ..., e N}

[0100] Events are recorded via EDR devices. Each event message 'e' is a structure that stores specific details of the event, such as timestamps, process trees (process name, process startup path, executable file hash, signature information, etc.), and asset information (hostname, IP address, responsible person information, etc.). The details stored may differ for different events. For example, an alert for malware infection might additionally store the file path of the suspected malware; an alert for remotely added tasks might additionally store the remote IP address and port, but not the process tree information.

[0101] The alarm information R can include information such as the alarm rules triggered by the alarm event. An alarm rule is a brief text description that is ultimately triggered by the execution of at least one event. For example, the alarm rule "Persistence: Registry and Startup Folder" indicates that a process is attempting to modify the registry or startup folder, which matches the characteristics of malware persistence; the alarm rule "Disguise: Abnormal Filename, Common Behavior in Phishing Attacks" indicates that a file with an abnormal filename, such as "resume.pdf.exe", has appeared in the environment, which matches the characteristic of using deceptive filenames in phishing attacks.

[0102] The following two scenarios are specific examples:

[0103] like Figure 5a Scenario a: A user in a Windows system opens the API Fox software (Apifox.exe) through Explorer (explorer.exe), causing the EDR device to suspect that Apifox.exe may be malware and create an alert. In this case, Event Information List E contains only one event, with process tree information, and the alert rule is "Persistence: Registry and Startup Folder".

[0104] like Figure 5bThe illustrated scenario b: the user in the Windows system, through the resource manager (explorer.exe) decompressed a compressed package, decompressed a malicious software "free game.exe" from the compressed package, and then ran "free game.exe" through the resource manager, causing the host to be code executed and detected by the EDR device, creating an alarm. In this case, the event information list E contains two events, there are 2 process tree information, and the alarm rule is "code execution: malicious file".

[0105] Step 302, determine the analysis granularity of the alarm event chain, and determine the corresponding prompt template according to the analysis granularity.

[0106] It can be understood that the alarm event chain obtained from the EDR device is generally stored in the form of a JSON structure, which is not convenient for large language models to understand, and therefore needs to be converted into natural language.

[0107] In different alarm event chains, the number of events may differ greatly, for example, most alarm event chains have only one alarm, but some alarm rules may have many corresponding process trees. However, the context length of large language models is limited, and the maximum context window length is commonly 2048, 4096 tokens. If the converted natural language input is too long, it will cause the large language model to not have enough tokens to output results.

[0108] Therefore, different analysis granularities are used in this embodiment, such as process granularity or event granularity. Specifically, a threshold of the number of processes or the number of process trees can be set in the security traceability system. When the number of processes or the number of process trees corresponding to each event is less than the threshold, process granularity is used for analysis, and the analysis granularity is fine; when the number of processes or the number of process trees corresponding to each event exceeds the threshold, event granularity is used, and the analysis granularity is coarse. After determining the analysis granularity, a prompt template corresponding to the analysis granularity can be selected.

[0109] For example Figure 6 If the alarm is close to the attack starting point, such as the host responsible person receiving a phishing email and downloading the phishing attachment therein, triggering the alarm, the alarm event chain is short, and process granularity can be used for analysis; if the alarm occurs after the attack is deepened, such as using multiple software or system exploits after running the phishing attachment, the alarm event chain is long, and there are many process trees, so event granularity is selected due to the limitation of the context. Among them, the process tree is composed of a process or a plurality of processes triggered in sequence.

[0110] In addition to the process granularity and the event granularity described above, the analysis granularity can also use other standards, such as whether the length of the text formatted in the following step 303 exceeds a threshold, and the like. If a large language model with a long context is used, for example, some existing models have a context length of 32K (32*1024=32768), the analysis granularity can also not be distinguished.

[0111] In step 303, the alarm event chain is formatted according to the prompt template to obtain a scenario analysis prompt, which is used to instruct the large language model to generate a candidate scenario set.

[0112] After determining the prompt template, the scenario analysis prompt formatted according to the prompt template can include the following two parts: analysis requirements and alarm information.

[0113] Among them:

[0114] (1) The analysis requirement is an analysis step that the large language model needs to follow. In the analysis requirement, the steps of scenario analysis and the analysis granularity are described.

[0115] Specifically, as shown in Figure 7a the steps of scenario analysis described in the analysis requirement can include:

[0116] The large language model is required to output the candidate scenario name that may occur, for example, the candidate scenario name is "user active operation", which refers to the operation of the host responsible person without being induced, which is a kind of false alarm scenario.

[0117] According to the required analysis granularity, it is described what specifically happens in this candidate scenario. For example, when the process granularity is selected for analysis, the large language model will analyze each process in the process tree according to the order of the processes in the process tree, and analyze how each process is created in this scenario. For example, some processes may be created by user operation, and some processes may be created because of vulnerability exploitation.

[0118] Further, the analysis requirement can also describe that all possible results of the scenario analysis are limited to a set, i.e., a candidate scenario set, and the large language model is required to select the possible candidate scenarios. Specifically, in the scenario analysis prompt, a set of optional candidate scenarios S = {s1, s2, s3} is described, and each candidate scenario describes the name of the candidate scenario and the content of the candidate scenario. The scenario analysis prompt requires the large language model to select a candidate scenario from the candidate scenario set as a candidate scenario that is actually possible to occur. For example, if the candidate scenario set S = {“host responsibility person’s harmless behavior”, “harmless behavior of software or system”, “real attack behavior”}, then the result of the scenario analysis of the large language model will select one or more from the three; the large language model can also be required to select not to select, and analyze each candidate scenario.

[0119] It should be noted that the set of optional candidate scenarios described in the scenario analysis prompt can be pre-set in the security tracing system 12. Specifically, the user can pre-set some related information (including the name and description information, etc.) of the candidate scenarios in the security tracing system 12 through the candidate scenario setting interface provided by the security tracing system 12.

[0120] (2) The alarm information refers to the content after the alarm event chain is converted into natural language, i.e., the scenario analysis prompt, which shows the key information for analyzing the alarm. It can be as shown in Table 1 or Table 2:

[0121]

[0122]

[0123] Table 1: The case where the alarm event chain includes a single event

[0124]

[0125]

[0126] Table 2: The case where the alarm event chain includes multiple events

[0127] Further, in other embodiments, in order to further assist the large language model in security tracing, other information such as information in other security systems can also be added to the scenario analysis prompt.

[0128] Step 304: calling the large language model, and the large language model performs scenario analysis according to the scenario analysis prompt to obtain a candidate scenario set.

[0129] Specifically, after inputting the above scenario analysis prompt into a large language model (LLM), a candidate scenario set can be obtained. Further, the large language model can provide a scenario analysis interface to the EDR device and display information of the candidate scenario set in the scenario analysis interface, as shown in Figure 8

[0130] To accelerate inference, the large language model can use a service deployment mode of vLLM (a service deployment framework of a large language model) + batch (indicating that multiple samples are packaged together and input into the large language model for inference each time).

[0131] In step 305, the alarm information and the candidate scenario set obtained in the above steps are spliced to obtain an investigation prompt, which is used to instruct the large language model to generate an investigation step and candidate inference information.

[0132] Specifically, the alarm information formatted in step 303 and the candidate scenario set can be spliced to obtain the investigation prompt. As shown in Figure 7b In the investigation prompt, the large language model can also be instructed to give the investigation step and the candidate inference information. Specifically:

[0133] (1) Investigation step: mainly the investigation step generated after obtaining multiple candidate scenarios through scenario analysis, which is used to analyze which candidate scenario is the real scenario. The investigation step generally has multiple investigation operations, each of which can correspond to a specified optional tool, and the name of the optional tool and the specific information that can be obtained by the tool can be described in the investigation prompt. For example, the description of the tool corresponding to an investigation operation “ask the responsible person” is shown in Table 3 below. In the investigation prompt, the investigation operation corresponding to different candidate scenarios can also be specified as shown in Table 4 below.

[0134]

[0135] Table 3

[0136]

[0137] Table 4

[0138] In this embodiment, since the candidate scenario set is known, the investigation operation of each candidate scenario in the candidate scenario set can be specified in the investigation prompt. Specifically, the investigation operation can be directly spliced in the investigation prompt, or the name of the candidate scenario can be extracted through a regular expression, and then the corresponding investigation operation is spliced.

[0139] In addition, in the investigation prompt, it can be limited that each investigation operation finally obtains “yes” or “no” as the investigation result, which is convenient for subsequent inference.​

[0140] (2) Candidate reasoning information: refers to how to determine the final scenario from the candidate scenario set according to the investigation results of the investigation operation, and requires the large language model to determine which candidate scenario corresponds to the investigation results of each investigation operation.

[0141] For ease of understanding, the following is a generation example. In this example, the scenario list S = S' = {"host responsibility person's harmless behavior", "software or system's harmless behavior", "real attack behavior"}, and through the answers to the two investigation steps, it can be determined which scenario occurred.

[0142] Step 306, after inputting the investigation prompt to the large language model, the corresponding investigation steps and candidate reasoning information are obtained. Further, the investigation steps and candidate reasoning information can also be displayed through the scenario analysis interface provided by the EDR device, as shown in Figure 8 .

[0143] Step 307, the alarm information, candidate scenario set, and investigation steps and candidate reasoning information are spliced to obtain the investigation code prompt, which is used to instruct the large language model to generate function call code, which can be Python code, etc.

[0144] Specifically, the function call code LLM Agent can be called through a Python function as an interface, and can be called in the following Table 5 call format:

[0145]

[0146] Table 5

[0147] Where the tool name and the content of the called tool are of string type, and the return value is of boolean type, as shown in Figure 9 , where:

[0148] (1) The tool name refers to the specific tool that needs to be used by LLM Agent. The optional tool list can include: "query process list", "query malware list", "ask operation and maintenance personnel", "ask host responsibility person", and other tools corresponding to investigation operations.

[0149] (2) The task content refers to the task target that needs to be completed by LLM Agent using the tool. When generating the task target, the large language model will generate questions that can be answered using similar "yes" / "no", "exist" / "not exist".

[0150] (3) bool type return value refers to the survey operation, because the problem can be answered by a simple "yes" / "no", the large language model will convert the result into the corresponding True / False return. The return value directly participates in the inference decision in the code.

[0151] In this embodiment, the large language model (LLM) is required to write function call code, which can alleviate the illusion problem of LLM, and will not appear inconsistent with the context before and after, forget key information, etc. In addition, in order to avoid the interference of other output content before and after the function call code with the code execution, the separator such as " is required before and after the function call code in the survey code prompt. <start> ”," <end>", enclose the generated function call code, and facilitate subsequent extraction of the function call code.

[0152] At step 308, after inputting the investigation code prompt into the large language model, the function call code is obtained. Further, the function call code can be extracted through a regular expression.

[0153] At step 309, the function call code is run to obtain a running result, and according to the running result, it is determined that the real scene corresponding to the alarm event chain is a candidate scene in the candidate scene set.

[0154] Specifically, the function call code can include the call code of each tool and the implementation code of each tool. The process of running the function call code is mainly the process of performing each investigation operation according to the tool name and the call content of each tool, and obtaining the investigation result of each investigation operation.

[0155] When the investigation result of each investigation operation is obtained, based on the candidate recommendation information, it can be determined which candidate scene in the candidate scene set is the final real scene, and the traceability of the alarm event chain is realized.

[0156] During the running of the function call code, each investigation operation in the process and the specific information obtained through each investigation operation can be provided through the above-mentioned scene analysis interface of the EDR device, as shown in the four investigation operations and the specific information obtained from the EDR device in each investigation operation. Figure 8

[0157] It can be seen that through the security traceability method in the embodiment, the following technical effects can be achieved:

[0158] High professionalism: by using a large language model with security expert knowledge, the security traceability process is more professional than that involving security operation personnel, the standard is more unified, and the dependence on professional personnel can be reduced, and the cost of security traceability is low.

[0159] High accuracy: the final security traceability is based on real investigation operations, not just relying on certain features of the alarm information.

[0160] Strong interpretability: the large language model interprets the alarm information in a natural language manner, and the interpretation process is transparent and has strong interpretability.

[0161] High analysis efficiency: since the scheme of the embodiment greatly reduces the dependence on manual operation, the degree of automation of security traceability is significantly improved, and even if a large number of alarm events occur in a short period of time, the analysis can be performed in time.

[0162] The embodiment of the application also provides a security traceability system, and a structure diagram thereof is shown in​ Figure 10 may specifically include:

[0163] An event chain acquisition unit 20 is configured to acquire an alarm event chain, which includes at least one event and alarm information arranged in an occurrence order.

[0164] A scenario analysis unit 21 is configured to perform scenario analysis on the at least one event and the alarm information in the alarm event chain acquired by the event chain acquisition unit 20, to obtain a candidate scenario set, which includes a plurality of candidate scenarios in which the alarm event chain is located.

[0165] A call generation unit 22 is configured to generate an investigation step and candidate reasoning information according to the candidate scenario set obtained by the scenario analysis unit 21, the investigation step being used to describe a plurality of investigation operations, and the candidate reasoning information being used to select a final scenario from the candidate scenario set according to investigation results of the plurality of investigation operations.

[0166] A code generation unit 23 is configured to generate a function call code according to the plurality of investigation operations in the investigation step and the candidate reasoning information obtained by the call generation unit 22.

[0167] A real determination unit 24 is configured to determine a candidate scenario in the candidate scenario set as a real scenario of the alarm event chain according to a running result of the function call code generated by the code generation unit 23.

[0168] Further, the security traceability system of the embodiment can further include:

[0169] An analysis granularity unit 25 is configured to determine an analysis granularity of the scenario analysis according to attribute information of the alarm event chain; and the scenario analysis unit 21 is specifically configured to perform scenario analysis according to the at least one event, the alarm information, and the analysis granularity in the alarm event chain.

[0170] The analysis granularity unit 25 is specifically configured to determine the analysis granularity of the scenario analysis as a process granularity if a number of processes or a number of process trees contained in the event in the alarm event chain is less than a preset value, determine the analysis granularity of the scenario analysis as an event granularity if the number of processes or the number of process trees contained in the event is greater than or equal to the preset value, and determine the analysis granularity of the scenario analysis as an interface granularity if the event contains an interface call.

[0171] The scenario analysis unit 21 is specifically configured to determine a prompt template corresponding to the analysis granularity; format at least one event and alarm information in the alarm event chain according to the prompt template to obtain a scenario analysis prompt; and call a preset first large language model to obtain the candidate scenario set according to the scenario analysis prompt.

[0172] In another aspect, the scenario analysis unit 21 is further configured to determine description information of at least one process contained in each event; and determine a plurality of candidate scenarios in which the alarm event chain is located according to the description information of at least one process corresponding to each event and the alarm information. In this regard, when determining the plurality of candidate scenarios, the scenario analysis unit 21 specifically includes at least one attribute description information in the description information of any process, and the alarm information includes a first process involved in the alarm. For any process involved in the alarm event chain, at least one attribute description of the any process is selected to obtain selected attribute description of the any process. The selected attribute description of the any process and the alarm description based on the first process are combined to obtain combined attribute description of the alarm event chain. The corresponding candidate scenario is determined according to the combined attribute description.

[0173] Further, the calling generation unit 22 is specifically configured to determine, for each candidate scenario in the candidate scenario set, a corresponding investigation operation as an operation of obtaining specific information used to determine whether the candidate scenario is a final scenario; and combine the generated investigation operations to form the investigation step.

[0174] The calling generation unit 22 is further configured to obtain an investigation prompt according to the candidate scenario set and the alarm information; and call a preset second large language model to obtain the investigation step and candidate reasoning information according to the investigation prompt.

[0175] The code generation unit 23 is specifically configured to obtain an investigation code prompt according to a plurality of investigation operations in the investigation step and the candidate reasoning information; and call a preset third large language model to obtain the function call code according to the investigation code prompt.

[0176] Further, the security traceability system in the embodiment can further include:

[0177] The interface providing unit 26 is configured to provide a scenario analysis interface, and sequentially display, in the scenario analysis interface, the candidate scenario set obtained by the scenario analysis unit 21, the investigation step and candidate reasoning information, and a running result of the function call code.

[0178] In the security traceability system of the embodiment, when the alarm event chain is acquired, a scenario analysis is directly performed according to at least one event in the alarm event chain and alarm information to obtain a candidate scenario set, and then investigation steps including a plurality of investigation operations are generated according to the candidate scenario set, candidate reasoning information for selecting a final scenario from the candidate scenario set based on investigation results of the plurality of investigation operations is generated, function call code is generated according to the investigation steps and the candidate reasoning information, and finally a candidate scenario in the candidate scenario set is determined as a real scenario of the alarm event chain according to a running result of the function call code. The security traceability of each alarm event can be automatically implemented in the whole process, and the efficiency and accuracy of the security traceability are improved.

[0179] The embodiment of the present application also provides a terminal device, a structure diagram of which is shown in Figure 11 The terminal device can have a large difference due to different configurations or performances, and can include one or more central processing units (CPUs) 30 (for example, one or more processors) and a memory 31, one or more storage media 32 (for example, one or more mass storage devices) storing application programs 321 or data 322. The memory 31 and the storage media 32 can be temporary storage or persistent storage. The programs stored in the storage media 32 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the terminal device. Further, the central processing unit 30 can be configured to communicate with the storage media 32 and execute a series of instruction operations in the storage media 32 on the terminal device.

[0180] Specifically, the application programs 321 stored in the storage media 32 include security traceability application programs, and the programs can include the event chain acquisition unit 20, the scenario analysis unit 21, the call generation unit 22, the code generation unit 23, the real determination unit 24, the analysis granularity unit 25 and the interface providing unit 26 in the security traceability system described above, and details are not described herein. Further, the central processing unit 30 can be configured to communicate with the storage media 32 and execute a series of operations corresponding to the security traceability application programs stored in the storage media 32 on the terminal device.

[0181] The terminal device can also include one or more power supplies 33, one or more wired or wireless network interfaces 34, one or more input and output interfaces 35, and / or one or more operating systems 323, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM and the like.

[0182] The steps performed by the secure traceability system in the above method embodiments can be based on the Figure 11 The structure of the terminal device is shown.

[0183] Further, the embodiments of the present application also provide a computer readable storage medium storing a plurality of computer programs, the computer programs being adapted to be loaded and executed by a processor to perform the secure traceability method performed by the secure traceability system.

[0184] In another aspect, the embodiments of the present application also provide a terminal device comprising a processor and a memory.

[0185] The memory is configured to store a plurality of computer programs, the computer programs being configured to be loaded and executed by the processor to perform the secure traceability method performed by the secure traceability system; and the processor is configured to implement each of the plurality of computer programs.

[0186] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer readable storage medium, which can include read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0187] The secure traceability method, system, storage medium and terminal device provided by the embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples in this paper; the above embodiment descriptions are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description should not be understood as limiting the present application.< / end> < / start>

Claims

1. A method of secure provenance, characterized in that, The method comprises the following steps: obtaining an alarm event chain, the alarm event chain comprising at least one event and alarm information arranged in order of occurrence; performing scenario analysis on the at least one event and alarm information in the alarm event chain to obtain a candidate scenario set, the candidate scenario set comprising a plurality of candidate scenarios in which the alarm event chain is located; generating an investigation step and candidate reasoning information according to the candidate scenario set, the investigation step being used to describe a plurality of investigation operations, and the candidate reasoning information being used to select a final scenario from the candidate scenario set according to investigation results of the plurality of investigation operations; generating function call code according to the plurality of investigation operations in the investigation step and the candidate reasoning information; determining a candidate scenario in the candidate scenario set as a real scenario of the alarm event chain according to a running result of the function call code.

2. The method of claim 1, wherein, Before the scenario analysis is performed on the at least one event and alarm information in the alarm event chain, the method further comprises the following steps: determining an analysis granularity of the scenario analysis according to attribute information of the alarm event chain; the scenario analysis is performed on the at least one event and alarm information in the alarm event chain according to the at least one event, alarm information and analysis granularity.

3. The method of claim 2, wherein, The analysis granularity of the scenario analysis is determined according to the attribute information of the alarm event chain, and specifically comprises the following steps: if a number of processes or a number of process trees contained in an event in the alarm event chain is less than a preset value, determining that an analysis granularity of the scenario analysis is a process granularity; if the number of processes or the number of process trees contained in the event is greater than or equal to the preset value, determining that the analysis granularity of the scenario analysis is an event granularity; if the event contains an interface call, determining that the analysis granularity of the scenario analysis is an interface granularity.

4. The method of claim 2, wherein, The scenario analysis is performed on the at least one event and alarm information in the alarm event chain according to the at least one event, alarm information and analysis granularity, and specifically comprises the following steps: determining a prompt template corresponding to the analysis granularity; formatting the at least one event and alarm information in the alarm event chain according to the prompt template to obtain a scenario analysis prompt; calling a preset first large language model, the first large language model obtaining the candidate scenario set according to the scenario analysis prompt.

5. The method of claim 1, wherein, The scenario analysis is performed on the at least one event and alarm information in the alarm event chain to obtain a candidate scenario set, and specifically comprises the following steps: determining description information of at least one process contained in each event; determining a plurality of candidate scenarios in which the alarm event chain is located according to the description information of at least one process corresponding to each event and the alarm information.

6. The method of claim 5, wherein, The description information of any process comprises at least one attribute description information, and the alarm information comprises a first process involved in the alarm; and the plurality of candidate scenarios in which the alarm event chain is located are determined according to the description information of at least one process corresponding to each event and the alarm information, and specifically comprises the following steps: selecting at least one attribute description of any process involved in the alarm event chain to obtain selected attribute description of the any process; Combining a selected attribute description of any process involved in the alarm event chain and an alarm description based on the first process, a combined attribute description of the alarm event chain is obtained; According to the combined attribute description, a corresponding candidate scenario is determined.

7. The method according to any one of claims 1 to 6, characterized in that, The investigation step generated according to the candidate scenario set specifically includes: For each candidate scenario in the candidate scenario set, the corresponding investigation operation is determined as an operation of obtaining specific information used to determine whether the candidate scenario is the final scenario; The generated investigation operation is combined to form the investigation step.

8. The method according to any one of claims 1 to 6, wherein, The investigation step and candidate reasoning information generated according to the candidate scenario set specifically include: According to the candidate scenario set and the alarm information, an investigation prompt is obtained; A preset second large language model is called, and the second large language model obtains the investigation step and candidate reasoning information according to the investigation prompt.

9. The method according to any one of claims 1 to 6, wherein, The function call code is generated according to the multiple investigation operations in the investigation step and the candidate reasoning information, specifically including: According to the multiple investigation operations in the investigation step and the candidate reasoning information, an investigation code prompt is obtained; A preset third large language model is called, and the third large language model obtains the function call code according to the investigation code prompt.

10. The method of any one of claims 1 to 6, wherein, The method further includes: Providing a scenario analysis interface, the scenario analysis interface sequentially displays the candidate scenario set, the investigation step and candidate reasoning information, and the running result of the function call code.

11. A secure provenance system, characterized in that, It includes: An event chain acquisition unit is configured to acquire an alarm event chain, the alarm event chain including at least one event and alarm information arranged in order of occurrence; A scenario analysis unit is configured to perform scenario analysis according to at least one event and alarm information in the alarm event chain to obtain a candidate scenario set, the candidate scenario set including multiple candidate scenarios in which the alarm event chain is located; A call generation unit is configured to generate an investigation step and candidate reasoning information according to the candidate scenario set, the investigation step being used to describe multiple investigation operations, and the candidate reasoning information being used to select a final scenario from the candidate scenario set according to the investigation results of the multiple investigation operations; A code generation unit is configured to generate function call code according to the multiple investigation operations in the investigation step and the candidate reasoning information; A real determination unit is configured to determine a candidate scenario in the candidate scenario set as a real scenario of the alarm event chain according to the running result of the function call code.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by the processor to perform the security tracing method according to any one of claims 1-10.

13. A terminal device, comprising: It includes a processor and a memory; The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the security tracing method according to any one of claims 1-10; the processor is used to realize each computer program in the plurality of computer programs.