SOAR script generation method, device and equipment for generating large model based on retrieval enhancement

Through the large-model method based on search enhancement generation, combined with the script knowledge base and the big model, the accurate SOAR scripts are dynamically generated, which solves the problems of high orchestration complexity and maintenance difficulties in the existing technology, and achieves efficient response to the threat of rapid change.

CN120336869APending Publication Date: 2025-07-18NSFOCUS INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510328187.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing SOAR script generation method is difficult to quickly adapt to changing security threats, the orchestration interaction is complex, and the maintenance and update of large-scale scripts are difficult to maintain and update, resulting in security protection lag and high costs.

Method used

The large model method based on search enhancement generation is adopted, and semantic analysis and intention recognition is performed by obtaining the user input script text description, the script knowledge base is used for matching and vectorized similarity calculation, and the large model is used to generate accurate SOAR scripts, and the knowledge base is dynamically expanded to adapt to rapidly changing threat scenarios.

Benefits of technology

Simplifies the complexity of interface orchestration, improves the accuracy and adaptability of generating scripts, reduces maintenance costs, and can better deal with rapidly changing security threats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336869A_ABST
    Figure CN120336869A_ABST
Patent Text Reader

Abstract

The invention provides an SOAR script generation method, device and equipment based on a retrieval enhancement generation large model, and the method comprises the steps: obtaining a script text description inputted by a user, and carrying out the semantic analysis and intention recognition of the script text description, and obtaining key intention information; the key intention information is matched with script text descriptions in first structured data stored in a script knowledge base, and the first structured data comprises script text descriptions generated based on processing logic of different scripts and corresponding script files; when it is determined that the target script text description with the similarity larger than a first similarity threshold exists, a target script file corresponding to the target script text description is inquired and output. By utilizing the method, the retrieval enhancement and understanding ability of the large model is combined, a more accurate SOAR script can be generated, and meanwhile, through the continuously expanded script knowledge base, the newly generated script is more suitable for a rapidly changing security threat scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security orchestration and automated response, and particularly to a method, apparatus, and device for generating SOAR playbooks based on a retrieval-augmented generation large model. Background Art

[0002] In the field of security orchestration and automated response, in a SOAR (Security Orchestration, Automation and Response) platform, SOAR playbooks help security operation teams quickly respond to various security threats by orchestrating security tools and processes. In the prior art, SOAR playbooks are usually generated based on interface orchestration or by using a fine-tuned large model to generate playbooks. However, in the face of evolving threats and complex cyberattacks, traditional playbook generation methods often struggle to quickly adapt to changes, and for the response handling of a certain type of security event, security expertise is often required, making orchestration more difficult.

[0003] Among them, SOAR playbooks usually involve the interaction of multiple response actions and tools. Using interface orchestration has high complexity, high learning costs, and low efficiency. While using a large model fine-tuned with existing SOAR playbooks for pre-trained models has difficulties in scalability and maintenance and updates. Because of the expansion of the security infrastructure and the continuous change of the threat environment, the number and complexity of SOAR playbooks will also be continuously updated. If the model is not fine-tuned in time, the generated playbooks cannot effectively respond to the latest threats, resulting in the lag of security protection, and frequent fine-tuning is costly. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, and device for generating SOAR playbooks based on a retrieval-augmented generation large model, which is used to solve the problems of high complexity in the orchestration interaction of existing SOAR playbooks and the inability of playbooks directly generated based on large models to adapt to newly emerging security events or respond to rapidly changing threats.

[0005] In a first aspect, an embodiment of this application proposes a method for generating a SOAR playbook based on a retrieval-augmented generation large model, including: Obtain the playbook text description input by the user, perform semantic parsing and intent recognition on the playbook text description to obtain key intent information; Match the key intent information with the playbook text descriptions in the first structured data stored in the playbook knowledge base. The first structured data includes playbook text descriptions generated based on the processing logics of different playbooks and the corresponding playbook files; When it is determined that there is a target playbook text description with a similarity greater than the first similarity threshold, query and output the target playbook file corresponding to the target playbook text description.

[0006] In some possible embodiments, the method further includes: When it is determined that there is no target script text description with a similarity greater than the first similarity threshold, splitting the corresponding key intent information into multiple key action node information; Matching the multiple key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base, where the second structured data includes script action text descriptions generated based on different script actions and corresponding script action information; When it is determined that there is a target script action text description with a similarity greater than the second similarity threshold for each key action node information, querying the respective target script action information corresponding to the target script action text description; Inputting the respective target script action information into a large model, and using the large model to generate a target script file according to a predefined template.

[0007] In some possible embodiments, the method further includes: When it is determined that there is no target action script text description with a similarity greater than the second similarity threshold for any key action node information, adjusting the key action node information; Matching the adjusted key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base again.

[0008] In some possible embodiments, the script knowledge base includes a first relational database and a first vector database, where the first relational database is used to store the first structured data, and the first vector database is used to store the first structured vector data obtained by vectorizing the script text descriptions in the first structured data; Among them, matching the key intent information with the script text descriptions in the first structured data stored in the script knowledge base includes: Vectorizing the key intent information to obtain the vectorized data of the key intent information; Calculating the similarity between the vectorized data of the key intent information and the first structured vector data in the script knowledge base.

[0009] In some possible embodiments, the script action knowledge base includes a second relational database and a second vector database, where the second relational database is used to store the second structured data, and the second vector database is used to store the second structured vector data obtained by vectorizing the script action text descriptions in the second structured data; Among them, matching the multiple key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base includes: Vectorize the multiple key action node information respectively to obtain the vectorized data of each of the multiple key action node information; Calculate the similarity between the vectorized data of the multiple key action node information and the second structured vector data in the script action knowledge base.

[0010] In some possible embodiments, the first structured data further includes a script ID, and the script ID is associated with the first structured vector data after vectorizing the corresponding script text description; The second structured data further includes a script action ID, and the script action ID is associated with the second structured vector data after vectorizing the corresponding script action text description.

[0011] In some possible embodiments, the first relational database in the script knowledge base is constructed in the following manner: Obtain the script text description and the script file from the static preset script library, and process the script text description and the script file into structured data and store it in the first relational database in the script knowledge base; Obtain the target script file generated by the large model from the dynamic custom script library and parse it to obtain the target script text description, and process the target script text description and the target script file into structured data and store it in the first relational database in the script knowledge base.

[0012] In some possible embodiments, the second relational database in the script action knowledge base is constructed in the following manner: Based on the different script files stored in the first relational database in the script knowledge base, parse and extract the script files to obtain a corresponding plurality of different script actions; Based on the plurality of different script actions, generate corresponding script action text descriptions and script action information; Process the script action text description and the script action information into structured data and store it in the second relational database in the script action knowledge base.

[0013] In a second aspect, an SOAR script generation device based on a retrieval-enhanced generation large model according to an embodiment number of the present application includes: An intent analysis module, configured to obtain the script text description input by the user, perform semantic parsing and intent recognition on the script text description, and obtain key intent information; An intent matching module, configured to match the key intent information with the script text descriptions in the first structured data stored in the script knowledge base, where the first structured data includes script text descriptions generated based on processing logics of different scripts and corresponding script files; A script output module, configured to query and output a target script file corresponding to the target script text description when it is determined that there is a target script text description whose similarity is greater than the first similarity threshold.

[0014] In a third aspect, an embodiment of the present application further provides a SOAR script generation device based on a retrieval-enhanced generative large model, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the SOAR script generation method based on the retrieval-enhanced generative large model according to any one of the above first aspects.

[0015] The present application provides a SOAR script generation method, apparatus and device based on a retrieval-enhanced generative large model. The method includes: obtaining a script text description input by a user, performing semantic parsing and intent recognition on the script text description to obtain key intent information; matching the key intent information with the script text descriptions in the first structured data stored in the script knowledge base, where the first structured data includes script text descriptions generated based on processing logics of different scripts and corresponding script files; querying and outputting a target script file corresponding to the target script text description when it is determined that there is a target script text description whose similarity is greater than the first similarity threshold. Using this method, the retrieval enhancement and understanding capabilities of the large model are combined, enabling more accurate SOAR scripts to be generated. At the same time, through the continuously expanding script knowledge base, the newly generated scripts are more suitable for rapidly changing security threat scenarios.

[0016] Other features and advantages of the present application will be described in the subsequent specification, and will, in part, be obvious from the specification, or be learned through the implementation of the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 Schematic diagram of the RAG architecture in the embodiments of the present application; Figure 2 Schematic diagram of the process of the SOAR script generation method based on the retrieval-augmented generation large model in the embodiments of the present application; Figure 3 Flowchart of matching key intent information in the embodiments of the present application; Figure 4 Flowchart of a method when there is no target script text description with a similarity greater than the first similarity threshold in the embodiments of the present application; Figure 5 Schematic diagram of the construction process of a second relational database in the embodiments of the present application; Figure 6 Flowchart of matching key action node information in the embodiments of the present application; Figure 7 Flowchart of a method when there is no target script action text description with a similarity greater than the second similarity threshold in the embodiments of the present application; Figure 8 Flowchart of script generation using the SOAR script generation method of a retrieval-augmented generation large model in the embodiments of the present application; Figure 9 Schematic diagram of a SOAR script generation device based on the retrieval-augmented generation large model in the embodiments of the present application; Figure 10 Schematic diagram of a SOAR script generation device based on the retrieval-augmented generation large model in the embodiments of the present application. Detailed implementation manners

[0019] To further illustrate the technical solutions provided in the embodiments of the present application, the following will be described in detail with reference to the accompanying drawings and specific implementation manners.

[0020] The SOAR script is a workflow that solidifies the security operation and response processes. It can automate and coordinate multiple tasks and tools when security incidents occur to effectively respond to and handle these incidents. In view of the problems of high complexity in SOAR script orchestration interaction in the prior art and the inability of scripts directly generated based on large models to adapt to newly emerging security incidents or cope with rapidly changing threats, the embodiments of the present application propose a SOAR script generation method based on a retrieval-augmented generation large model.

[0021] The method is based on the RAG (Retrieval-Augmented Generation) architecture, as Figure 1As shown in the figure, it includes a LUI (Language User Interface), a large model, and a database. It interacts with users through the LUI. The LUI supports natural language conversations. Then, through the large model and in combination with the accumulated database, it can quickly adjust and generate script files, and can adapt to security threat scenarios with high-frequency changes. In the embodiments of the present application, the large model can but is not limited to using existing large models in the field of network security and general large models, and no specific limitations are made in the embodiments of the present application.

[0022] In the embodiments of the present application, a SOAR script generation method based on a retrieval-enhanced generation large model, as Figure 2 shown, includes: Step 201, obtain the script text description input by the user, perform semantic parsing and intention recognition on the script text description, and obtain key intention information; Optimize before retrieval. First, the large model will be called to perform semantic parsing and intention recognition based on the context information when the user inputs and the user's original input, convert the script text description of the user's original input into a description suitable for retrieving the script knowledge base, and extract the user's key intention for generating a more accurate script.

[0023] Step 202, match the key intention information with the script text descriptions in the first structured data stored in the script knowledge base. The first structured data includes script text descriptions generated based on the processing logics of different scripts and the corresponding script files; The script knowledge base is constructed by a static preset script library and a dynamic custom script library, includes a large number of script text descriptions and the corresponding script files, and can store common attack response scripts and scripts generated by users according to customization.

[0024] Step 203, when it is determined that there is a target script text description with a similarity greater than the first similarity threshold, query and output the target script file corresponding to the target script text description.

[0025] When a target script text description with a similarity greater than the preset first similarity threshold is matched, it indicates that there is a corresponding script in the script knowledge base. Querying and outputting the corresponding script file in the script knowledge base can enable users to obtain accurate SOAR scripts.

[0026] In some possible embodiments, the above script knowledge base includes a first relational database, and the first relational database is used to store the first structured data and is constructed in the following two ways: Method 1: Obtain the script text description and script file from the static preset script library, process the script text description and script file into structured data, and store it in the first relational database in the script knowledge base; Method 2: Obtain the target script file generated by the large model from the dynamic custom script library and parse it to obtain the target script text description. After processing the target script text description and target script file into structured data, store it in the first relational database in the script knowledge base.

[0027] In the embodiments of this application, the script knowledge base stores common attack response scripts and scripts customized by users. With the use of the large model for script accumulation, the scripts in the script knowledge base will continue to increase and improve. By dynamically constructing the script knowledge base, more accurate and complex scripts can be generated.

[0028] As shown in Table 1 below, it is an example of processing the script text description and script file into the first structured data.

[0029] Table 1 First Structured Data

[0030] In some possible embodiments, the above script knowledge base further includes a first vector database, which is used to store the first structured vector data obtained by vectorizing the script text description in the first structured data, and the script ID in the first structured data is associated with the first structured vector data obtained by vectorizing the corresponding script text description. In the embodiments of this application, the method for vectorization includes, but is not limited to, using the existing text embedding model m3e-base, which is not specifically limited in the embodiments of this application.

[0031] In the embodiments of this application, match the key intent information with the script text description in the first structured data stored in the script knowledge base, as Figure 3 shown, including the following steps: Step 301: Vectorize the key intent information to obtain the vectorized data of the key intent information; Step 302: Calculate the similarity between the vectorized data of the key intent information and the first structured vector data in the script knowledge base.

[0032] Specifically, obtain all script text descriptions supported in the current script knowledge base. The script text descriptions include script scenes and script descriptions. Merge the script scenes in the script text descriptions as entities, and merge the script descriptions as keywords. After vectorization, calculate the similarity with the vectorized data of the key intent information. Among them, for the specific calculation method of similarity, it includes but is not limited to using the existing cosine similarity calculation method, which is not specifically limited in the embodiments of this application.

[0033] In some possible embodiments, the method is as Figure 4 shown, and further includes: Step 401, when it is determined that there is no target script text description with a similarity greater than the first similarity threshold, split the corresponding key intent information into multiple key action node information; Step 402, match the multiple key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base. The second structured data includes script action text descriptions generated based on different script actions and corresponding script action information; Step 403, when it is determined that there is a target script action text description with a similarity greater than the second similarity threshold for each key action node information, query the corresponding target script action information for the target script action text description; Step 404, input each target script action information into a large model, and use the large model to generate a target script file according to a predefined template.

[0034] When no target script text description with a similarity greater than the preset first similarity threshold is matched, it means that there is no corresponding script in the script knowledge base, and the user's input needs to be processed secondary. That is, the identified key intent information is split into multiple key action node information, and the key action node information represents the key script actions in the script.

[0035] In the embodiments of this application, input the target script action information into the large model, and the large model can generate a corresponding target script file according to a predefined template. Among them, the large model is trained according to a predefined template, and its training method includes the following steps: Step 1, obtain a sample set composed of multiple training samples. The training samples include script action information and target sample files; Step 2, input the training samples into the large model, learn the script action information in the training samples, and generate a script file according to a predefined template based on the script action information, and adjust the parameters of the large model with the goal of outputting the target script file in the training samples.

[0036] In some possible embodiments, the script action knowledge base includes a second relational database, and the second relational database is used to store the second structured data. As Figure 5 shown, the second relational database is constructed in the following manner: Step 501, based on different script files stored in the first relational database in the script knowledge base, parse and extract the script files to obtain a corresponding plurality of different script actions; Step 502, based on the plurality of different script actions, generate corresponding script action text descriptions and script action information; Step 503, process the script action text description and script action information into structured data and store them in the second relational database in the script action knowledge base.

[0037] The parsing and extracting of the script files to obtain a corresponding plurality of different script actions includes: parsing each script file and extracting the script action information in the script, where the script action information includes at least one of the following: action classification, parameter configuration, action type. As shown in Table 2 below, it is an example of processing script action text descriptions and script action information into second structured data.

[0038] Table 2 Second Structured Data

[0039] In some possible embodiments, the script actions are supported by application plug-ins. When a new extended application plug-in is installed and registered, based on the script actions supported by the application plug-in to be extended, corresponding script action text descriptions and script action information are generated, and the script action text descriptions and script action information are processed into structured data and stored in the second relational database in the script action knowledge base.

[0040] In some possible embodiments, the script action knowledge base further includes a second vector database, and the second vector database is used to store second structured vector data obtained by vectorizing the script action text descriptions in the second structured data, and the script action IDs in the second structured data are associated with the second structured vector data obtained by vectorizing the corresponding script action text descriptions. In the embodiments of the present application, the method for vectorization includes, but is not limited to, using the existing text embedding model m3e-base, and no specific limitation is made in the embodiments of the present application.

[0041] In the embodiments of the present application, the multiple key action node information is respectively matched with the script action text descriptions in the second structured data stored in the script action knowledge base. As Figure 6 shown, it includes the following steps: Step 601: Vectorize each of the multiple key action node information to obtain the vectorized data of each of the multiple key action node information. Step 602: Calculate the similarity between the vectorized data of the multiple key action node information and the second structured vector data in the script action knowledge base.

[0042] Specifically, obtain all the script action text descriptions supported in the current script action knowledge base. The script text description includes an action name and an action description. Combine the action names in the script action text description as entities, and combine the action descriptions as keywords. After vectorization, calculate the similarity with the vectorized data of the key action node information. The method of calculating the similarity with the above-mentioned calculation of the key intention information and the script text description is the same, including but not limited to using the existing cosine similarity calculation method, which is not specifically limited in the embodiments of the present application.

[0043] In some possible embodiments, the method is as Figure 7 shown, and further includes: Step 701: When it is determined that there is no target action script text description with a similarity greater than the second similarity threshold for any key action node information, adjust the key action node information. Step 702: Match the adjusted key action node information with the script action text description in the second structured data stored in the script action knowledge base again.

[0044] In the embodiments of the present application, if the adjusted key action node information still cannot find a matching target script action text description in the script action knowledge base, continue to adjust until all key action node information finds corresponding target script action text descriptions.

[0045] In some possible embodiments, if no matching target script action text description is found after adjustment, feedback to the user is provided, prompting the user to adjust the original input to obtain a more accurate custom script file.

[0046] In some possible embodiments, based on the echo instruction, the large model can also provide feedback on the previously generated script file, echoing the user's original input and the generated script file.

[0047] The following gives a complete flowchart of a SOAR script generation method for generating a script using a retrieval-enhanced generation large model in the embodiments of the present application, as Figure 8 shown, including: Step 1: Obtain the script text description input by the user. Step 2: Use the large model to perform semantic parsing and intention recognition on the script text description to obtain key intention information. Step 3: Vectorize the key intent information to obtain the vectorized data of the key intent information; Step 4: Calculate the similarity between the vectorized data of the key intent information and the first structured vector data in the script knowledge base; Step 5: Determine whether the similarity is greater than the first similarity threshold; Step 6a: When it is determined that there is vectorized data of the target script text description with a similarity greater than the first similarity threshold, obtain the script ID associated with the vectorized data of the target script text description, and query and output the corresponding script file in the first relational database of the script knowledge base based on the script ID.

[0048] Step 6b: When it is determined that there is no target script text description with a similarity greater than the first similarity threshold, use the large model to split the corresponding key intent information into multiple key action node information; Step 7: Vectorize each of the multiple key action node information to obtain the vectorized data of each of the multiple key action node information; Step 8: Calculate the similarity between the vectorized data of the multiple key action node information and the second structured vector data in the script action knowledge base; Step 9: Determine whether the similarity is greater than the second similarity threshold; Step 10a: When it is determined that there is a target script action text description with a similarity greater than the second similarity threshold for any key action node information, obtain the script action ID associated with the vectorized data of the target script action text description, and query the corresponding script action information in the second structured database of the script action knowledge base based on the script action ID; Step 10b: When it is determined that there is no target action script text description with a similarity greater than the second similarity threshold for any key action node information, adjust the key action node information, and match the adjusted key action node information again with the script action text description in the second structured data stored in the script action knowledge base; Step 11: Input the script action information corresponding to all the key action node information queried into the large model, and use the large model to generate a target script file according to a predefined template.

[0049] A SOAR script generation method based on a retrieval-enhanced generative large model in an embodiment of the present application can dynamically construct a script knowledge base and a script action knowledge base according to different customer environments, combines the retrieval enhancement and understanding capabilities of the large model, and can generate more accurate SOAR scripts. It not only simplifies the complex interaction of interface arrangement, but also, through the continuously expanding script knowledge base, enables the newly generated scripts to be more applicable to rapidly changing security threat scenarios.

[0050] Based on the same inventive concept, an embodiment of the present application also proposes a SOAR script generation device based on a retrieval-enhanced generative large model, as Figure 9 shown, including: An intent analysis module 901, configured to obtain the script text description input by the user, perform semantic parsing and intent recognition on the script text description, and obtain key intent information; An intent matching module 902, configured to match the key intent information with the script text descriptions in the first structured data stored in the script knowledge base, where the first structured data includes script text descriptions generated based on the processing logics of different scripts and the corresponding script files; A script output module 903, configured to query and output the target script file corresponding to the target script text description when it is determined that there is a target script text description with a similarity greater than the first similarity threshold.

[0051] In some possible implementation manners, a SOAR script generation device based on a retrieval-enhanced generative large model according to the present application includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps in the SOAR script generation method based on a retrieval-enhanced generative large model according to various exemplary embodiments of the present application as described above.

[0052] Next, a SOAR script generation device 100 according to this embodiment of the present application will be described with reference to Figure 10 The device 100 shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application. Figure 10

[0053] Figure 10 ​​As shown, a SOAR scenario generation device 100 based on a retrieval-augmented generation large model is presented in the form of a general-purpose electronic device. The components of a SOAR scenario generation device 100 based on a retrieval-augmented generation large model may include, but are not limited to: the at least one processor 101, the at least one memory 102, and a bus 103 that connects different system components (including the memory 102 and the processor 101).

[0054] The bus 103 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local area bus using any bus structure in a variety of bus structures.

[0055] The memory 102 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1021 and / or cache memory 1022, and may further include read-only memory (ROM) 1023.

[0056] The memory 102 may also include a program / utilities 1025 having a set (at least one) of program modules 1024. Such program modules 1024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0057] A SOAR scenario generation device 100 based on a retrieval-augmented generation large model may also communicate with one or more external devices 104 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with a SOAR scenario generation device 100 based on a retrieval-augmented generation large model, and / or communicate with any device (such as a router, a modem, etc.) that enables the SOAR scenario generation device 100 based on a retrieval-augmented generation large model to communicate with one or more other electronic devices. This communication may be carried out through an input / output (I / O) interface 105. And, a SOAR scenario generation device 100 based on a retrieval-augmented generation large model may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 106. As shown in the figure, the network adapter 106 communicates with other modules for a SOAR scenario generation device 100 through the bus 103. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with a SOAR scenario generation device 100, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0058] In some possible embodiments, various aspects of a SOAR scenario generation method based on a retrieval-augmented generation large model provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in a SOAR scenario generation method based on a retrieval-augmented generation large model according to various exemplary embodiments described above in this specification.

[0059] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0060] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for generating SOAR scripts based on a retrieval-augmented generative large model, characterized in that It includes: Obtain the script text description input by the user, perform semantic parsing and intention recognition on the script text description to obtain key intention information; Match the key intention information with the script text descriptions in the first structured data stored in the script knowledge base. The first structured data includes script text descriptions generated based on processing logics of different scripts and corresponding script files; When it is determined that there is a target script text description with a similarity greater than the first similarity threshold, query and output the target script file corresponding to the target script text description.

2. The method according to claim 1, wherein It also includes: When it is determined that there is no target script text description with a similarity greater than the first similarity threshold, split the corresponding key intention information into multiple key action node information; Match the multiple key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base. The second structured data includes script action text descriptions generated based on different script actions and corresponding script action information; When it is determined that there is a target script action text description with a similarity greater than the second similarity threshold for each key action node information, query the respective target script action information corresponding to the target script action text description; Input the respective target script action information into a large model, and use the large model to generate a target script file according to a predefined template.

3. The method according to claim 2, characterized in that It also includes: When it is determined that there is no target action script text description with a similarity greater than the second similarity threshold for any key action node information, adjust the key action node information; Match the adjusted key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base again.

4. The method according to any one of claims 1, wherein The script knowledge base includes a first relational database and a first vector database. The first relational database is used to store the first structured data, and the first vector database is used to store the first structured vector data obtained by vectorizing the script text descriptions in the first structured data; Among them, matching the key intention information with the script text descriptions in the first structured data stored in the script knowledge base includes: Vectorize the key intention information to obtain the vectorized data of the key intention information; Calculate the similarity between the vectorized data of the key intention information and the first structured vector data in the script knowledge base.

5. The method according to claim 2 or 3, wherein The script action knowledge base includes a second relational database and a second vector database. The second relational database is used to store the second structured data, and the second vector database is used to store the second structured vector data obtained by vectorizing the script action text descriptions in the second structured data; Among them, matching the multiple key action node information with the script action text descriptions in the second structured data stored in the script action knowledge base includes: Vectorize the multiple key action node information respectively to obtain the vectorized data of each of the multiple key action node information; Calculate the similarity between the vectorized data of the multiple key action node information and the second structured vector data in the script action knowledge base.

6. The method according to claim 1 or 2, wherein The first structured data further includes a script ID, and associates the first structured vector data obtained by vectorizing the script ID and the corresponding script text description; The second structured data further includes a script action ID, and associates the second structured vector data obtained by vectorizing the script action ID and the corresponding script action text description.

7. The method according to claim 1 or 2, characterized in that, The first relational database in the script knowledge base is constructed in the following manner: Obtain the script text description and script file from the static preset script library, and process the script text description and script file into structured data and store it in the first relational database in the script knowledge base; Obtain the target script file generated by the large model from the dynamic custom script library and parse it to obtain the target script text description, and process the target script text description and target script file into structured data and store it in the first relational database in the script knowledge base.

8. The method according to claim 2 or 3, characterized in that, The second relational database in the script action knowledge base is constructed in the following manner: Based on the different script files stored in the first relational database in the script knowledge base, parse and extract the script files to obtain a corresponding plurality of different script actions; Based on the plurality of different script actions, generate corresponding script action text descriptions and script action information; Process the script action text description and script action information into structured data and store it in the second relational database in the script action knowledge base.

9. A SOAR script generation device based on a retrieval-augmented generation large model, characterized in that, Comprising: An intention analysis module, configured to obtain the script text description input by the user, perform semantic parsing and intention recognition on the script text description, and obtain key intention information; An intention matching module, configured to match the key intention information with the script text description in the first structured data stored in the script knowledge base, where the first structured data includes script text descriptions generated based on the processing logics of different scripts and corresponding script files; A script output module, configured to query and output the target script file corresponding to the target script text description when it is determined that there is a target script text description with a similarity greater than the first similarity threshold.

10. A SOAR script generation device based on a retrieval-augmented generation large model, characterized in that, Comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the SOAR script generation method based on the retrieval-enhanced generation large model according to any one of claims 1-8.

Citation Information

Cited By

  • Text generation method, electronic equipment and computer readable storage medium

    CN121524308A