A language model-based reverse shell detection method, device, equipment, medium and product for cloud security protection and business risk identification
Through a language model-based method, prompt words are generated and their natural language processing capabilities are used to detect external input and transmission paths in code files, the applicability and scalability of detecting script-like rebound shells across programming languages in the existing technology is solved, and more efficient rebound shell detection is achieved.
Patent Information
- Application Number
- CN202411688778.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing security protection products are difficult to effectively identify code files in script-like rebound shells, especially because code files can be written by different types of programming languages, resulting in poor applicability and scalability of traditional detection methods.
The language model-based method is adopted to obtain external input, target execution function and transmission path information in the code file by generating prompt words, and use the natural language processing capabilities of the language model for detection to realize rebound shell detection across programming languages.
It improves the accuracy and applicability of rebound shell detection, can effectively identify rebound shell behavior in different programming language environments, and reduces the detection blind spots of coding bypass and complex methods.
Smart Images

Figure CN119538266B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a language model-based reverse shell detection method, apparatus, electronic device, computer-readable storage medium, and computer program product for cloud security protection and business risk identification. Background Art
[0002] With the continuous development of computer technology, security protection products for security detection have emerged as the times require. Security protection products can detect physical computing devices such as computers and hosts, or virtual computing devices such as containers, so as to ensure the running security. In practical applications, security protection products can include cloud workload protection platform (CWPP), host-based intrusion detection system (HIDS), endpoint detection and response (EDR), container security platform (CSP), etc.
[0003] Security protection products can perform various detections on computing devices. For example, security protection products can detect reverse shells. Among them, a reverse shell can be understood as a technique that enables a target computing device (also called the controlled end) to actively connect to the attacker's computing device (also called the control end), so that the control end can remotely control the controlled end.
[0004] Specifically, a listening port is set on the control end, and a shell command is executed on the controlled end to create a reverse connection to the listening port on the control end on the controlled end, so as to realize the connection between the controlled end and the control end. In a script-based reverse shell, the shell command executed on the controlled end can be in the form of a code file, that is, the controlled end executes the code file to establish a reverse connection from the controlled end to the control end, so as to realize the remote control of the controlled end by the control end.
[0005] Since the code files in script-based reverse shells are usually embedded in normally running system processes, it is difficult to identify them through traditional signature detection or traffic analysis methods. At the same time, since code files can be written in different types of programming languages, the detection of script-based reverse shells is particularly complex, and there is an urgent need in the industry for a flexible and highly applicable security detection method. Summary of the Invention
[0006] The present application provides a language model-based reverse shell detection method for cloud security protection and business risk identification. This method can accurately determine whether a code file has a reverse shell risk, and moreover, it can achieve cross-programming language reverse shell detection, improving scalability and applicability. The present application also provides an apparatus, an electronic device, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0007] In a first aspect, the present application provides a language model-based reverse shell detection method for cloud security protection and business risk identification, and the method includes:
[0008] In response to detecting a target file event of a first computing device, obtaining a first code file related to the target file event; wherein, the first code file is used to execute in the first computing device;
[0009] Generating a first prompt;
[0010] Wherein, the first prompt includes: the first code file; information indicating the detection of the following target data: external input in the first code file, a target execution function in the first code file, and a transmission path of the external input in the first code file; and information indicating a detection method for detecting the first code file based on the target data;
[0011] Sending the first prompt to a first language model and receiving a detection result of the first code file returned by the first language model.
[0012] In a second aspect, the present application provides a language model-based reverse shell detection apparatus for cloud security protection and business risk identification, and the apparatus includes:
[0013] An obtaining module, configured to obtain a first code file related to the target file event in response to detecting a target file event of a first computing device; wherein, the first code file is used to execute in the first computing device;
[0014] A generating module, configured to generate a first prompt; wherein, the first prompt includes: the first code file; information indicating the detection of the following target data: external input in the first code file, a target execution function in the first code file, and a transmission path of the external input in the first code file; and information indicating a detection method for detecting the first code file based on the target data;
[0015] A communication module, configured to send the first prompt to a first language model and receive a detection result of the first code file returned by the first language model.
[0016] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory, so that the electronic device executes the language model-based reverse shell detection method for cloud security protection and business risk identification as described in the first aspect or any implementation manner of the first aspect.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions direct an electronic device to execute the language model-based reverse shell detection method for cloud security protection and business risk identification as described in the first aspect or any implementation manner of the first aspect.
[0018] In a fifth aspect, the present application provides a computer program product containing instructions, which, when running on an electronic device, causes the electronic device to execute the language model-based reverse shell detection method for cloud security protection and business risk identification as described in the first aspect or any implementation manner of the first aspect.
[0019] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
[0020] From the above technical solutions, it can be seen that the present application has the following advantages:
[0021] The present application provides a language model-based reverse shell detection method for cloud security protection and business risk identification. In response to detecting a target file event of a first computing device, the method obtains a first code file related to the target file event, where the first code file is used to be executed in the first computing device. Then, a first prompt word is generated, where the first prompt word includes the first code file. Information on the following target data is indicated to be detected: external input in the first code file, a target execution function in the first code file, and a transmission path of the external input in the first code file. And information indicating a detection method for detecting the first code file based on the target data is sent to a first language model, and a detection result of the first code file returned by the first language model is received.
[0022] In this method, for a code file executed in a first computing device, by leveraging the natural language processing ability of a language model, external input, data flow paths, and command execution behaviors in the code file are detected. On the one hand, in combination with the specific information of the code file, it is accurately determined whether the code file is related to reverse shell behavior. On the other hand, cross-programming language reverse shell detection is achieved by means of a language model, enhancing scalability and applicability. Brief Description of the Drawings
[0023] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below.
[0024] Figure 1 It is a schematic flowchart of a language model-based reverse shell detection method for cloud security protection and business risk identification provided by an embodiment of the present application;
[0025] Figure 2 It is a schematic flowchart of a language model-based reverse shell detection method for cloud security protection and business risk identification provided by an embodiment of the present application;
[0026] Figure 3 It is a schematic structural diagram of a language model-based reverse shell detection device for cloud security protection and business risk identification provided by an embodiment of the present application;
[0027] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0028] The terms "first" and "second" in the embodiments of the present application are only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0029] First, some technical terms and application scenarios involved in the embodiments of the present application will be introduced.
[0030] With the continuous development of computer technology, security protection products for performing security detection and then ensuring the security of physical computing devices such as computers and hosts or virtual computing devices such as containers have emerged as the times require. Security protection products can perform multi-faceted security detection for various operating scenarios. For example, the security protection product can be a cloud workload protection platform (CWPP), and the CWPP can detect host security and network security. Another example is that the security protection product can be a host-based intrusion detection system (HIDS), and the HIDS can perform security detection on the behavior and status of a computer system. Another example is that the security protection product can be an endpoint detection and response (EDR), and the EDR can perform security detection on the system-level behavior of endpoints. Another example is that the security protection product can be a container security platform (CSP), and the CSP can provide multi-faceted security detection for containers.
[0031] In some examples, the security protection product can detect reverse shells. Among them, a reverse shell can be understood as an attack method that, through technical means, enables the target computing device (also called the controlled end) to actively connect to the attacker's computing device (also called the control end), so that the control end can remotely control the controlled end.
[0032] Specifically, a listening port is set on the control end. By executing a shell command on the controlled end, a reverse connection to the listening port on the control end is created on the controlled end to achieve the connection between the controlled end and the control end. In a script-type reverse shell, the shell command executed on the controlled end can be in the form of a code file, that is, the controlled end executes the code file to establish a reverse connection from the controlled end to the control end, realizing the remote control of the controlled end by the control end.
[0033] Since the code files in script-based reverse shells are usually embedded in normally running system processes, it is difficult to identify them through traditional signature detection or traffic analysis methods. In related technologies, taint analysis technology is usually used to detect reverse shells. Among them, the core idea of taint analysis technology is: for the code files of the controlled end, mark the external inputs (also called tainted data) that may come from untrusted sources, detect the propagation paths of the tainted data in the code files, and when the tainted data passes through critical functions (also called sink functions), it indicates that the code file may be used to establish a reverse connection and trigger an alarm related to the reverse shell.
[0034] Usually, when using taint analysis technology to detect reverse shells, the code files are traversed through an abstract syntax tree (AST) to generate a list of characters (tokens) corresponding to the code files. During the detection process, the externally configured inputs in the token list (such as user inputs, network requests, etc.) are marked as tainted data, and the propagation paths of the tainted data are recorded by detecting the propagation rules (such as assignment functions, etc.) configured in the token list. When the tainted data passes through the critical functions (such as command execution functions, file operation functions, etc.) configured in the token list, an alarm is triggered to indicate that the code file may be related to reverse shell behavior.
[0035] However, the above method of detecting reverse shells based on the abstract syntax tree has the following defects: First, the above method depends on pre-configured tainted data, propagation rules, and critical functions, and the real-time performance of security detection is poor. Second, the anti-bypass ability of the above method is weak, and it is difficult to effectively detect the propagation paths for encoding functions, decoding functions, encryption functions, etc. in the code files. Moreover, since code files can be written in different types of programming languages, the above method needs to configure different tainted data, propagation rules, and critical functions for different types of programming languages respectively, which greatly increases the complexity of security detection and has poor applicability and scalability.
[0036] In view of this, the present application provides a language model-based reverse shell detection method for cloud security protection and business risk identification. In response to detecting a target file event of a first computing device, the method obtains a first code file related to the target file event, where the first code file is used to be executed in the first computing device. Then, a first prompt is generated, where the first prompt includes the first code file; information for instructing to detect the following target data: external inputs in the first code file, target execution functions in the first code file, and the transmission paths of the external inputs in the first code file; and information for instructing the detection method for detecting the first code file based on the target data. The first prompt is sent to a first language model, and a detection result of the first code file returned by the first language model is received.
[0037] In this method, for the code file executed in the first computing device, by virtue of the natural language processing ability of the language model, the external inputs, data flow paths, and command execution behaviors in the code file are detected. On the one hand, combined with the specific information of the code file, it is accurately determined whether the code file is related to the reverse shell behavior. On the other hand, the reverse shell detection across programming languages is realized by means of the language model, improving the scalability and applicability.
[0038] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, the following will be described with reference to the accompanying drawings. Refer to Figure 1 The flowchart of a language model-based reverse shell detection method for cloud security protection and business risk identification provided by an embodiment of the present application as shown. The method specifically includes:
[0039] S101: In response to detecting a target file event of a first computing device, obtain a first code file related to the target file event.
[0040] In the embodiments of the present application, security detection is performed on script-based reverse shells. Considering that in script-based reverse shells, the controlled end executes shell commands in the form of running code files to establish a reverse connection from the controlled end to the control end. Therefore, in the embodiments of the present application, by detecting system file events, code files related to the first computing device are obtained, and then security detection is performed on the code files.
[0041] Among them, the first computing device can be understood as the device for performing security detection. In other words, in the embodiments of the present application, security detection is performed on whether there is a reverse shell behavior in the first computing device. In some embodiments, the first computing device may be an entity computing device such as a computer or a host. In other embodiments, the first computing device may also be a virtual computing device such as a virtual machine or a container. The embodiments of the present application do not limit this.
[0042] The target file event can be understood as a file event generated in the first computing device. For example, the target file event can be a file creation event in the first computing device. For another example, the target file event can also be a file modification event for an existing file in the first computing device.
[0043] The first code file can be understood as a file associated with the target file event. For example, when the target file event is a file creation event, the first code file can be a code file created in the first computing device. For another example, when the target file event is a file modification event, the first code file can be a code file to be modified in the first computing device.
[0044] In the embodiment of the present application, the first code file is used to be executed in the first computing device, that is to say, the first code file can be a code file running in the first computing device.
[0045] The first code file may include executable instructions for running in the first computing device. For example, the first code file can be a script file. The embodiment of the present application does not limit the programming language of the first code file. For example, the programming language of the first code file can be Java, Python, PHP, etc.
[0046] It should be noted that the embodiment of the present application does not limit the method for detecting the target file event of the first computing device. In some embodiments, based on the fanotify technology, the file directory of the first computing device can be detected to implement the detection of the target file event of the first computing device. In other embodiments, the target file event of the first computing device can also be detected through an external file detection tool. In other embodiments, when the security detection method provided by the embodiment of the present application is executed by a security protection product, the target file event of the first computing device can also be detected through the file detection function provided by the security protection product.
[0047] In some embodiments, considering that the number of target file events detected for the first computing device may be large, in order to improve the security detection efficiency, the first code file can also be screened and filtered. Specifically, as Figure 2 shown, at least one of the following operations is performed on the first code file: matching the first code file with a file set and determining the file size of the first code file.
[0048] Among them, the file set includes multiple files. For example, the file set may include multiple files that have passed security detection in the first computing device (such as multiple business files). In other words, the file set can be understood as a whitelist file set. In this way, by matching the first code file with the file set, it is determined whether the first code file belongs to the known secure files in the first computing device. If the first code file hits the file set, that is, the matching is successful, the first code file is discarded and no subsequent security detection processing is required.
[0049] In some embodiments, the file hash of the first code file can be determined, and the file hash of the first code file is compared with the file hashes of multiple files in the file set. By virtue of the uniqueness of the file hash, the matching of the first code file with the file set is achieved.
[0050] On the other hand, considering that the file size of the code file used to establish a reverse connection is usually small, and it consumes too much computing resources to perform subsequent security detection processing on code files with a large file size, the file size of the first code file can also be determined, and the first code file is screened based on the file size of the first code file. Specifically, if the file size of the first code file is not less than the set threshold, the first code file is discarded and no subsequent security detection processing is required. In this way, computing resources are saved and more targeted security detection is achieved.
[0051] The embodiments of the present application do not limit the set threshold for the file size. In practical applications, it can be set according to the file upload rate and the file processing rate to avoid situations such as processing timeouts.
[0052] S102: Generate a first prompt.
[0053] In the embodiments of the present application, security detection is performed with the help of a first language model. Among them, the first language model has natural language processing capabilities, can understand the meaning of natural language, and process different types of natural language tasks. For example, the first language model can be a deep learning model trained using text data. In other words, by using the natural language processing capabilities of the first language model, the idea of detecting a reverse shell for the first computing device is understood, and security detection for the first code file is achieved.
[0054] Specifically, when implemented, the first language model performs security detection on the first code file based on the method of prompt learning. Among them, a prompt can be used to guide the language model to perform specific outputs in generative tasks (such as text generation tasks, question-and-answer tasks, and dialogue tasks). By configuring the prompt, the language model is helped to understand the background and requirements of the task, and without the need to retrain the language model, the language model can process different types of natural language processing tasks, increasing the scalability and flexibility of the language model.
[0055] In an embodiment of the present application, the first prompt word may include: a first code file; information indicating the detection of the following target data: external input in the first code file, a target execution function in the first code file, and the transmission path of the external input in the first code file; and information indicating the detection method for detecting the first code file based on the target data.
[0056] That is to say, the first prompt word includes three parts of information: the first part is the first code file, the second part is the information indicating the first language model to detect the target data in the first code file, and the third part is the information indicating the first language model to generate a detection result based on the target data.
[0057] Among them, the target data specifically includes external input, a target execution function, and a propagation path. That is to say, the first prompt word instructs the first language model to perform taint analysis on the first code file, identify the external input existing in the first code file, the propagation path of the external input in the first code file, and the target execution function related to command execution.
[0058] The information indicating the detection of the external input in the first code file in the first prompt word may include: information indicating the detection of the input function in the first code file, and information indicating the detection of the external resource call in the first code file. In this way, for the cases of explicit input functions and indirect calls to external resources, detection can be achieved.
[0059] The information indicating the detection of the target execution function in the first code file in the first prompt word may include: information indicating the detection of command execution in the first code file, and information indicating the detection of file operations in the first code file. In other words, the target execution function may be a sink function in taint analysis.
[0060] The information indicating the detection of the transmission path of the external input in the first code file in the first prompt word may include: information indicating the detection of the propagation path of the external input in the first code file after at least one of the following operations: assignment operation, encoding operation, decoding operation, encryption operation, and conversion operation.
[0061] That is to say, by configuring different propagation methods of the external input in the first prompt word, the first language model can, based on the prompting ability of the first prompt word, detect the propagation path of the external input in the first code file through different propagation methods, avoid the defects that are easily missed in traditional taint analysis for operations such as encoding and decoding, encryption, etc., and effectively cope with complex bypass techniques such as encoding bypass, multi-file combination, and traffic encryption to reduce the detection blind area.
[0062] Further, since the first prompt includes information indicating the detection method for the first language model to detect the first code file based on the target data, the first language model can, based on the prompting ability of the first prompt, determine whether the first code file is used to establish a reverse connection and whether it is related to the behavior of a reverse shell, thereby realizing the security detection of the first code file.
[0063] In some possible implementation manners, considering the actual process of detecting a reverse shell using taint analysis technology, the detection method for detecting the first code file based on the target data can be: in response to meeting the following conditions, determining that the first code file fails the detection: there is an external input and a target execution function in the first code file, and based on the transmission path of the external input in the first code file, determining that the input parameter of the target execution function comes from the external input.
[0064] In other words, the first language model can, based on the prompting ability of the first prompt, determine whether an external input and a target execution function are detected in the first code file. When an external input and a target execution function are detected in the first code file, and the propagation path of the external input indicates that the input parameter of the target execution function comes from the external input, that is, the propagation path of the external input includes the target execution function, it indicates that when the target execution function of the first code file is executed, the data related to the external input will be executed together, and there is a possibility of establishing a reverse connection. Therefore, there is a possibility that the first code file has a reverse shell and fails the detection.
[0065] In this way, by configuring the specific detection method for the first code file in the first prompt, based on the prompting ability of the first prompt, the first language model is informed how to perform security detection on the first code file, so that the first language model can automatically identify whether there is content related to a reverse shell in the first code file and determine whether the first code file passes the security detection.
[0066] As described above, in some embodiments, screening and filtering are performed on the first code file. In this case, in response to the existence of at least one of the following situations, a first prompt is generated: the first code file fails to match successfully with the file set and the file size of the first code file is less than a set threshold.
[0067] In some embodiments, in order to further improve the accuracy of security detection, auxiliary information can also be retrieved based on the retrieval-augmented generation (RAG) technology, so that the first language model can refer to the auxiliary information to generate a more accurate detection result. Specifically, continue as Figure 2As shown, obtain auxiliary information representing natural language content from the security detection knowledge base, and generate a first prompt word including the auxiliary information. Among them, the auxiliary information includes: information describing the external input in the code file, information describing the determined propagation path, and information describing the input target execution function in the code file.
[0068] Among them, the security detection knowledge base stores multiple pieces of auxiliary information related to detecting reverse shells using taint analysis technology. By configuring the auxiliary information in the first prompt word, the first language model can more comprehensively learn how to identify external inputs in the first code file, how to determine the propagation path of external inputs in the first code file, and how to identify the target execution function in the first code file, making the detection of the first code file more accurate.
[0069] Moreover, since the auxiliary information in the security detection knowledge base is described in natural language and not limited to a specific programming language, thus, with the natural language processing ability of the first language model, learn the analysis ideas of identifying external inputs, determining propagation paths, and identifying target execution functions, and in the first code files of different programming languages, the detection of target data can be realized, achieving cross-language and cross-environment security detection.
[0070] Illustrated with an example, the first prompt word can be as follows:
[0071] "You are an expert in detecting reverse shells in code files, specializing in analyzing potential reverse shell vulnerabilities in code files, especially those involving external inputs, command execution, and data transfer, and generating detection results for code files through external inputs, command execution, and data transfer.
[0072] The code file is as follows:
[0073] {{First code file}}
[0074] Skill 1: Detect external input
[0075] Detect whether there is a function in the code file that receives external input
[0076] Possible techniques can be found in the security detection knowledge base
[0077] Skill 2: Detect command execution
[0078] Detect whether there is a function for command execution in the code
[0079] Possible techniques can be found in the security detection knowledge base
[0080] Skill 3: Detect data transfer
[0081] Check whether the parameters of the execution command are passed in the middle and finally come from external input
[0082] Possible techniques that may be used can be found in the security detection knowledge base
[0083] The security detection knowledge base is as follows:
[0084] a. Variable assignment, related to Skill 3: Assigning external input to a new variable, causing taint propagation.
[0085] b. Parameter passing, related to Skill 3: When a function is called, external input is propagated as a parameter to the function interior.
[0086] c. Encoding / decoding, related to Skill 3: Data is converted through encoding (such as Base64, URL encoding) or decoding, and taint propagates accordingly.
[0087] d. Type / format conversion, related to Skill 3: Conversion of data types or formats, such as converting a string to an integer, and taint continues to propagate.
[0088] e. Collection operations, related to Skill 3: Tainted data is stored in data structures such as arrays and lists and propagates through access and operations.
[0089] f. Input / output binding, related to Skill 3: External input interacts with the program through standard input / output, file reading and writing, etc. and propagates.
[0090] g. File and environment variable operations, related to Skill 3: External input contaminates files or environment variables, affecting subsequent operations of the program.
[0091] h. Pipes / streams, related to Skill 3: Tainted data is passed through pipes or streams to other processes or program modules.
[0092] i. Serialization / deserialization, related to Skill 3: Taint still propagates during the serialization and deserialization of data.
[0093] j. Regular expression replacement, related to Skill 3: When replacing with regular expressions, data is modified but taint continues to propagate.
[0094] k. Closures and callback functions, related to Skill 3: The state of external variables carried in closures or callback functions propagates taint.”
[0095] Among them, "The code file is as follows: {{The first code file}}" is the first code file included in the first prompt, "Skill 1: Detect external input", "Skill 2: Detect command execution", and "Skill 3: Detect data transfer" are the information indicating the target data to be detected in the first prompt, "Generate the detection result of the code file through external input, command execution, and data transfer" is the information indicating the detection method for detecting the first code file based on the target data in the first prompt, and the information in the "Security Detection Knowledge Base" is the auxiliary information in the first prompt.
[0096] In the embodiment of the present application, by configuring the first prompt, the first language model can perform accurate and comprehensive taint analysis on the first code file in combination with the first prompt.
[0097] S103: Send the first prompt to the first language model and receive the detection result of the first code file returned by the first language model.
[0098] By calling the first language model and inputting the first prompt into the first language model, the first language model performs taint analysis on the first code file under the prompting ability of the first prompt, identifies the external input and target execution functions in the first code file, detects the propagation path of the external input in the first code file, and generates the detection result of the first code file based on the detected above-mentioned target data.
[0099] Specifically, the detection result of the first code file may include at least one of the following: information indicating whether the first code file passes the detection, information describing the analysis process of generating the detection result, and code fragments in the first code file related to the detection result.
[0100] That is to say, the detection result of the first code file returned by the first language model may carry multiple pieces of information. Among them, the information indicating whether the first code file passes the detection can be understood as a judgment conclusion. For example, the first code file passes the detection or the first code file fails the detection. The first code file passing the detection means that the first code file is not a code file for establishing a reverse connection, and there is no rebound shell behavior in the first code file.
[0101] The information describing the analysis process of generating the detection result can be understood as a judgment idea, and the code fragments in the first code file related to the detection result can be understood as basis code fragments. In other words, the first language model can present the idea and basis of the judgment result together, indicating the specific analysis process of the first language model.
[0102] The detection result of the first code file can be returned in the form of a structure. For example, the detection result of the first code file can be as follows:
[0103]
[0104] In this way, by leveraging the natural language processing capabilities of the first language model, the overall semantics and code structure of the first code file are analyzed to determine whether the first code file is used to establish a reverse connection and whether there is a risk of a reverse shell. Without the need to adapt to different programming languages, cross-language and cross-environment security detection is achieved.
[0105] Furthermore, continue as Figure 2 shown. In response to the detection result indicating that the first code file fails the detection, an alarm event is generated according to the detection result, and the information associated with the alarm event is presented.
[0106] Among them, the information associated with the alarm event may include the detection result of the first code file, the information of the first computing device, the handling information for dealing with the reverse shell, etc. In this way, when the first code file fails the detection and there may be a reverse shell behavior in the first code file, the alarm event is triggered in a timely manner so that the user (such as a security operation personnel) can perform alarm handling in a timely manner to ensure the running security of the first computing device.
[0107] Furthermore, since the natural language model is used for security detection in the embodiments of the present application, considering that the detection result of the first code file generated by the first language model may deviate from the actual needs of the user, it is also possible to support the user (such as a security operation personnel) to judge the alarm event. Continue as Figure 2 shown. Receive a determination operation for the alarm event. In response to the determination operation indicating that the information associated with the alarm event is incorrect, perform at least one of the following operations: update the information associated with the alarm event, update the security detection knowledge base, and add the first code file to the file set.
[0108] That is to say, if the alarm event generated according to the detection result is correct, no additional operation is required. If the alarm event generated according to the detection result is incorrect, for example, there are errors in the judgment conclusion, attribution, etc. of the first code file, the user can correct it based on the information associated with the currently presented alarm event to form updated information associated with the alarm event.
[0109] In addition, if the alarm event generated according to the detection result is incorrect, for example, there is an error in the judgment idea of the first code file, the user can correct and revise multiple pieces of auxiliary information in the security detection knowledge base related to detecting reverse shells using taint analysis technology to form an updated security detection knowledge base.
[0110] In addition, if the alarm event generated according to the detection result is incorrect, for example, the first code file belongs to a business file, the first code file can be added to the file set to update the whitelist file set.
[0111] In this method, for the code file executed in the first computing device, by virtue of the natural language processing ability of the language model, the external input, data flow path, and command execution behavior in the code file are detected. On the one hand, in combination with the specific information of the code file, it is accurately determined whether the code file is related to the reverse shell behavior. On the other hand, the reverse shell detection across programming languages is realized by means of the language model, improving the scalability and applicability.
[0112] As described above in combination with Figure 1 and Figure 2 A detailed introduction to the language model-based reverse shell detection method for cloud security protection and business risk identification provided by the embodiments of the present application has been given. Next, the devices and equipment provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0113] See Figure 3 The structural schematic diagram of the language model-based reverse shell detection device for cloud security protection and business risk identification shown in
[0114] An acquisition module 301, configured to acquire a first code file related to the target file event in response to detecting a target file event of the first computing device; wherein, the first code file is used to be executed in the first computing device;
[0115] A generation module 302, configured to generate a first prompt word; wherein, the first prompt word includes: the first code file; information indicating to detect the following target data: external input in the first code file, target execution functions in the first code file, and the transmission path of the external input in the first code file; and, information indicating the detection method for detecting the first code file based on the target data;
[0116] A communication module 303, configured to send the first prompt word to the first language model and receive the detection result of the first code file returned by the first language model.
[0117] In some possible implementation manners, the detection method for detecting the first code file based on the target data includes:
[0118] In response to satisfying the following conditions, it is determined that the first code file fails the detection: there is external input and a target execution function in the first code file, and based on the transmission path of the external input in the first code file, it is determined that the input parameter of the target execution function comes from the external input.
[0119] In some possible implementations, the information indicating the transmission path of the external input in the first code file includes:
[0120] Information indicating the propagation path of the external input in the first code file after at least one of the following operations: assignment operation, encoding operation, decoding operation, encryption operation, and conversion operation.
[0121] In some possible implementations, the apparatus 30 further includes a retrieval module, and the retrieval module is configured to:
[0122] Obtain auxiliary information characterizing natural language content from a security detection knowledge base; wherein, the auxiliary information includes: information describing the external input in the code file, information describing the determined propagation path, and information describing the input target execution function in the code file;
[0123] The first prompt word further includes: the auxiliary information.
[0124] In some possible implementations, the apparatus 30 further includes a preprocessing module, and the preprocessing module is configured to:
[0125] Perform at least one of the following operations on the first code file: match the first code file with a file set and determine the file size of the first code file; wherein, the file set includes multiple files;
[0126] The generating module 302 is specifically configured to:
[0127] Generate a first prompt word in response to the existence of at least one of the following situations: the first code file fails to match successfully with the file set and the file size of the first code file is less than a set threshold.
[0128] In some possible implementations, the detection result of the first code file includes at least one of the following: information indicating whether the first code file passes the detection, information describing the analysis process of generating the detection result, and code fragments in the first code file related to the detection result.
[0129] In some possible implementations, the apparatus 30 further includes an alarm module, and the alarm module is configured to:
[0130] In response to the detection result indicating that the first code file fails to pass the detection, generate an alarm event according to the detection result;
[0131] Present information associated with the alarm event.
[0132] In some possible implementations, the apparatus 30 further includes an update module, and the update module is configured to:
[0133] Receive a determination operation for the alarm event;
[0134] In response to the determination operation indicating that there is an error in the information associated with the alarm event, perform at least one of the following operations: update the information associated with the alarm event, update the security detection knowledge base, and add the first code file to the file set.
[0135] The language model-based reverse shell detection device 30 for cloud security protection and business risk identification according to an embodiment of the present application can correspondingly execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the language model-based reverse shell detection device 30 for cloud security protection and business risk identification are respectively for implementing Figure 1 The corresponding processes of the respective methods in the illustrated embodiments are not described herein again for the sake of brevity.
[0136] An embodiment of the present application also provides an electronic device. This electronic device is specifically used to implement the functions of the language model-based reverse shell detection device 30 for cloud security protection and business risk identification as shown in Figure 3 the illustrated embodiments.
[0137] Figure 4 A structural schematic diagram of an electronic device 400 is provided. As shown in Figure 4 the figure, the electronic device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other through the bus 401.
[0138] The bus 401 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0139] The processor 402 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0140] The communication interface 403 is used for external communication. For example, the communication interface 403 can be used to communicate with a terminal.
[0141] The memory 404 may include a volatile memory, such as a random access memory (RAM). The memory 404 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0142] The memory 404 stores executable code, and the processor 402 executes the executable code to perform the foregoing language model-based reverse shell detection method for cloud security protection and business risk identification.
[0143] Specifically, in the case of implementing Figure 3 the illustrated embodiment, and Figure 3 when each module or unit of the language model-based reverse shell detection device 30 for cloud security protection and business risk identification described in the embodiment is implemented by software, the software or program code required to execute the functions of each module / unit in Figure 3 can be partially or entirely stored in the memory 404. The processor 402 executes the program code corresponding to each unit stored in the memory 404 to perform the foregoing language model-based reverse shell detection method for cloud security protection and business risk identification.
[0144] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the foregoing language model-based reverse shell detection method for cloud security protection and business risk identification applied to the language model-based reverse shell detection device 30.
[0145] The embodiment of the present application further provides a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the process or function described in the embodiment of the present application is generated in whole or in part.
[0146] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer or data center to another website, computer or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0147] When the computer program product is executed by a computer, the computer executes any of the aforementioned rebound shell detection methods based on language models for cloud security protection and business risk identification. The computer program product may be a software installation package, and when any of the aforementioned rebound shell detection methods based on language models for cloud security protection and business risk identification is needed, the computer program product may be downloaded and executed on a computer.
[0148] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to each embodiment of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0149] The units involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.
[0150] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
[0151] In the context of the embodiments of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0152] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method section.
[0153] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0154] It should also be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0155] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination thereof. The software module may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0156] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A language model-based reverse shell detection method for cloud security protection and business risk identification, characterized in that, The method includes: In response to detecting a target file event of a first computing device, obtaining a first code file related to the target file event; wherein, the first code file is a script file for execution in the first computing device; Generating a first prompt; Wherein, the first prompt includes: the first code file; information indicating detection of the following target data: external input in the first code file, a target execution function in the first code file, and a transmission path of the external input in the first code file; and, information indicating a detection method for detecting the first code file based on the target data; the detection method for detecting the first code file based on the target data includes: in response to meeting the following conditions, determining that the first code file fails the detection: there is external input and a target execution function in the first code file, and based on the transmission path of the external input in the first code file, determining that the input parameter of the target execution function originates from the external input; Sending the first prompt to a first language model, and receiving a detection result of the first code file returned by the first language model, where the detection result represents whether the first code file has a reverse shell behavior.
2. The method according to claim 1, characterized in that The information indicating detection of the transmission path of the external input in the first code file includes: Information indicating detection of a propagation path of the external input in the first code file after at least one of the following operations: assignment operation, encoding operation, decoding operation, encryption operation, and conversion operation.
3. The method according to claim 1, wherein The method further includes: Obtaining auxiliary information representing natural language content from a security detection knowledge base; wherein, the auxiliary information includes: information describing external input in the code file, information describing determination of the propagation path, and information describing input to the target execution function in the code file; The first prompt further includes: the auxiliary information.
4. The method according to claim 1, characterized in that The method further includes: Performing at least one of the following operations on the first code file: matching the first code file with a file set and determining the file size of the first code file; wherein, the file set includes multiple files; The generating of the first prompt includes: In response to at least one of the following situations existing, generating a first prompt: the first code file fails to match successfully with the file set and the file size of the first code file is less than a set threshold.
5. The method according to claim 1, wherein The detection result of the first code file includes at least one of the following: information representing whether the first code file passes the detection, information describing the analysis process of generating the detection result, and a code snippet in the first code file related to the detection result.
6. The method according to any one of claims 1 to 5, characterized in that The method further includes: In response to the detection result representing that the first code file fails the detection, generating an alarm event according to the detection result; Presenting information associated with the alarm event.
7. The method according to claim 6, characterized in that, The method further includes: Receiving a determination operation for the alarm event; In response to the determination operation indicating that there is an error in the information associated with the alarm event, perform at least one of the following operations: update the information associated with the alarm event, update the security detection knowledge base, and add the first code file to the file set.
8. A language model-based reverse shell detection device for cloud security protection and business risk identification, characterized in that, The device includes: An acquisition module, configured to acquire a first code file related to the target file event in response to detecting a target file event of a first computing device; wherein, the first code file is a script file for execution in the first computing device. A generation module, configured to generate a first prompt; wherein, the first prompt includes: the first code file; information indicating detection of the following target data: external inputs in the first code file, target execution functions in the first code file, and the transmission path of the external inputs in the first code file; and, information indicating the detection method for detecting the first code file based on the target data; the detection method for detecting the first code file based on the target data includes: determining that the first code file fails the detection in response to the following conditions being met: there are external inputs and target execution functions in the first code file, and based on the transmission path of the external inputs in the first code file, determining that the input parameters of the target execution function are derived from the external inputs. A communication module, configured to send the first prompt to a first language model and receive the detection result of the first code file returned by the first language model, where the detection result indicates whether there is a reverse shell behavior in the first code file.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory; The processor is configured to execute instructions stored in the memory, causing the electronic device to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Includes instructions that direct the electronic device to execute the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, The computer program product includes computer-readable instructions for implementing the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Code risk detection method and device, electronic equipment and computer storage medium
CN118673497A