Method and apparatus for determining Trojan horse
Through multi-layer neural network analysis of files and sessions, the problem of low accuracy of WebShell Trojan detection is solved, and more efficient detection results are achieved.
Patent Information
- Application Number
- CN202110144449.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-02-02
AI Technical Summary
In the prior art, the accuracy of WebShell Trojan detection is low, resulting in frequent false alarms and missed reports.
The first neural network model is used to analyze the file, the second neural network model is used to analyze the session, and the analysis results are combined to determine whether the file and the session are related to the Trojan horses, and the third neural network model is integrated and judged.
Improve the accuracy of WebShell Trojan detection and reduce false alarms and missed reports.
Smart Images

Figure CN114840850B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communications, and in particular, to a method and apparatus for determining a Trojan horse. Background Art
[0002] With the rapid development of the network, web applications have penetrated into all aspects of life, especially being most widely used in enterprises. Although various web applications bring a lot of conveniences, they also become the things most easily exposed to attackers. Some people with ulterior motives will find the applications that are exposed outside the enterprise and can be invaded. WebShell is a web attack method often used by attackers. The attacker can use WebShell to remotely control the web server to perform some illegal operations. In this case, WebShell is a backdoor Trojan program. In order to discover and prevent such attacks, researchers have proposed various methods for detecting WebShell Trojans. One is to detect WebShell script files, and the other is to detect the traffic characteristics during the WebShell communication process. However, both of these methods have certain limitations, resulting in many false positives and false negatives.
[0003] In view of the problem of low accuracy in detecting WebShell Trojans in the related art, there is currently no effective solution. Summary of the Invention
[0004] Embodiments of the present invention provide a method and apparatus for determining a Trojan horse, so as to at least solve the problem of low accuracy in detecting WebShell Trojans in the related art.
[0005] According to an embodiment of the present invention, there is provided a method for determining a Trojan horse, including: analyzing a first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple sets of first data, and each set of first data includes: a Trojan horse file; analyzing a first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple sets of second data, and each set of second data includes: a session based on a Trojan horse file; and determining whether the first file is a Trojan horse file and whether the first session is a session based on a Trojan horse file according to the first analysis result and the second analysis result.
[0006] Optionally, determining whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result includes: when a second identity identifier exists in a target cache, determining whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result, where the second identity identifier is the identity identifier of a second file requested in the first session.
[0007] Optionally, when a second identity identifier exists in a target cache, determining whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result includes: determining a first feature vector according to the first analysis result and the second analysis result; analyzing the first feature vector using a third neural network model to obtain a third analysis result, where the third neural network model is trained using multiple groups of third data, and each group of third data includes: a first feature vector; determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result.
[0008] Optionally, determining a first feature vector according to the first analysis result and the second analysis result includes: performing an OR operation on a second feature vector of the first file and a third feature vector of the first session to obtain the first feature vector, where the first analysis result includes the second feature vector and the second analysis result includes the third feature vector.
[0009] Optionally, determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result includes: determining the maximum value among a first probability value, a second probability value, and a third probability value, where the first analysis result includes the first probability value, the second analysis result includes the second probability value, and the third analysis result includes the third probability value; when the maximum value is greater than or equal to a first preset value, determining that the first file is a Trojan file and the first session is a session based on a Trojan file.
[0010] Optionally, the method further includes: when a second identity identifier does not exist in a target cache, if the second probability value in the second analysis result is greater than or equal to a second preset value, determining that the first session is a session based on a Trojan file, where the second identity identifier is the identity identifier of a second file requested in the first session.
[0011] Optionally, the method further includes: when a first probability value in the first analysis result is greater than or equal to a third preset value, determining that the first session is a session based on a Trojan file.
[0012] Optionally, the method further includes at least one of the following: associatively storing a first identity identifier of the first file and the first analysis result in a target cache; associatively storing a second identity identifier of a second file in the first session request and the second analysis result in the target cache.
[0013] According to another embodiment of the present invention, there is provided a Trojan determination device, including: a first analysis module configured to analyze a first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple sets of first data, and each set of first data includes: a Trojan file; a second analysis module configured to analyze a first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple sets of second data, and each set of second data includes: a session based on a Trojan file; a determination module configured to determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result.
[0014] According to still another embodiment of the present invention, there is further provided a storage medium storing a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0015] According to still another embodiment of the present invention, there is further provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0016] Through the present invention, since a first analysis result is obtained by analyzing a first file using a first neural network model, and a second analysis result is obtained by analyzing a first session using a second neural network model, it is determined whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result. Therefore, the problem of low accuracy in WebShell Trojan detection can be solved, and the effect of improving the accuracy of WebShell Trojan detection can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a mobile terminal for a method of determining a Trojan horse according to an embodiment of the present invention;
[0019] Figure 2 is a flowchart for determining a Trojan horse according to an embodiment of the present invention;
[0020] Figure 3 is a schematic diagram of a CNN network structure according to an alternative embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of an RNN network structure according to an alternative embodiment of the present invention;
[0022] Figure 5 is a schematic diagram of an MLP network structure according to an alternative embodiment of the present invention;
[0023] Figure 6 is a framework diagram of a WebShell detection system according to an alternative embodiment of the present invention;
[0024] Figure 7 is a structure block diagram of a device for determining a Trojan horse according to an embodiment of the present invention. Detailed implementation manners
[0025] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0027] The method embodiment provided in Embodiment 1 of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of determining a Trojan horse according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal 10 may include one or more ( Figure 1 only one is shown in Figure 1 the processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1more or fewer components shown, or having a configuration different from that shown in Figure 1 that shown.
[0028] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining a Trojan in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal 10 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0029] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0030] In this embodiment, a method for determining a Trojan running on the above mobile terminal is provided. Figure 2 is a flowchart of the determination of a Trojan according to an embodiment of the present invention, as shown in Figure 2 shown, and the process includes the following steps:
[0031] Step S202: Analyze the first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple groups of first data, and each group of first data includes: Trojan files;
[0032] Step S204: Analyze the first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple groups of second data, and each group of second data includes: sessions based on Trojan files;
[0033] Step S206: Determine whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result.
[0034] Through the above steps, since the first file is analyzed using the first neural network model to obtain the first analysis result, and the first session is analyzed using the second neural network model to obtain the second analysis result, it is determined whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result. Therefore, the problem of low accuracy in WebShell Trojan detection can be solved, and the effect of improving the accuracy of WebShell Trojan detection can be achieved.
[0035] Optionally, the execution subject of the above steps may be a terminal or the like, but is not limited thereto.
[0036] As an optional implementation, WebShell is a command execution environment existing in the form of web page files such as asp, php, jsp, or cgi, and can also be called a web page backdoor. After a hacker invades a website, they usually mix asp or php backdoor files with normal web page files in the WEB directory of the website server, and then can use a browser to access the asp or php backdoor to obtain a command execution environment to achieve the purpose of controlling the website server. Generally, an attacker needs to go through 3 stages to achieve their attack purpose using WebShell:
[0037] 1. Upload a WebShell file to a specific web server;
[0038] 2. The WebShell client sends commands or code to the server (through http);
[0039] 3. The server obtains the data passed by the WebShell client and uses this data as parameters to execute the WebShell file, and returns the execution result to the client if necessary.
[0040] The first stage only needs to be executed once, while the second and third stages can be executed multiple times, and the WebShell file can be requested by different clients in different places to execute specified commands.
[0041] As an optional implementation, the above first neural network model may be a convolutional neural network model CNN, such as Figure 3 shown in the schematic diagram of the CNN network structure according to an optional embodiment of the present invention. The CNN network structure includes a max pooling layer Max-pool, a convolution layer Convolution, a fully connected layer Dense, and a hidden layer. The CNN network is trained using a large number of WebShell files and normal files. Its output is a probability value between 0 and 1, which can be used as the probability that a file is a WebShell file, denoted as P f . Pf The larger it is, the greater the probability that the file is a WebShell file. The last hidden layer of the network can be used as the feature vector extracted from the file, denoted as V here. f Output the first file to the CNN network to obtain the first analysis result P. f and V f The file analysis module will use the file name as the key to store the corresponding P f and V f into the cache layer.
[0042] As an alternative implementation, the above second neural network model can be a recurrent neural network RNN or its variants such as long short-term memory network LSTM and GRU network. As Figure 4 shown in the figure is the schematic diagram of the RNN network structure according to an alternative embodiment of the present invention. The RNN is trained using a large number of http sessions based on WebShell and normal http sessions. The output of the network is also a probability value ranging from 0 to 1, denoted as P here. h P h The closer P is to 1, the greater the likelihood that the http session is a WebShell session. The last hidden layer of the network can be used as the feature vector of the http session, denoted as V here. h Output the first session to the RNN network to obtain the second analysis result P h and V h .
[0043] As an alternative implementation, based on the first analysis result P f and V f output by the CNN network, and the second analysis result P h and V h output by the RNN network, it can be determined whether the first file is a WebShell Trojan file and whether the first session is a session of a WebShell Trojan.
[0044] Optionally, determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result includes: when there is a second identity identifier in the target cache, determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result, where the second identity identifier is the identity identifier of the second file requested by the first session.
[0045] As an optional implementation, the second file is the file of the http session request, and the second identity identifier can be the file name of the http session request. Check in the cache to see if the file name of the http session request exists. If the file name and its corresponding P f and V f are found, then it is necessary to analyze together with V h and P h to determine whether the first file is a WebShell trojan file and whether the first session http session is the session of the WebShell trojan file.
[0046] Optionally, when the second identity identifier exists in the target cache, determining whether the first file is a trojan file and whether the first session is a session based on the trojan file according to the first analysis result and the second analysis result includes: determining a first feature vector according to the first analysis result and the second analysis result; using a third neural network model to analyze the first feature vector to obtain a third analysis result, where the third neural network model is trained using multiple groups of third data, and each group of third data includes: a first feature vector; determining whether the first file is a trojan file and whether the first session is a session based on the trojan file according to the first analysis result, the second analysis result and the third analysis result.
[0047] As an optional implementation, the first feature vector V is obtained according to V f and V h . The third neural network model can be a multi-layer perceptron or a fully connected neural network MLP. As shown in Figure 5 , it is a schematic diagram of the MLP network structure according to an optional embodiment of the present invention. The MLP network includes an input layer Input Layer, a hidden layer Hidden Layer, and an output layer OutputLayer. The MLP is trained through the feature vector V. Its output is a probability value between 0 and 1, indicating the probability that the input is generated by a WebShell attack, denoted here as the third analysis result P i .
[0048] Optionally, determining the first feature vector according to the first analysis result and the second analysis result includes: performing an OR operation on the second feature vector of the first file and the third feature vector of the first session to obtain the first feature vector, where the first analysis result includes the second feature vector and the second analysis result includes the third feature vector.
[0049] As an optional implementation, for the second feature vector V f and the third feature vector V hThe first eigenvector V obtained by performing an OR operation.
[0050] Optionally, determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result includes: determining the maximum value among a first probability value, a second probability value, and a third probability value, where the first probability value is included in the first analysis result, the second probability value is included in the second analysis result, and the third probability value is included in the third analysis result; and determining that the first file is a Trojan file and the first session is a session based on the Trojan file when the maximum value is greater than or equal to a first preset value.
[0051] As an optional implementation, then obtain the probability P of the integrated analysis through the following formula:
[0052] P = max{P i , P f , P h}
[0053] where P i is the third probability value, P f is the first probability value, and P h is the second probability value. Calculate the maximum value among them and denote it as P.
[0054] Here, another probability threshold is set, denoted as the first preset value T, and this value can be determined according to the actual situation. For example, it can be 0.4, 0.3, etc. If P is greater than T, it indicates that the first file is a WebShell file and the first session is an http session based on WebShell.
[0055] Optionally, the method further includes: when the second identity identifier does not exist in the target cache, if the second probability value in the second analysis result is greater than or equal to a second preset value, determining that the first session is a session based on the Trojan file, where the second identity identifier is the identity identifier of the second file requested by the first session.
[0056] As an optional implementation, the second file is the file requested by the http session, and the second identity identifier can be the file name of the http session request. Check whether the file name of the http session request exists in the cache. If the file name is not found in the cache, it can indicate that the file requested by the session is not uploaded to the server through the network. At this time, only the second probability value P h can be relied on for judgment. If P h is greater than a probability threshold T H , it indicates that the corresponding first session is a session based on WebShell. T HIt can be determined according to the actual situation, for example, it can be 0.4, 0.3, etc.
[0057] Optionally, the method further includes: when the first probability value in the first analysis result is greater than or equal to a third preset value, determining that the first session is a session based on a Trojan file.
[0058] As an alternative implementation, the first probability value P f The larger it is, the greater the probability that the file is a WebShell file. A third preset value T can be set F , if P f is greater than a probability threshold T F , it indicates that the corresponding first file is a WebShell file. T F It can be determined according to the actual situation, for example, it can be 0.4, 0.3, etc.
[0059] Optionally, the method further includes at least one of the following: associatively storing the first identity identifier of the first file and the first analysis result in a target cache; associatively storing the second identity identifier of the second file in the first session request and the second analysis result in the target cache.
[0060] As an alternative implementation, the corresponding first probability value P can be used with the file name of the first file as the key f and the second feature vector V f and put into the cache layer. Use the file name of the second file in the first session request as the key and put the corresponding second probability value P h and the third feature vector V h into the cache layer.
[0061] As an alternative implementation, the present application will be described below through a specific implementation. As Figure 6 shown is the framework diagram of the WebShell detection system according to an alternative embodiment of the present invention. The detection system from bottom to top is a network traffic acquisition layer, a restoration layer, a primary analysis layer, an integrated analysis layer, and a quick analysis layer. The data also flows from bottom to top.
[0062] The network traffic acquisition layer is used to acquire the network traffic entering and leaving the server. The network traffic includes the first file and the first session. It is used by the upper restoration layer. Usually, it is served by a probe deployed at the network entrance and exit of the server.
[0063] The restoration layer is divided into two independent modules: uploading file restoration and http session restoration. The uploading file restoration is used to analyze the behavior of uploading files from the network to the server and restore the uploaded files, where the uploaded files include the first file. Then it is handed over to the WebShell file analysis module to analyze the first file. The http session restoration is used to identify the http traffic in the network and restore the http session, and then handed over to the http session analysis module to analyze the first session.
[0064] The primary analysis layer is used to extract features from the raw data (the first uploaded file and the first session) and perform a primary analysis. The primary analysis layer contains two independent modules: the WebShell file analysis module and the http session analysis module, which are used to analyze the first file and the first session obtained by the restoration layer respectively. From the process of using WebShell for attacks, it can be seen that the behavior of uploading the WebShell file precedes the http session based on this WebShell, and the WebShell file only needs to be uploaded once, and then multiple http sessions based on this WebShell can be created. For a system that only analyzes from the network side, this causes an out-of-sync problem. To solve this problem, a cache layer is needed to store the results of file analysis.
[0065] The task of the primary analysis layer is to extract features and perform a preliminary analysis. Since the deep neural network has the ability to both extract features and perform analysis (regression, classification), the two analysis modules in this layer are both built based on the deep neural network. The file analysis module uses a convolutional neural network CNN (corresponding to the first neural network model) to extract the features of the file and perform a logistic regression.
[0066] The CNN used by the file analysis module needs to be trained offline using a large number of obtained WebShell files and normal files. The first analysis result output by it includes a value from 0 to 1, which can be used as the probability that the file is a WebShell file, denoted here as the first probability value P f . P f The larger the value, the greater the probability that the file is a WebShell file. The first analysis result also includes the feature vector extracted from the first file by the last hidden layer of the network, denoted here as the second feature vector V f . The file analysis module will use the file name as the key to put the corresponding P f and V f into the cache layer. Since the http session is sequential data, the second neural network model that can be used here can be RNN (recurrent neural network) or its variants such as LSTM, GRU, etc.
[0067] The RNN needs to be trained offline using a large number of http sessions based on WebShells and normal http sessions. The second analysis result of the network includes a second probability value, which is a value between 0 and 1, denoted here as P h . P h The closer it is to 1, the greater the likelihood that the first session is a WebShell session. The second analysis result of the network also includes the feature vector of the first session output by the last hidden layer, denoted here as the third feature vector Vh. The http session analysis module will pass the file name of the second file in the first session and V h and P h to the upper-level integrated analysis module for analysis.
[0068] The integrated analysis layer will perform integrated analysis on the analysis results from the primary analysis layer. First, it uses the file name of the second file obtained from the http session analysis module to search in the cache. If not found, it can indicate that the requested file was not uploaded to the server via the network. In this case, only P h can be used for judgment. If P h is greater than a probability threshold T H , it indicates that the corresponding http session is a WebShell-based session. If the file name of the second file and its corresponding P f and V f are found in the cache, they need to be analyzed together with V h and P h . They will be concatenated to form a new feature vector V:
[0069] V = V f ∨V h
[0070] where ∨ is the OR operation, and the first feature vector V is obtained by performing the OR operation on V f and V h .
[0071] An MLP (also known as a multi-layer perceptron or fully connected neural network) is used for integrated analysis. This MLP needs to be trained offline using the feature vector V obtained by the above method. Its output is a value between 0 and 1, representing the probability that the input is generated by a WebShell attack, denoted here as the third probability value P i .
[0072] Then, the integrated analysis probability P is obtained through the following formula:
[0073] P = max{P i , P f , P h}
[0074] Another probability threshold T will be set here. If P is greater than T, it indicates that the corresponding first file is a WebShell file, and the corresponding first session is an http session based on WebShell. Then the integrated analysis module will pass the corresponding file to the quick analysis module and mark the http session as a session based on WebShell.
[0075] At the integrated analysis layer, some files analyzed and determined to be WebShell will be passed to the quick analysis module, and the quick analysis module stores these data in its cache. With these data, the quick analysis module can simply detect using the files requested in the http session. If the requested file of the http session in the traffic appears in the cache of the quick analysis module, the http session can be directly marked as a WebShell session.
[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0077] In this embodiment, a device for determining a Trojan is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware, is also possible and contemplated.
[0078] Figure 7 is a structural block diagram of the device for determining a Trojan according to an embodiment of the present invention, as Figure 7As shown, the device includes: a first analysis module 72 for analyzing a first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple sets of first data, and each set of first data includes: a Trojan file; a second analysis module 74 for analyzing a first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple sets of second data, and each set of second data includes: a session based on a Trojan file; a determination module 76 for determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result.
[0079] Optionally, the above device is further configured to determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result when a second identity identifier exists in a target cache, where the second identity identifier is the identity identifier of a second file requested by the first session.
[0080] Optionally, the above device is further configured to determine a first feature vector according to the first analysis result and the second analysis result; analyze the first feature vector using a third neural network model to obtain a third analysis result, where the third neural network model is trained using multiple sets of third data, and each set of third data includes: a first feature vector; determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result.
[0081] Optionally, the above device is further configured to perform an OR operation on a second feature vector of the first file and a third feature vector of the first session to obtain the first feature vector, where the first analysis result includes the second feature vector, and the second analysis result includes the third feature vector.
[0082] Optionally, the above device is further configured to determine the maximum value among a first probability value, a second probability value, and a third probability value, where the first probability value is included in the first analysis result, the second probability value is included in the second analysis result, and the third probability value is included in the third analysis result; determine that the first file is a Trojan file and the first session is a session based on the Trojan file when the maximum value is greater than or equal to a first preset value.
[0083] Optionally, the above device is further configured to, when the second identity identifier does not exist in the target cache, determine that the first session is a session based on a Trojan file if the second probability value in the second analysis result is greater than or equal to a second preset value, where the second identity identifier is the identity identifier of a second file requested by the first session.
[0084] Optionally, the above device is further configured to determine that the first session is a session based on a Trojan file when the first probability value in the first analysis result is greater than or equal to a third preset value.
[0085] Optionally, the above device is further configured to associatively store the first identity identifier of the first file and the first analysis result in the target cache; associatively store the second identity identifier of the second file requested by the first session and the second analysis result in the target cache.
[0086] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.
[0087] A storage medium is further provided, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0088] Optionally, in this embodiment, the above storage medium can be configured to store a computer program for executing the following steps:
[0089] S1. Analyze a first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple groups of first data, and each group of first data includes: Trojan files;
[0090] S2. Analyze a first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple groups of second data, and each group of second data includes: sessions based on Trojan files;
[0091] S3. Determine whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result.
[0092] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0093] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0094] Optionally, the above electronic device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the above processor, and the input / output devices are connected to the above processor.
[0095] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0096] S1, Analyze the first file using a first neural network model to obtain a first analysis result. Among them, the first neural network model is trained using multiple groups of first data, and each group of first data includes: Trojan files;
[0097] S2, Analyze the first session using a second neural network model to obtain a second analysis result. Among them, the second neural network model is trained using multiple groups of second data, and each group of second data includes: sessions based on Trojan files;
[0098] S3, Determine whether the first file is a Trojan file and whether the first session is a session based on a Trojan file according to the first analysis result and the second analysis result.
[0099] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated here.
[0100] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0101] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining a Trojan horse, characterized in that, Including: Analyze the first file using a first neural network model to obtain a first analysis result, where the first neural network model is trained using multiple sets of first data, and each set of first data includes: Trojan files; Analyze the first session using a second neural network model to obtain a second analysis result, where the second neural network model is trained using multiple sets of second data, and each set of second data includes: sessions based on Trojan files; When a second identity identifier exists in the target cache, determine a first feature vector according to the first analysis result and the second analysis result, where the second identity identifier is the identity identifier of the second file requested by the first session, the second file is the file requested by an HTTP session, and the second identity identifier is the file name of the HTTP session request; Analyze the first feature vector using a third neural network model to obtain a third analysis result, where the third neural network model is trained using multiple sets of third data, and each set of third data includes: first feature vectors; Determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result; When the second identity identifier does not exist in the target cache, if the second probability value in the second analysis result is greater than or equal to a second preset value, determine that the first session is a session based on a Trojan file, where the second identity identifier is the identity identifier of the second file requested by the first session.
2. The method according to claim 1, characterized in that, Determining the first feature vector according to the first analysis result and the second analysis result includes: Performing an OR operation on the second feature vector of the first file and the third feature vector of the first session to obtain the first feature vector, where the first analysis result includes the second feature vector and the second analysis result includes the third feature vector.
3. The method according to claim 2, wherein Determining whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result includes: Determine the maximum value among a first probability value, a second probability value, and a third probability value, where the first analysis result includes the first probability value, the second analysis result includes the second probability value, and the third analysis result includes the third probability value; When the maximum value is greater than or equal to a first preset value, determine that the first file is a Trojan file and the first session is a session based on the Trojan file.
4. The method according to claim 1, wherein The method further includes: When the first probability value in the first analysis result is greater than or equal to a third preset value, determine that the first session is a session based on the Trojan file.
5. The method according to claim 1, wherein The method further includes at least one of the following: Associatively store the first identity identifier of the first file and the first analysis result in the target cache; Associatively store the second identity identifier of the second file requested by the first session and the second analysis result in the target cache.
6. A determining device for a Trojan horse, characterized in that, Including: A first analysis module, configured to analyze a first file using a first neural network model to obtain a first analysis result, wherein the first neural network model is trained using multiple groups of first data, and each group of first data includes: a Trojan file; A second analysis module, configured to analyze a first session using a second neural network model to obtain a second analysis result, wherein the second neural network model is trained using multiple groups of second data, and each group of second data includes: a session based on a Trojan file; A determination module, configured to determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result and the second analysis result; The determination module is further configured to determine a first feature vector according to the first analysis result and the second analysis result when a second identity identifier exists in a target cache, wherein the second identity identifier is an identity identifier of a second file requested by the first session, the second file is a file requested by an HTTP session, and the second identity identifier is the file name of the HTTP session request; Analyze the first feature vector using a third neural network model to obtain a third analysis result, wherein the third neural network model is trained using multiple groups of third data, and each group of third data includes: a first feature vector; Determine whether the first file is a Trojan file and whether the first session is a session based on the Trojan file according to the first analysis result, the second analysis result, and the third analysis result; The determination module is further configured to determine that the first session is a session based on a Trojan file if a second probability value in the second analysis result is greater than or equal to a second preset value when the second identity identifier does not exist in the target cache, wherein the second identity identifier is an identity identifier of a second file requested by the first session.
7. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the program, when executed by a terminal device or a computer, performs the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Webshell detection method and device based on artificial intelligence
CN111614599A