Detection rule establishment method and apparatus, and file detection method and apparatus
By generating detection rules based on decision tree models in the cloud computing system, combined with dynamic detection methods, the problems of low accuracy and high labor cost in the existing technology are solved, efficient and low-cost malicious file detection are achieved, and the security of the cloud computing system is improved.
Patent Information
- Application Number
- PCT/IB2024/063098
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-24
AI Technical Summary
The prior art has problems with high labor costs, low detection accuracy and vulnerable to attackers when detecting malicious files in cloud computing systems, especially when using static feature detection and intelligent algorithm model detection.
By obtaining the attribute information of multiple preset type entities of the file in the running environment, a detection rule based on the decision tree model is generated, and rules for detecting malicious files are generated using binary tree structures and logical expressions, and files are run in a sandbox environment in combination with dynamic detection methods, recording behaviors and generating detection rules.
It improves the accuracy of detecting malicious files, reduces labor costs and update time, enhances the interpretability of detection rules, reduces false detection and missed detection, and improves the security of cloud computing systems.
Smart Images

Figure IB2024063098_24072025_PF_FP_ABST
Abstract
Description
[0001] This disclosure claims priority to Chinese patent application number 202410069347.5, filed with the Patent Office of China on January 17, 2024, entitled "Method for Establishing Detection Rules, Method for Detecting Files, and Apparatus," the entire contents of which are incorporated herein by reference. Technical Field: This disclosure relates to the field of cloud computing technology, and more particularly to a method and apparatus for establishing detection rules, as well as a method and apparatus for detecting files. Background: File security is increasingly important in cloud computing systems. The presence of malicious files in a cloud computing system poses a significant threat to the integrity of the cloud computing system. Therefore, it is necessary to perform security checks on files in the cloud computing system. If a malicious file is detected in the cloud computing system, the malicious file can be deleted from the cloud computing system to improve the security of the cloud computing system. SUMMARY: This disclosure provides a method and apparatus for establishing detection rules, as well as a method and apparatus for detecting files. In a first aspect, a method for establishing a detection rule is shown, including: obtaining operation information of multiple files running in a running environment, the operation information of the files including multiple preset types of entities involved in the process of the files running in the running environment, the multiple preset types of entities including the file name and the process name for processing the file; determining first operation information and second operation information in the operation information of the multiple files, the file name in the first operation information is the file name of the malicious file, and the file name in the second operation information is not the file name of the malicious file; obtaining attribute information of multiple preset types of entities in the first operation information, obtaining attribute information of multiple preset types of entities in the second operation information; generating a detection rule for detecting malicious files based on the attribute information of the multiple preset types of entities in the first operation information and the attribute information of the multiple preset types of entities in the second operation information. In a second aspect, a method for detecting a file is shown, comprising: obtaining running information of a file to be detected, the running information of the file to be detected including multiple entities of preset types involved in the process of the file to be detected running in a running environment, the multiple entities of preset types including at least the file name of the file to be detected and the name of a process for processing the file to be detected; obtaining attribute information of the multiple entities of preset types in the running information of the file to be detected; and detecting whether the file to be detected is a malicious file based on the attribute information and a pre-generated detection rule for detecting malicious files.In a third aspect, a device for establishing detection rules is shown, including: a first acquisition module, used to obtain operation information of multiple files running in a running environment, the operation information of the files including multiple preset types of entities involved in the process of the files running in the running environment, the multiple preset types of entities including the file name and the process name for processing the file; a determination module, used to determine the first operation information and the second operation information in the operation information of the multiple files, the file name in the first operation information is the file name of the malicious file, and the file name in the second operation information is not the file name of the malicious file; a second acquisition module, used to obtain attribute information of multiple preset types of entities in the first operation information, and obtain attribute information of multiple preset types of entities in the second operation information; a generation module, used to generate a detection rule for detecting malicious files based on the attribute information of multiple preset types of entities in the first operation information and the attribute information of multiple preset types of entities in the second operation information. In a fourth aspect, a device for detecting files is provided, comprising: a third acquisition module for acquiring execution information of a file to be detected, the execution information of the file to be detected including: multiple entities of preset types involved in the execution of the file to be detected in an execution environment, the multiple entities of preset types including at least: the file name of the file to be detected and the name of the process used to process the file to be detected; a fourth acquisition module for acquiring attribute information of the multiple entities of preset types in the execution information of the file to be detected; and a detection module for detecting whether the file to be detected is malicious based on the attribute information and pre-generated detection rules for detecting malicious files. In a fifth aspect, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in any of the aforementioned aspects. In a sixth aspect, a non-transitory computer-readable storage medium is provided, wherein when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the method described in any of the aforementioned aspects. In a seventh aspect, a computer program product is provided, wherein when the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is enabled to execute the method described in any of the aforementioned aspects.Compared to the prior art, the present disclosure offers the following advantages: If the accuracy of detecting malicious behavior in a file under inspection based on detection rules needs to be further improved, the detection rules are automatically generated logically based on information related to the actual operation of the file in the execution environment. Therefore, the detection rules are often white-box and interpretable. This means that the reason for detecting malicious behavior in the file under inspection based on the detection rules can generally be determined. For example, the reason why the detection rules failed to detect malicious behavior in a malicious file can generally be determined. Therefore, updating the detection rules to address false positives does not require regenerating detection rules. Instead, new detection rules can be added to existing detection rules, or some existing detection rules can be modified. On the one hand, since the detection rules can be added or modified directly based on existing detection rules, the time required to regenerate detection rules can be saved and update efficiency can be improved. On the other hand, manual data collection for generating detection rules is eliminated, thereby reducing labor costs. On the other hand, since detection rules don't need to be regenerated, adding new detection rules to existing ones will not change or impact the existing ones, nor will it reduce the detection accuracy of the existing ones. Alternatively, modifying some of the existing detection rules can improve the detection accuracy of the modified detection rules corresponding to those rules. Furthermore, since the existing detection rules (unmodified detection rules) remain unchanged, they will not be affected and their detection accuracy will not be reduced. After using detection rules for a long time to detect malicious behavior in files, resistance to the detection rules may develop, potentially leading to a reduction in the detection accuracy of the detection rules. For example, an attacker can deliver a large number of malicious files to test the detection rules. If it is found that the malicious behavior of certain types of malicious files is not detected as malicious behavior by the detection rules, the attacker can determine that the malicious behavior of certain types of malicious files can bypass the detection of the detection rules, and then specifically construct a large number of malicious files of these certain types, and make these certain types of malicious files perform malicious behavior to attack the cloud computing system. The attacker can also fight against the detection rules and avoid being detected by the detection rules for their malicious behavior, resulting in a decrease in the detection accuracy of the detection rules, thereby affecting the security of the cloud computing system.Furthermore, because the accuracy of detecting malicious behavior in files under inspection based on detection rules decreases, it is often necessary to improve the accuracy of detecting malicious behavior in files under inspection based on detection rules, for example, by updating the detection rules. However, as previously described, the present disclosure can improve update efficiency and reduce labor costs without reducing the accuracy of existing detection rules. BRIEF DESCRIPTION OF THE DRAWINGS FIG1 is a flowchart of the steps of a method for establishing detection rules according to the present disclosure; FIG2 is a schematic diagram of a binary tree structure according to the present disclosure; FIG3 is a flowchart of the steps of a method for obtaining operation information according to the present disclosure; FIG4 is a flowchart of the steps of a method for determining operation information according to the present disclosure; FIG5 is a schematic diagram of a knowledge graph according to the present disclosure; FIG6 is a schematic diagram of a knowledge graph according to the present disclosure; FIG7 is a flowchart of the steps of a method for detecting files according to the present disclosure; FIG8 is a block diagram of a device for establishing detection rules according to the present disclosure; FIG9 is a block diagram of a device for detecting files according to the present disclosure; and FIG10 is a block diagram of a device according to the present disclosure. To make the above-mentioned objectives, features, and advantages of the present disclosure more readily apparent, the present disclosure is further described in detail below with reference to the accompanying drawings and specific embodiments. When detecting whether a file to be detected is a malicious file, one approach can be to detect whether the file to be detected is malicious based on its static features. For example, the file to be detected can be disassembled to obtain certain inherent features (e.g., inherent attributes) of the file to be detected. These features are then used as static features of the file to be detected. The static features of the file to be detected are then tested to determine whether they possess inherent features of a malicious file. If the static features of the file to be detected possess inherent features of a malicious file, the file to be detected can be determined to be malicious. Alternatively, if the static features of the file to be detected do not possess inherent features of a malicious file, the file to be detected can be determined to be a safe file rather than a malicious file. A large number of technical personnel can be pre-assigned to collect statistics on the inherent features of malicious files on the market and specifically identify which inherent features are commonly found in malicious files, thereby developing a detection model tailored to these inherent features. In this way, when testing whether the static features of a file to be detected have inherent characteristics of a malicious file, the inherent features of the file to be detected can be matched with the detection model. If the inherent features of the file to be detected match the detection model, it can be determined that the static features of the file to be detected have inherent characteristics of a malicious file. Alternatively, if the inherent features of the file to be detected do not match the detection model, it can be determined that the static features of the file to be detected do not have inherent characteristics of a malicious file. However, the inventors have discovered that the above approach has the following drawbacks:
[0002] 1. The process of "counting the inherent characteristics of malicious files on the market, specifically identifying which inherent characteristics are common in malicious files, and developing detection models based on the inherent characteristics that malicious files often have" often requires the participation of a large number of technical personnel and has high labor costs.
[0003] 2. To prevent malicious files from being detected, malicious file developers often hide the inherent characteristics of malicious files through encryption, packing, and / or obfuscation (for example, by converting the inherent characteristics into garbled characters or symbols). This can result in an inability to count all the inherent characteristics of malicious files on the market, leading to omissions in the counted inherent characteristics of malicious files. This, in turn, can lead to omissions in the targeted identification of inherent characteristics commonly found in malicious files. Consequently, the developed detection model may miss at least some of the inherent characteristics of malicious files on the market. Consequently, the developed detection model may not include at least some of the inherent characteristics of malicious files on the market, resulting in an incomplete detection model with low generalizability. Furthermore, if the file to be detected is malicious and its inherent features are hidden through encryption, packing, and / or obfuscation, it may be impossible to extract at least some of the inherent features of the file to be detected. If this is impossible, the aforementioned method will likely fail to match the detection model, resulting in the file not being detected as malicious, i.e., a missed detection. Furthermore, although the aforementioned method cannot extract some of the inherent features of the encrypted, packed, and / or obfuscated file, the inventors have discovered that the file requires the inherent features of the file during execution. Therefore, the file will first execute the self-decryption code within the file to automatically decrypt, unpack, and / or deobfuscate the file, thereby obtaining the inherent features of the file, which can then be used during execution. In light of this, another method has been proposed for detecting whether the file to be detected is malicious based on dynamic detection during actual file execution. For example, a file to be tested is delivered to a sandbox environment, run within the sandbox environment, and the file's behavior within the sandbox environment (e.g., actions performed by the file within the sandbox environment) is recorded. The file's behavior within the sandbox environment can then be tested for malicious behavior. If malicious behavior is detected, the file can be determined to be malicious. Alternatively, if no malicious behavior is detected, the file can be determined to be safe rather than malicious. In one example, dynamic detection methods may include those based on intelligent algorithm models, for example, detecting whether a file to be tested is malicious based on an intelligent algorithm model.For example, a file can be delivered to a sandbox environment and run within it. Based on the sandbox environment's run report, the file's process call sequence (including the file's behavior, etc.) in the sandbox environment can be extracted. These process call sequences can then be vectorized to obtain vectorized features, and an intelligent algorithm model can be trained based on these vectorized features. Subsequently, to detect whether the behavior of the file to be tested in the sandbox environment contains malicious behavior, the file to be tested can be delivered to the sandbox environment and run within it. The process call sequence of the file to be tested in the sandbox environment can be recorded and then input into the trained intelligent algorithm model. The trained intelligent algorithm model then processes the process call sequence of the file to be tested in the sandbox environment to obtain a detection result indicating whether the behavior of the file to be tested in the sandbox environment contains malicious behavior. However, the inventors have discovered that this alternative method for detecting whether the behavior of the file to be tested contains malicious behavior has low accuracy. Sometimes, the low accuracy of detecting malicious behavior in files under inspection may be attributed to the intelligent algorithm model itself. Therefore, it may be necessary to improve the accuracy of detecting malicious behavior in files under inspection based on the intelligent algorithm model, for example, by updating the intelligent algorithm model. However, intelligent algorithm models are often black boxes and therefore lack interpretability. In other words, it is often impossible to determine why the intelligent algorithm model detects malicious behavior in files under inspection. For example, it is often impossible to determine why the intelligent algorithm model fails to detect malicious behavior in a malicious file. Therefore, updating the intelligent algorithm model to address false positives or missed detections often requires manually collecting appropriate training data based on the missed detections and retraining the intelligent algorithm model using the collected training data. However, the process of manually collecting appropriate training data and retraining the intelligent algorithm model is often time-consuming, resulting in low update efficiency. Furthermore, manually collecting appropriate training data requires the participation of a large number of technical personnel, resulting in high labor costs. On the other hand, retraining an intelligent algorithm model using newly collected training data may also introduce new false detections and missed detections, which may in turn reduce the detection accuracy of the retrained intelligent algorithm model. Furthermore, after using an intelligent algorithm model for a long time to detect malicious behavior in files to be detected, it may sometimes develop resistance to the intelligent algorithm model, resulting in a decrease in the intelligent algorithm model's detection accuracy.For example, an attacker can submit a large number of malicious files to test the intelligent algorithm model. If the malicious behavior of certain types of malicious files is not detected as malicious by the intelligent algorithm model, the attacker can determine that the malicious behavior of these types of malicious files can bypass the detection of the intelligent algorithm model. The attacker can then construct a large number of malicious files of these types and cause them to perform malicious behaviors to attack the cloud computing system. This can counter the intelligent algorithm model and prevent it from detecting their malicious behavior, resulting in a decrease in the detection accuracy of the intelligent algorithm model, which in turn affects the security of the cloud computing system. Furthermore, the low accuracy of detecting malicious behavior in the behavior of the detected file based on the aforementioned alternative method may sometimes be attributed to the intelligent algorithm model itself. Therefore, it may be necessary to improve the accuracy of detecting malicious behavior in the behavior of the detected file based on the intelligent algorithm model, for example, by updating the intelligent algorithm model. However, as mentioned above, updating is inefficient, labor-intensive, and may also result in a decrease in detection accuracy. Therefore, to address the above-mentioned issues, the solution of the present disclosure is proposed. Before introducing the solutions of this disclosure, we will first explain some of the technical terms that may be involved. Knowledge graph: A knowledge base called a semantic network, i.e., a knowledge base with a directed graph structure. A knowledge graph consists of at least two entities (also called nodes), and different entities may have relationships. Binary file: A computer file format that stores data in binary form. MD5, or Message-Digest Algorithm 5, is a widely used cryptographic hash function that generates a 128-bit (16-byte) hash value to ensure the integrity and consistency of information transmission. A file's MD5 is its signature or identifier, and different files have different MD5s. Sandbox: A sandbox simulates a host environment, providing isolation and security, allowing malware to run within it and recording its operational behavior. The decision tree model is a decision analysis method that uses a tree structure to determine the probability of the expected net present value being greater than or equal to zero, based on the known probabilities of various scenarios. This allows for project risk assessment and feasibility assessment. A decision tree is a tree-like structure whose decision branches resemble the branches of a tree. A decision tree consists of a root node, internal nodes, and leaf nodes.Each decision tree has a single root node. Each internal node represents a test on an attribute, each branch represents a test output, and each leaf node represents a category. Decision tree generation generally begins with the root node, selects the corresponding attribute, then chooses a split point for the node's corresponding attribute, and then splits the node based on the split point. The decision tree generates multiple child nodes by selecting features and corresponding split points. If the value in a node belongs only to a certain category (or has a small variance), no further child node splitting is performed. FIG1 illustrates a method for establishing detection rules in the present disclosure. This method is applied to electronic devices in cloud computing scenarios. Electronic devices in cloud computing scenarios can include physical devices or virtual devices. Physical devices include servers or terminals, while virtual devices can include virtual machines hosted on physical devices. Virtual machines can include Elastic Cloud Server (ECS), etc. The method may include: in step S101, obtaining execution information of multiple files running in an execution environment; the file execution information includes: multiple preset entities of various types involved in the execution of the file in the execution environment; the multiple preset entities of various types including at least: the file name and the name of the process used to process the file. Furthermore, in one embodiment, in addition to the process name of the process used to process the file and the file name, the multiple preset entities of various types may also include: the file path, the file MD5 (the file MD5 can be used as a file identifier; different files may have the same file name, but different MD5s are used to uniquely identify the file), an identifier of the action executed by the file, the external port to which the process is connected, the external IP (Internet Protocol) address to which the process is connected, and the event name of an event created by the process (the event may include a login event, a payment event, or a call event). The file path may include the file storage path in the electronic device. The identifier of the action executed by the file may include "reading file information," "accessing the registry," and "calling a system function," etc., and examples are not given here. The behavior identifier is used to indicate the type of behavior, etc., and is used to uniquely identify the type of behavior. The behavior executed by the file is a dynamic action. The behavior identifier can be used to quantify the behavior executed by the file.The identifier of the file's executed behavior can be obtained based on an ATTCK (Adversarial Tactics, Techniques, and Common Knowledge) matrix compiled in advance by technical personnel. The ATTCK matrix includes a mapping relationship between behavior description information and behavior identifiers. Thus, the description information of the file's executed behavior can be obtained (e.g., obtained from behavior logs), and then the identifier of the file's executed behavior can be indexed into the ATTCK matrix based on the description information of the file's executed behavior. In the present disclosure, an execution environment is loaded on an electronic device, and files execute within the execution environment. Therefore, specific execution information of the file within the execution environment is often stored in the relevant log information of the electronic device. For example, while a file executes within the execution environment, the electronic device collects specific execution information of the file within the execution environment in real time and stores it in the relevant log information of the electronic device. Thus, execution information of multiple files executing within the execution environment can be obtained based on the relevant log information of the electronic device. Files in the present disclosure may include binary files, etc. The details of this step can be found in the embodiments described below and will not be described in detail here. In step S102, first and second operation information are determined from the operation information of multiple files. The file name in the first operation information is the file name of a malicious file, while the file name in the second operation information is not the file name of a malicious file. The purpose of this step is to divide the operation information of multiple files into two categories: first operation information and second operation information. The file name in the first operation information is an entity of a preset type, while the file name in the second operation information is an entity of a preset type. In one embodiment, for the operation information of any file, at least the file name is included among the multiple entities of the preset types in the operation information of the file. The file may or may not be a malicious file. If the file is a malicious file, the file name is the file name of a malicious file, and the first operation information can be obtained based on the operation information of the file, for example, using the operation information of the file as the first operation information. Alternatively, if the file is not a malicious file, the file name is not the file name of a malicious file, and the second operation information can be obtained based on the operation information of the file, for example, using the operation information of the file as the second operation information. Malicious files are files that may bring security risks to the cloud computing system. For example, malicious files include files that may execute malicious behaviors. The malicious behaviors executed by malicious files may bring security risks to the cloud computing system.Alternatively, in another embodiment, first operation information may be determined from the operation information of multiple files, and then the operation information from the operation information of the multiple files other than the first operation information may be used as the second operation information. For details of this step, please refer to the embodiments described below and will not be described in detail here. In step S103, attribute information of multiple preset entities of the first operation information is obtained, and attribute information of multiple preset entities of the second operation information is obtained. In one embodiment, for any piece of first operation information, the preset entities of the first operation information may include: the file name, the process name of the process processing the file, the file path, the file's MD5, an identifier of the action executed by the file, the external port to which the process connects, the external IP address to which the process connects, and the event name of an event created by the process. For the entity "file name", the attribute information of the "file name" may be characters in the "file name". For example, if a file is named "lover.exe", where "lover" is the file name, the file name "lover" is an entity of the preset type, and the attribute information of the file name "lover" may be "lover". For the entity "Process Name," if the cmd line command line information of the process corresponding to "Process Name" contains the command 'cat,' the label of "Process Name" is set to command_cat. If the cmd line command line information of the process corresponding to "Process Name" contains the keyword "password," the label of "Process Name" is set to keyword_password. Each label of "Process Name" is used as attribute information of "Process Name." For the entity "File Path," the file path can be labeled. The purpose of labeling file paths is to map various paths to fixed labels. For example, multiple different preset paths can be set in advance, and the keywords in different preset paths may not all be the same. For "File Path," it is possible to determine which preset path's keywords are included in the characters in the file path. Based on the determined keywords, the label of the file path can be obtained and used as the attribute information of the file path. For example, keywords for multiple preset paths may include: [' / dev / mem', ' / dev / tcp', ' / dev / udp', ' / dev / kmem', ' / dev / null',
[0004] Vinet / tcp If the file path is Vdev / mem / temp', and the file path Vdev / mem / temp' contains the keyword " / dev / mem", then the file path Vdev / mem / temp's label can be 'path / dev / mem', and the label format can be path_{}, where {} contains the keyword. For the entity "file MD5", the attribute information of "file MD5" can be the characters in "MD5". For example, the MD5 of a file is "XXXX XXXX", where "XXXX XXXX" is MD5. The file MD5 "XXXX XXXX" is an entity of a preset type, then the attribute information of the file MD5 "XXXX XXXX'" can be
[0005] "XXXX > . XXXX". For the entity "Identifier of the behavior executed by the file", the attribute information of "Identifier of the behavior executed by the file" can be the characters in "Identifier". For example, if a file performs the behavior of accessing the registry, the identifier of the behavior executed by the file is "Accessing the registry", where "Accessing the registry" is the identifier. The identifier of the behavior executed by the file "Accessing the registry" is an entity of a preset type, and the attribute information of the identifier of the behavior executed by the file "Accessing the registry" can be "Accessing the registry". For the entity "External port", "External port" is an actual port. The port segment where the external port is located can be determined from multiple pre-set port segments. Different port segments have different labels. The label of the determined port segment can be obtained and then determined as the attribute information of the external port. For example, ports can be divided into three categories. The first category: Well-known ports: (0-1023) are tightly bound to certain services. Communication on these ports generally indicates the protocol of a certain service. (For example, port 80 is used for HTTP (Hypertext Transfer Protocol) communications, port 21 is assigned to FTP (File Transfer Protocol) services, port 25 is assigned to SMTP (Simple Mail Transfer Protocol) services, port 135 is assigned to RPC (Remote Procedure Call) services, and so on.) The second category: Registered ports (1024-49151): These are loosely associated with a number of services. Many services are bound to these ports, and these ports are also used for many other purposes. (The system handles dynamic ports starting at 1024.) The third category: Dynamic / private ports (49152-65535): These ports should generally not be assigned to services. In practice, machines typically assign dynamic ports starting at 1024. However, there are exceptions: some RPC ports start at 32768. For the entity "external IP (Internet Protocol) address", the "external IP address" is an actual IP address. The address segment where the external IP address is located can be determined from multiple different address segments set in advance. Different address segments have different labels. The label of the determined address segment can be obtained, and then the label of the determined address segment can be determined as the attribute information of the external IP address.For example, an address segment may have a label of Link_local (link-local address), an address segment may have a label of Loopback (loopback address), an address segment may have a label of Multicast (multicast), an address segment may have a label of Reserved (reserved), an address segment may have a label of Unspecified (unspecified), an address segment may have a label of Private (private), and an address segment may have a label of Public (public). For the entity "Event Name," the attribute information of the "Event Name" may be characters in the "Event Name." For example, if an event is a login event and its event name is "LOGIN_EVENT," and the event name "LOGIN_EVENT" is a preset type entity, the attribute information of the event name "LOGIN_EVENT" may be "LOGIN_EVENT." The same process is repeated for each of the other first and second operation information, which will not be described in detail here. This process obtains attribute information of multiple entities of the preset types in each of the first and second operation information. In step S104, detection rules for detecting malicious files are generated based on the attribute information of the multiple entities of the preset types in the first and second operation information. In one embodiment, this step can be implemented through the following process, including:
[0006] 1041. Train a decision tree model based on the attribute information of multiple entities of preset types in the first operation information and the attribute information of multiple entities of preset types in the second operation information to obtain a binary tree structure. For any piece of first operation information, since the file names in the multiple entities of preset types in the first operation information are malicious file names, negative sample training data can be obtained based on the attribute information of each entity of preset types in the first operation information. For example, the attribute information of each entity of preset types in the first operation information is organized into a sequence in a specific order and used as negative sample training data. The specific order can be the order between entities of each preset type, or can be pre-set, for example, specified by a technician. The same process is repeated for each other piece of first operation information. Thus, for each piece of first operation information, multiple pieces of negative sample training data can be obtained. For any piece of second operation information, since the file names in the multiple entities of preset types in the second operation information are not malicious file names, positive sample training data can be obtained based on the attribute information of each entity of preset types in the second operation information. For example, the attribute information of each preset entity type in the second operation information is organized into a sequence in a specific order and used as positive training data. The same process is repeated for each other piece of second operation information. Thus, for each piece of second operation information, several pieces of positive training data can be obtained. A decision tree model can then be trained based on the positive training data and the negative training data to obtain a binary tree structure. Each node in the binary tree structure represents a piece of attribute information, and different nodes represent different pieces of attribute information. Any piece of attribute information may be located in the negative training data, or in the positive training data, or in both the negative and positive training data. In one example, the binary tree structure obtained by training can be shown in FIG2 . FIG2 includes nodes A through K, each of which represents a piece of attribute information, and each of which represents different pieces of attribute information. In the binary tree structure, leaf nodes are accompanied by results indicating whether they match the negative training data and / or whether they match the positive training data. In Figure 2, node A is connected to node B, node A is connected to node C, node B is connected to node D, node B is connected to node E, node D is connected to node G, node D is connected to node H, node H is connected to node K, node E is connected to node I, node E is connected to node J, and node C is connected to node F. oNode A is the root node, and nodes G, K, I, J, and F are leaf nodes. Leaf node G only hits all negative training data, leaf node K only hits all negative training data, leaf node I only hits all negative training data, leaf node J hits both negative and positive training data, and node F only hits all positive training data.
[0007] 1042. Generate detection rules for detecting malicious files based on the binary tree structure. For example, in one embodiment, the binary tree structure can be traversed using a pre-order traversal method to generate detection rules for detecting malicious files. For example, the binary tree structure can be traversed using a pre-order traversal method starting from a leaf node to generate detection rules for detecting malicious files. The detection rules contain logical expressions. The detection rules contain attribute information. The attribute information is at least part of the attribute information of multiple preset types of entities in the first operation information and the attribute information of multiple preset types of entities in the second operation information. For example, the attribute information is part of or all of the attribute information of multiple preset types of entities in the first operation information and the attribute information of multiple preset types of entities in the second operation information. At least part of the attribute information is connected by logical symbols, such as "logical AND," "logical OR," and "logical NOT." In one embodiment, a "logical AND" or "logical NOT" logical symbol can be set for each node in the binary tree structure to obtain a binary tree structure with logical symbols. For example, for any node in a binary tree structure, if the node is the left branch of its parent node, the node representing the left branch does not contain the attribute information represented by the node. Therefore, the logical symbol "logical NOT" can be added to the left of the node. Alternatively, if the node is the right branch of its parent node, the node representing the right branch contains the attribute information represented by the node. Therefore, the logical symbol "logical AND" can be added to the left of the node. The same applies to each other node in the binary tree structure, resulting in a binary tree structure with logical symbols. All nodes other than the root node in the binary tree structure with logical symbols have the logical symbol "logical AND" or "logical NOT." Alternatively, in another embodiment, before setting the logical symbol of "logical AND" or "logical NOT" for each node in the binary tree structure, the logical symbol of "left any match" may be set for each node in the binary tree structure. For example, the logical symbol of "left any match" may be added to the left side of each node to improve the generalization of the detection rules generated subsequently. Then, the logical symbol of "logical AND" or "logical NOT" may be set for each node in the binary tree structure that has the logical symbol of "left any match" set. For example, the logical symbol of "logical AND" or "logical NOT" may be added to the left side of each node in the binary tree structure that has the logical symbol of "left any match" set, thereby obtaining a binary tree structure with logical symbols.Then, from all leaf nodes in the binary tree structure with logical symbols, only leaf nodes that match all negative training data can be selected. Starting from any selected leaf node, the binary tree structure is traversed using a pre-order traversal method (recursively traversing the left node first, then the right node). During the traversal, logical expressions are continuously accumulated and established until the root node and all selected leaf nodes are traversed. Finally, a final detection rule is generated. When establishing a logical expression between a node and its parent node, a logical expression can be established between the node and its parent node using the logical symbol "logical AND." Furthermore, when establishing a logical expression between a node and its sibling node (the node and its sibling node share the same parent node), a logical expression can be established between the node and its sibling node using the logical symbol "logical OR." After traversing from the leaf nodes to the root node and all selected leaf nodes, a logical expression containing the logical symbols "logical AND," "logical NOT," and "logical OR" is generated and used as a detection rule. For example, according to the aforementioned method, the logical symbol "logical AND" or "logical NOT" is set for each node in the binary tree structure, resulting in a binary tree structure with logical symbols. All nodes other than the root node in the binary tree structure with logical symbols have the logical symbol "logical AND" or "logical NOT." The root node in the binary tree structure with logical symbols is located at level 1, and level N is the level where the node farthest from the root node is located. Leaf nodes may be located at level N, level N-1, level N-2, and so on, where N is greater than or equal to 2 or 3. For example, in a binary tree structure with logical symbols, one of the leaf nodes is N-1-A (N-1 is the layer number of the leaf node, A is the number of the leaf node in layer N-1, and different nodes in the same layer have different numbers). Leaf node N-1-A is the leaf node that only matches all negative training data. Thus, traversal can be performed from the leaf nodes toward the root node, for example, to the parent node N-2-A of leaf node N-1 (N-2 is the layer number of node N-2-A, and A is the number of node N-2-A in layer N-2). In the present disclosure, the relationship between two nodes in a parent-child relationship can be a "logical AND" relationship. Thus, a logical expression 1 can be established between leaf node N-1-A and its parent node N-2-A, connected by the logical symbol "logical AND": [(leaf node N-1-A) && (node N-2-A)]"&&" represents the logical AND symbol. If, in addition to the branch from leaf node N-1-A, there are other branches among the child nodes of node N-2-A, and none of the leaf nodes in those branches contain leaf nodes that only match all negative training data, or if there are no other branches, traversal can continue from node N-2-A toward the root node. Alternatively, if, in addition to the branch from leaf node N-1-A, there are other branches among the child nodes of node N-2-A, and none of the leaf nodes in those branches contain leaf nodes that only match all negative training data, a logical expression can be established between node N-2-A and the nodes in the other branches. Traversal can then continue from node N-2-A toward the root node. For example, the nodes in other branches include node N-1-B (N-1 is the layer number of leaf node N-1-B, and B is the number of leaf node N-1-B in layer N-1) and leaf node NA (N is the layer number of leaf node NA, and A is the number of leaf node NA in layer N). Leaf node NA is a leaf node that only matches all negative training data. Node N-1-B is the parent node of leaf node NA, and node N-1-B is the child node of node N-2-A. Thus, logical expression 2 can be established, connecting nodes N-2-A, N-1-B, and NA using the logical AND symbol: [(node N-2-A) && (node N-1-B) && (node NA)] . Leaf node N-1-A and node N-1-B are siblings. The relationship between two nodes in a sibling relationship can be a logical OR relationship. After logical expressions are established for all branches of the leaf nodes of node N-2-A that have only negative sample training data that hit all, logical expressions can be established for all branches of the leaf nodes of node N-2-A that have only negative sample training data that hit all, and the branches can be connected using the logical symbol "logical OR" to obtain logical expression N-2-A. For example, from the perspective of node N-2-A, a logical expression N-2-A can be established between logical expression 1 and logical expression 2, connected using the logical symbol "logical OR" (the merged logical expression can be named after the highest-level node in the merged logical expression): [ (logical expression 1). (Logical expression 2)] oIn an example, the logical expression N-2-A can be: [ (leaf node N-1-A) && (node [ (node N-2-A) && (node N-1-B) && (node NA) ]]. After the "||" represents the logical symbol "logic", the traversal can continue from node N-2-A to the root node. For example, traversing to node N-3-A (N-3 is the layer number of node N-3-A, and A is a node of node N-3-A in the N-3 layer). In the present disclosure, the relationship between two nodes in a parent-child relationship can be a "logical AND" relationship. Thus, a logical expression 3 can be established between logical expression N-2-A and node N-3-A, the parent node of node N-2-A, using the logical AND symbol: [(logical expression N-2-A) && (node N-3-A)] . If, in the branches of node N-3-A's child nodes, there are other branches besides the branch of leaf node N-2-A, and none of the leaf nodes in these other branches contain leaf nodes that only match all negative training data, or if there are no other branches, traversal can continue from node N-3-A toward the root node. Alternatively, if, in the branches of node N-3-A's child nodes, there are other branches besides the branch of leaf node N-2-A, and none of the leaf nodes in these other branches contain leaf nodes that only match all negative training data, then a logical expression can be established between node N-3-A and nodes in other branches. Then, traverse from node N-3-A toward the root node. For example, the nodes in other branches include node N-2-B (N-2 is the layer number of leaf node N-2-B, and B is the number of leaf node N-2-B in layer N-2) and leaf node N-1-C (N-1 is the layer number of leaf node N-1-C, and C is the number of leaf node N-1-C in layer N-1). Leaf node N-1-C is the leaf node that only hits all negative training data. Node N-2-B is the parent node of leaf node N-1-C, and node N-2-B is the child node of node N-3-A. Thus, we can establish the logical expression 4 connecting nodes N-3-A, N-2-B, and N-1-C using the logical AND symbol: [(node N-3-A) && (node N-2-B) && (node N-1-C)] The leaf nodes N-2-A and N-2-B are in a sibling relationship. The relationship between the two nodes in the sibling relationship can be a "logical OR" relationship. After establishing logical expressions for all branches of the leaf nodes of node N-3-A that have only negative sample training data that hit all, logical expressions can be established for all branches of the leaf nodes of node N-3-A that have only negative sample training data that hit all. Logical symbols can be connected in a "logical OR" manner to obtain the logical expression N _For example, from the perspective of node N-3-A, we can establish a logical expression N-3-A between logical expression 3 and logical expression 4, connected by the logical symbol "logical OR" (the merged logical expression is named after the highest-level node in the merged logical expression): [ (logical expression 3) | | (logical expression 4) ] oIn one example, the logical expression N-3-A can be: [(logical expression N-2-A) && (node N-3-A)] | | [(node N-3-A) && (node N-2-B) && (node N-1-C)] . This process continues in this way until the root node of the binary tree structure with logical symbols is reached, and until all branches related to the root node and containing leaf nodes with only fully matched negative training data are incorporated into the logical expression, the resulting logical expression can be used as the detection rule. Sometimes, if the binary tree structure is complex, the number of nodes (attribute information) in the detection rule generated using the above method may be large, leading to complexity and inconvenience in subsequent detection rule updates. Therefore, in another embodiment of the present disclosure, the generated detection rule can be split into multiple sub-detection rules, with the number of nodes (attribute information) in each sub-detection rule being less than or equal to a preset value. The preset value may include 4, 5, or 6, etc., depending on actual circumstances and is not limited by this disclosure. Alternatively, during the detection rule generation process, when performing a pre-order recursive traversal of a binary tree structure with logical symbols, if a parent node is encountered, it is stored. When a sibling node is encountered, it is first determined whether the number of nodes in the generated detection rule is greater than or equal to a preset value. If so, upon encountering a sibling node, the pre-order traversal upward is performed in multiple branches. If not, the branches of the sibling node are connected using the logical symbol "logical OR" to the branches. In another embodiment of the present disclosure, an operating environment is loaded on an electronic device. Referring to FIG3 , step S101 includes: In step S201, at least log information of the electronic device is obtained. The log information of the electronic device includes log information generated by the electronic device during the execution of multiple files in the operating environment on the electronic device. The log information of the electronic device includes at least one of the following: log information of the electronic device's process, log information of the electronic device's network, log information of the electronic device's system, and log information related to storage of the electronic device. Of course, it is understood that, depending on actual circumstances, the log information of the electronic device may also include other types of log information of the electronic device, and this disclosure is not limited to this. In one embodiment, when a file is executed in an execution environment of an electronic device, the file is often executed through a process within the electronic device. The electronic device automatically obtains specific log information related to the electronic device process and records it in the log information of the electronic device process. In this way, the log information of the electronic device process can be obtained, facilitating the subsequent acquisition of execution information of multiple files running in the execution environment.In another embodiment, when a file is executed in an operating environment (e.g., a real operating environment) within an electronic device, in certain scenarios, the electronic device may interact with an external network for the file. The electronic device automatically obtains specific information related to the network interaction and records it in the electronic device's network log information. This allows the electronic device's network log information to be obtained, facilitating subsequent acquisition of operating information for multiple files executed within the operating environment. In yet another embodiment, when a file is executed within the operating environment within an electronic device, the operating environment also automatically obtains specific system logs from the electronic device (including information related to the file executed within the operating environment within the electronic device) and records them in the electronic device's system (operating environment) log information. This allows the electronic device's system log information to be obtained, facilitating subsequent acquisition of operating information for multiple files executed within the operating environment. Furthermore, intelligence information can be obtained, including externally counted malicious files and other malicious entities associated with the malicious files, facilitating subsequent acquisition of operating information for multiple files executed within the operating environment in combination with the log information. Furthermore, for the entity "Process Name," the cmd line command line information of the process corresponding to "Process Name" can be extracted. Implicit relationship mining can be performed on the command line to extract more entities related to the process corresponding to "Process Name" from the cmd line command line information. For example, entities of preset types, such as more IP addresses connected to the process, more file names processed by the process, more file paths processed by the process, system commands invoked by the process, and keywords, can be extracted. This allows for subsequent acquisition of operational information of multiple files running in the operational environment by combining it with log information and / or intelligence information. In step S202, operational information of multiple files running in the operational environment is generated based at least on the log information of the electronic device. In one embodiment, this step can be implemented through the following process, including:
[0008] 2021. Search for multiple file names in the log information of the electronic device. Each of the multiple file names found in the log information of the electronic device can be considered an entity of a preset type. In one embodiment, the log information of the electronic device includes information related to multiple files that have been run in the operating environment. The file name can be searched in the log information of the electronic device using keywords and used as an entity of the preset type, or the file name can be searched in the log information of the electronic device using semantic analysis and used as an entity of the preset type. This disclosure does not limit the method for searching for the file name in the log information. Specifically, the file name can be searched in each of the log information such as the log information of the electronic device's process, the log information of the electronic device's network, and the log information of the electronic device's system, and used as an entity of the preset type. Furthermore, for any of the multiple file names found, the following process 2022-2023 can be executed.
[0009] 2022. Extract at least one entity of a preset type associated with the file name from the log information of the electronic device. The at least one entity of the preset type includes at least the process name of the process processing the file corresponding to the file name. Different processes in the electronic device have different process names. Although the file name also belongs to the preset type of entity, the at least one entity of the preset type associated with the file name may not include the file name. The preset types of entities to be searched may be pre-calculated and set by a technician. In other words, the entities of the preset types may be pre-calculated and set by a technician. Furthermore, the at least one entity of the preset type may include at least one of the following: the file path, the file MD5, an identifier of the action executed by the file, the external port to which the process is connected, the external IP address to which the process is connected, and the event name of an event created by the process. The association relationship can be reflected as follows: the file corresponding to the file name has the path of the file corresponding to the file name; the file corresponding to the file name has the MD5 of the file corresponding to the file name; the MD5 of the file corresponding to the file name has an identifier of the action executed by the file corresponding to the file name; the process corresponding to the process name has the external port to which the process corresponding to the process name is connected; the process corresponding to the process name has the external IP address to which the process corresponding to the process name is connected; and the process corresponding to the process name has the event name of the event created by the process corresponding to the process name. In the present disclosure, the log information of the electronic device can reflect the association relationship between entities of preset types. The log information of the electronic device can be searched for at least one entity of the preset type associated with the file name by using keywords, or by using semantic analysis. The at least one entity of the preset type includes at least the process name of the process that processes the file corresponding to the file name. The present disclosure does not limit the method of searching the log information for at least one entity of the preset type associated with the file name.
[0010] 2023. Generate execution information of the file corresponding to the file name running in the execution environment based at least on the file name and at least one entity of a preset type. In one example, the file name and the at least one entity of the preset type can be combined to obtain the execution information of the file corresponding to the file name running in the execution environment. In another embodiment, the execution information of the file corresponding to the file name running in the execution environment can be generated based on the file name, the at least one entity of the preset type, and the association between the file name and the at least one entity of the preset type. In one example, the file name, the at least one entity of the preset type, and the association between the file and the at least one entity of the preset type can be combined to obtain the execution information of the file corresponding to the file name running in the execution environment. If there are two or more entities of the at least one preset type, the combination may further include the association between the two or more entities of the preset type. In one embodiment, for a computer, the entities in the file execution information may be stored in a data structure, and the association between the entities in the file execution information may be stored in a data structure, etc. Alternatively, in another embodiment, if it is necessary to display file execution information, a knowledge graph can be generated. The knowledge graph includes nodes and relationships, where nodes are entities in the file execution information, and relationships are associations between entities in the file execution information, for viewing. In the present disclosure, when determining the first execution information and the second execution information from the execution information of multiple files in step S102, in another embodiment of the present disclosure, for the execution information of any file, the file name can be searched among multiple preset types of entities in the file execution information to determine whether the found file name is a malicious file name. If the found file name is a malicious file name, the first execution information can be obtained based on the file execution information, for example, determining the file execution information as the first execution information. Alternatively, if the found file name is not a malicious file name, the second execution information can be obtained based on the file execution information, for example, determining the file execution information as the second execution information. The same applies to the execution information of each other file. In some embodiments, the file execution information may also include associations between multiple preset types of entities involved in the execution of the file in the execution environment. For example, the association relationship between the process name and the file name includes: the process corresponding to the process name is used to process the file corresponding to the file name.When a file is running in an execution environment, a process in an electronic device processes (loads, starts, deletes, sends, modifies, etc.) the file. Thus, a file has a file name, and a process has a process name. The association between the process name and the file name may include: the process corresponding to the process name is used to process the file corresponding to the file name. Thus, an association between entities may include: one entity owns another entity, one entity performs an action on another entity (for example, one entity processes another entity), or one entity belongs to another entity. This disclosure does not limit the types of associations between entities. When the multiple entities of the preset types further include a file path, a file MD5, an identifier of an action executed by the file, an external port connected to a process, an external IP address connected to a process, and an event name of an event created by a process, the associations between these multiple entities of the preset types may include at least one of the following: the file corresponding to the file name has the path of the file corresponding to the file name, the file corresponding to the file name has the MD5 of the file corresponding to the file name, the MD5 of the file corresponding to the file name has the identifier of an action executed by the file corresponding to the file name, the process corresponding to the process name has the external port connected to the process corresponding to the process name, the process corresponding to the process name has the external IP address connected to the process corresponding to the process name, and the process corresponding to the process name has the event name of an event created by the process corresponding to the process name. Specifically, referring to FIG. 4 , step S102 includes: In step S301, for the execution information of any file, based on the multiple entities of the preset types and the associations between the multiple entities of the preset types in the execution information of the file, a knowledge graph of the file is obtained. For example, assume that the multiple preset types of entities in the running information of the file corresponding to the file name include: the file name, the path of the file corresponding to the file name, the MD5 of the file corresponding to the file name, the identifier of the behavior executed by the file corresponding to the file name, the process name of the process that processes the file corresponding to the file name, the external port to which the process that processes the file corresponding to the file name is connected, the external IP address to which the process that processes the file corresponding to the file name is connected, and the login event created by the process that processes the file corresponding to the file name, etc.The associations between these multiple entities of preset types may include: the file corresponding to the file name has the path of the file corresponding to the file name; the file corresponding to the file name has the MD5 of the file corresponding to the file name; the MD5 of the file corresponding to the file name has an identifier of the action executed by the file corresponding to the file name; the process corresponding to the process name has the external port to which the process corresponding to the process name is connected; the process corresponding to the process name has the external IP address to which the process corresponding to the process name is connected; and the process corresponding to the process name has the event name of the event created by the process corresponding to the process name. The knowledge graph generated for the file corresponding to the file name is shown in FIG5 . In step S302, the file name in the knowledge graph of the file is detected to determine whether it is the file name of a malicious file. The file name in the knowledge graph of the file is an entity of the preset type. In one embodiment, the knowledge graph of the file can be traversed. For example, based on the associations between multiple entities of the preset types in the execution information of the file, the entity in the knowledge graph of the file as the starting node can be determined, and then the knowledge graph of the file can be traversed starting from the entity as the starting node. In the present disclosure, the knowledge graph shown in FIG5 can be improved with the help of external security detection results. External security detection results can be manual detection results by technicians, or they can be detection results detected using other automated detection tools, etc., which are not limited in this disclosure. External security detection results can indicate which entities are malicious entities. For example, assuming that the external security detection results indicate that the file name in the knowledge graph shown in FIG5 is the file name of a malicious file, the process name in the knowledge graph shown in FIG5 is the process name of a malicious process, and the login event in the knowledge graph shown in FIG5 is a malicious event, then a malicious result ABNORMAL (as a node) can be added to the knowledge graph shown in FIG5 . The association between the process name and the malicious result ABNORMAL is that the process corresponding to the process name has the malicious result ABNORMAL, the association between the file name and the malicious result ABNORMAL is that the file corresponding to the file name has the malicious result ABNORMAL, and the association between the login event and the malicious result ABNORMAL is that the login event has the malicious result ABNORMAL, thereby obtaining the knowledge graph shown in FIG6 . In the knowledge graph of the file shown in FIG. 5 or 6 , a circular node represents an entity, a line with an arrow connecting two entities indicates an association relationship between the two entities, and the text on the arrow indicates the type of the association relationship between the two entities.In the association relationship between two entities connected by an arrowed line and the other entity, the entity pointed to by the arrow is passive relative to the other entity, while the other entity is active relative to the entity pointed to by the arrow. In other words, the association relationship starts from the active entity and connects to the passive entity. As can be seen, the order of active entities often comes before the order of passive entities. Therefore, the entity that serves as the starting node in the knowledge graph of the file can be considered as the entity not pointed to by an arrow. For example, in the knowledge graph shown in Figures 5 or 6, the entity not pointed to by an arrow is the process name. Therefore, traversing the knowledge graph of the file can be started from the process name. Among them, the knowledge graph shown in Figure 6 can be traversed. In the process of traversing the knowledge graph shown in Figure 6, whenever an entity is traversed in the knowledge graph, it is determined whether the entity is a file name. If the entity is a file name, it can be determined whether the file name is the file name of a malicious file. For example, it can be determined whether the file name has an association with the malicious result ABNORMAL. If the file name has an association with the malicious result ABNORMAL, it can be determined that the file name is the file name of a malicious file, or if the file name does not have an association with the malicious result ABNORMAL, it can be determined that the file name is not the file name of a malicious file. If the file name is a malicious file, step S303 can be executed. However, the external security detection results are also generated based on known malicious entities. Therefore, using the external security detection results to detect whether the file name in the knowledge graph is a malicious file name can only determine that the file name of a known malicious file is a malicious file name. The external security detection results do not involve currently unknown malicious entities, and therefore do not involve the file names of unknown malicious files. Therefore, if a file name in the knowledge graph is the name of an unknown malicious file, the external security detection results cannot be used to determine that the file name in the knowledge graph is a malicious file name, resulting in an error. Therefore, to avoid errors, in another embodiment, when traversing the knowledge graph shown in FIG5 or 6, when an entity that is a file name is encountered, the knowledge graph can be input into a trained graph neural network model. The trained graph neural network model processes the knowledge graph, obtains a detection result on whether the file name in the knowledge graph is a malicious file name, and outputs the detection result.A graph neural network model can be trained in advance. For example, a positive sample knowledge graph and a negative sample knowledge graph are obtained. The file names in the positive sample knowledge graph are not malicious file names, while the file names in the negative sample knowledge graph are malicious file names. The initialized graph neural network model is trained based on the positive sample knowledge graph and the negative sample knowledge graph until the model parameters converge, resulting in a trained graph neural network model. The intelligence and learning capabilities of the graph neural network model can improve the accuracy of determining whether a file name in the knowledge graph is a malicious file name. However, sometimes, the knowledge graph of a file includes many entities of preset types, and the relationships between many of these entities are complex, resulting in a large amount of data in the knowledge graph of the file. For example, the knowledge graphs shown in Figures 5 and 6 are merely illustrative examples. In reality, the knowledge graphs of files are often complex and include many entities of preset types. For example, the knowledge graph of a file may include multiple processes, with one process starting another process, which in turn starts another process, which in turn starts another process, which in turn starts another process, and finally another process starts the file. Furthermore, in addition to launching another process, this process also launches other entities, such as creating other events, which may also involve other IP addresses. Another process also launches yet another process, which also launches other entities, such as creating other events, which may also involve other IP addresses. Yet another process also launches yet another process, which also launches other entities, such as creating other events, which may also involve other IP addresses. Yet another process also launches this file, which also launches other entities, such as creating other events, which may also involve other IP addresses. This results in a complex knowledge graph for this file, and in turn, a large amount of data in this knowledge graph. Therefore, if this file's knowledge graph is directly input into a trained graph neural network model, the model may take a long time to process the knowledge graph, and the large amount of data may lead to memory overflow. Therefore, in order to avoid the above problem, in another embodiment, when detecting whether the file name in the entity of the preset type in the knowledge graph of the file corresponding to the file name is the file name of a malicious file, the following process can be referred to:.
[0011] 3021. Determine a process name and a file name in the knowledge graph of the file. The association between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name. The process name and file name determined in the knowledge graph of the file can both be considered entities of a preset type.
[0012] 3022. Extract a subgraph from the knowledge graph of the file. The subgraph includes the determined process name and entities of a preset type within P hops after the determined process name. P is less than or equal to X, where X is the number of node hops between the determined process name and the last entity of the preset type in the knowledge graph of the file. The number of node hops between two entities is the sum of the number of layers in which the entities exist between the two entities and the value "1," that is, the number of hops required to reach one node from another. For example, P can include 3, 4, 5, or 6, etc., depending on actual circumstances and is not limited in this disclosure. The subgraph includes the determined process name and entities of the preset type within P hops after the determined process name in the knowledge graph of the file, but does not include entities in the knowledge graph of the file that are located before the determined process name. The subgraph may reflect an association relationship between a determined process name and an entity of a preset type located within P hops after the determined process name, and may also reflect an association relationship between entities of the preset type located within P hops after the determined process name. In addition, in the subgraph, the entity of the preset type located within P hops after the determined process name has a determined file name.
[0013] 3023. Input the subgraph into the trained graph neural network model so that the trained graph neural network model processes the subgraph and obtains a detection result indicating whether the determined file name in the subgraph is a malicious file name. The subgraph input into the trained graph neural network model contains the determined process name and entities of a preset type within P hops after the determined process name, but does not contain entities that precede the determined process name in the knowledge graph of the file. The subgraph contains the determined process name and the determined file name. In the knowledge graph of the file, entities that are farther from the determined process name and the determined file name have a more distant relationship with the determined process name and the determined file name, while entities that are closer to the determined process name and the determined file name have a closer relationship with the determined process name and the determined file name. Typically, the process corresponding to the determined process name is used to process the file corresponding to the determined file name, and the process corresponding to the determined process name is proactive with respect to the file corresponding to the determined file name. Therefore, in the knowledge graph of the file, entities that precede the determined process name often have a more distant relationship with the file corresponding to the determined file name. Therefore, the preset type of entities within P hops can focus on the two key pieces of information in the file's knowledge graph: the determined process name and the determined file name. They can also focus on other entities closely associated with the determined process name and file name, discarding entities further away from the determined process name and file name, as well as entities preceding the determined process name. This reduces the use of entities with more distant associations with the determined process name and file name. This allows for focusing on key information, while not affecting the accuracy of the trained graph neural network model's detection of whether the determined file name is a malicious file. This reduces processing time and the amount of data processed, minimizing memory overflows. If the file name in the file's knowledge graph is a malicious file name, in step S303, first execution information is obtained based on the file's knowledge graph. The file name in the file's knowledge graph is an entity of the preset type. In one embodiment, when obtaining the first execution information based on the file's knowledge graph, the file's knowledge graph can be used as the first execution information. However, sometimes, the knowledge graph of the file includes many entities of preset types, and the association relationships between many entities of preset types are very complex, resulting in a large amount of data in the knowledge graph of the file.For example, the knowledge graphs shown in Figures 5 and 6 are merely illustrative examples. In practice, file knowledge graphs are often complex and include many entities of predefined types. For example, a file's knowledge graph may include multiple processes: one process initiates another process, which in turn initiates another process, which in turn initiates another process, which in turn initiates another process, which in turn initiates the file. Furthermore, in addition to initiating another process, the first process also initiates other entities, such as creating other events, which may also involve other IP addresses. In addition to initiating yet another process, the second process also initiates other entities, such as creating other events, which may also involve other IP addresses. In addition to initiating yet another process, the third process also initiates other entities, such as creating other events, which may also involve other IP addresses. In addition to initiating the file, the third process also initiates other entities, such as creating other events, which may also involve other IP addresses. This results in a complex knowledge graph for the file, which in turn results in a large amount of data in the knowledge graph. Therefore, if the knowledge graph of the file is directly used as the first operational information, the amount of data in the first operational information may be very large. Subsequently, when the decision tree model is trained using this first operational information, the calculation of the first operational information may take a long time, and the large amount of data may lead to memory overflow. Therefore, to avoid the above issues, in another embodiment, when obtaining the first operational information based on the knowledge graph of the file, the following process can be used:
[0014] 3031. Determine the process name of the malicious process in the knowledge graph of the file.
[0015] 3032. In the knowledge graph of the file, select an entity of a preset type within N hops after the process name of the malicious process.
[0016] N is less than or equal to Y, where Y is the number of node hops between the process name of the malicious process and the entity of the preset type at the end of the knowledge graph of the file. For example, N may include 3, 4, 5, or 6, etc., and the specific value may be determined based on actual circumstances and is not limited in this disclosure.
[0017] 3033. Determine whether a file name of the malicious file exists in an entity of a preset type located within N hops after the process name of the malicious process.
[0018] 3034. If the file name of a malicious file is present in an entity of a preset type located within N hops after the process name of the malicious process, a subgraph is extracted from the knowledge graph for the file. The subgraph includes the process name of the malicious process and entities of the preset type located within N hops after the process name of the malicious process. The subgraph includes the file name of the malicious file, and the malicious process is used to process the malicious file. The subgraph includes the process name of the malicious process and entities of the preset type located within N hops after the process name of the malicious process in the knowledge graph, but does not include entities located before the process name of the malicious process in the knowledge graph. The subgraph may reflect an association relationship between the process name of a malicious process and an entity of a preset type located within N hops after the process name of the malicious process, and may reflect an association relationship between entities of a preset type located within N hops after the process name of the malicious process. In addition, in the subgraph, the entity of the preset type located within N hops after the process name of the malicious process has a file name of a malicious file, wherein the malicious process is used to process the malicious file, and the malicious file is the file corresponding to the file name. That is, there is an association relationship between the process name of the malicious process and the file name of the malicious file, and the malicious process corresponding to the process name is used to process the malicious file corresponding to the file name.
[0019] 3035. Obtain first operation information based on the subgraph. For example, the subgraph may be used as the first operation information. The first operation information includes the process name of the malicious process and entities of a preset type located within N hops after the process name of the malicious process, but does not include entities located before the process name of the malicious process in the knowledge graph. The first operation information may reflect the association between the process name of the malicious process and entities of the preset type located within N hops after the process name of the malicious process, and may also reflect the association between entities of the preset type located within N hops after the process name of the malicious process. Furthermore, in the first operation information, the entities of the preset type located within N hops after the process name of the malicious process include the file name of a malicious file, wherein the malicious process is used to process the malicious file. The malicious file is the file corresponding to the file name. That is, there is an association between the process name of the malicious process and the file name of the malicious file, and the malicious process corresponding to the process name is used to process the malicious file corresponding to the file name. The first running information contains the process name of a malicious process and the file name of a malicious file processed by the malicious process. In the knowledge graph, entities that are further away from the process name of the malicious process and the file name of the malicious file have a more distant relationship with them, while entities that are closer to the process name of the malicious process and the file name of the malicious file have a closer relationship with them. Generally, a process that processes a malicious file can be considered a malicious process, and malicious processes are proactive in processing malicious files. Therefore, in the knowledge graph, entities that precede the process name of the malicious process tend to have a more distant relationship with the malicious file. Therefore, the preset entity types within N hops can focus on the two key pieces of information in the knowledge graph: the process name of the malicious process and the file name of the malicious file. They can also focus on other entities closely associated with the process name of the malicious process and the file name of the malicious file. This discards entities that are further away from the process name of the malicious process and the file name of the malicious file, as well as entities that precede the process name of the malicious process and the file name of the malicious file. This reduces the use of entities that are more distantly associated with the process name of the malicious process and the file name of the malicious file. This allows for focusing on key information, reducing computational time and the amount of data required for computation without affecting the accuracy of subsequently generated detection rules, thereby minimizing memory overflows. It should be noted that if the process of steps S3021-3023 is executed in step S302 and a detection result is obtained indicating that the file name in the subgraph is the file name of a malicious file, then step 3035 can be directly executed in this step, and steps 3031-3034 can be skipped, thereby saving time and computing resources.If the file name in the knowledge graph of the file is not a malicious file name, in step S304, second operation information is obtained based on the knowledge graph of the file. The file name in the knowledge graph of the file is an entity of a preset type. In one embodiment, when obtaining the second operation information based on the knowledge graph of the file, the knowledge graph of the file can be used as the second operation information. However, sometimes, the knowledge graph of the file includes many entities of the preset type, and the relationships between many of these entities of the preset type are complex, resulting in a large amount of data in the knowledge graph of the file. For example, the knowledge graphs shown in Figures 5 or 6 are merely illustrative examples. In actual situations, the knowledge graph of a file is often complex and includes many entities of the preset type. For example, the knowledge graph of a file may include multiple processes, one process starting another process, which in turn starts another process, which in turn starts another process, which in turn starts another process, which in turn starts the file. In addition, in addition to starting another process, the process may also start other entities, such as creating other events, and may also involve other IP addresses. In addition to launching yet another process, another process also launches other entities, for example, creates other events, and may also involve other IP addresses. In addition to launching yet another process, another process also launches other entities, for example, creates other events, and may also involve other IP addresses. In addition to launching the file, yet another process also launches other entities, for example, creates other events, and may also involve other IP addresses. This results in a complex knowledge graph for the file, and in turn, a large amount of data in the knowledge graph. Therefore, if the knowledge graph of the file is directly used as the second operation information, the data volume of the second operation information may be very large. When the second operation information is subsequently used to train a decision tree model, the calculation of the second operation information may take a long time, and the large data volume may lead to memory overflow. Therefore, to avoid the above problems, in another embodiment, when obtaining the second operation information based on the knowledge graph of the file, the following process can be used:
[0020] 3041. Determine a process name and a file name in the knowledge graph of the file. The association between the determined process name and the determined file name includes: a process corresponding to the determined process name is used to process the file corresponding to the determined file name. The process name and file name determined in the knowledge graph of the file can both be considered entities of a preset type.
[0021] 3042. Extract a subgraph from the knowledge graph of the file, where the subgraph includes the determined process name and entities of a preset type within M hops after the determined process name.
[0022] M is less than or equal to Z, where Z is the number of node hops between the determined process name and the terminal entity of the preset type in the knowledge graph. For example, M may include 3, 4, 5, or 6, depending on actual circumstances and not limited in this disclosure. M may be greater than or equal to the aforementioned N. The subgraph contains the determined process name and an entity of the preset type within M hops after the determined process name in the knowledge graph of the file, but does not contain an entity before the determined process name in the knowledge graph of the file. The subgraph may reflect the association between the determined process name and the entity of the preset type within M hops after the determined process name, as well as the association between entities of the preset type within M hops after the determined process name. In the subgraph, the entity of the preset type within M hops after the determined process name contains the determined file name. The determined process name is the process name of a non-malicious process, and the determined file name is the file name of a non-malicious file. The non-malicious process is used to process non-malicious files. That is, there is an association between the process name of the non-malicious process and the file name of the non-malicious file, and the non-malicious process corresponding to the process name is used to process the non-malicious file corresponding to the file name. 3043. Obtain second operation information based on the subgraph. For example, the subgraph is used as the second operation information. The second operation information contains the process name of the non-malicious process and entities of a preset type within M hops after the process name of the non-malicious process, but does not contain entities that are located before the process name of the non-malicious process in the knowledge graph. The second operation information may reflect the association between the process name of the non-malicious process and entities of the preset type within M hops after the process name of the non-malicious process, as well as the association between entities of the preset type within M hops after the process name of the non-malicious process. In addition, in the second operation information, the entities of the preset type within M hops after the process name of the non-malicious process contain the file name of the non-malicious file, and the non-malicious process is used to process non-malicious files. The second running information includes the process name of a non-malicious process and the file name of a non-malicious file processed by the non-malicious process. In the knowledge graph, entities that are farther away from the process name of the non-malicious process and the file name of the non-malicious file have a more distant association with the process name of the non-malicious process and the file name of the non-malicious file, and entities that are closer to the process name of the non-malicious process and the file name of the non-malicious file have a closer association with the process name of the non-malicious process and the file name of the non-malicious file.Typically, a process processing a non-malicious file can be considered a non-malicious process. Non-malicious processes are proactive towards non-malicious files. Therefore, in the knowledge graph for that file, entities preceding the non-malicious process's name often have a more distant relationship with the non-malicious file. Therefore, the pre-set entity type within M hops can focus on two key pieces of information in the knowledge graph: the non-malicious process's name and the non-malicious file's file name. It can also focus on other entities closely associated with the non-malicious process's name and the non-malicious file's file name, discarding entities that are more distant from the non-malicious process's name and the non-malicious file's file name, as well as entities preceding the non-malicious process's name. This reduces the use of entities with more distant relationships with the non-malicious process's name and the non-malicious file's file name. This allows for a focus on key information, while not affecting the accuracy of subsequently generated detection rules. This can also reduce computational time and the amount of data required for computation, minimizing memory overflows and other issues. It should be noted that if the process of S3021 to S3023 is executed in step S302 and the detection result indicates that the file name in the subgraph is not a malicious file, then step 3044 can be directly executed in this step, and steps 3041 to 3043 can be skipped, thereby saving time and computing resources. If the operating environment is a sandbox environment, the following defects 1-3 may exist:
[0023] 1. A sandbox environment is an isolated operating environment. For security reasons, it is typically not connected to the internet. If a file is malicious, malicious behaviors that require an internet connection to trigger may not be triggered. In other words, the file will not perform malicious behaviors that require an internet connection in the sandbox environment, resulting in the sandbox environment failing to capture the malicious behaviors that require an internet connection. Because the sandbox environment fails to capture malicious behaviors that require an internet connection, detection rules fail to include information about the malicious behaviors that cannot be captured. Consequently, the generated detection rules omit information about some malicious behaviors of the malicious file, resulting in incompleteness and low generalization. Therefore, the detection rules generated based on the present disclosure have low accuracy in detecting whether a file's behavior is malicious.
[0024] 2. Currently, various anti-sandboxing methods exist. For example, if a file is malicious and has anti-sandboxing methods, the file will detect whether the file's execution environment is a sandbox environment. If the file's execution environment is not a sandbox environment, the file may perform malicious behavior. Alternatively, if the file's execution environment is a sandbox environment, the file will generally not perform malicious behavior. Therefore, if the file's execution environment is a sandbox environment, the malicious behavior of the file running in the sandbox environment will often not be triggered. In other words, the file running in the sandbox environment will not perform malicious behavior, resulting in the inability to capture the malicious behavior of the file running in the sandbox environment. Since the malicious behavior of the file running in the sandbox environment cannot be captured, the detection rules will not include relevant information about the malicious behavior that cannot be captured. As a result, the generated detection rules will omit relevant information about some malicious behaviors of the malicious file, resulting in incompleteness and low generalization. Therefore, the detection rules generated based on the present disclosure have low accuracy in detecting whether the file's behavior contains malicious behavior. 3. The sandbox environment is a simulated runtime environment, not a real one. If a file is malicious, it may sometimes require that its runtime environment be a real one (not a simulated one, but the one commonly used by the general public). For example, if the file detects that the runtime environment is a real one, it will determine that the environment meets the file's requirements, and the file will only perform malicious actions in a real environment. Alternatively, if the file detects that the environment is not a real one, it will determine that the environment does not meet the file's requirements, and the file will generally not perform malicious actions in a non-real environment. Therefore, if the file's runtime environment is not a real one, the malicious behavior of the file running in the real environment often fails to trigger. That is, the file running in the real environment does not perform malicious behavior, resulting in the inability to capture the malicious behavior of the file running in the real environment. Because the malicious behavior of the file running in the real environment cannot be captured, the detection rules do not include relevant information about the malicious behavior that cannot be captured. As a result, the generated detection rules omit relevant information about some malicious behaviors of the malicious file, resulting in incompleteness and low generalization of the generated detection rules. Therefore, the detection rules generated based on the present disclosure have low accuracy in detecting whether a file's behavior contains malicious behavior.To this end, in one embodiment, the runtime environment in the present disclosure may include a real runtime environment. A real runtime environment is not an isolated runtime environment; rather, it is an runtime environment that can interact with the outside world. A real runtime environment is not a simulated runtime environment. For example, a real runtime environment is not a sandbox environment. A real runtime environment may include an operating system installed on a real electronic device. Operating systems may include operating systems used by a wide range of manufacturers on the market, such as the Windows operating system, the Android operating system, or the Linux operating system. Thus, the detection rules in the present disclosure are decoupled from the simulated runtime environment (e.g., the sandbox environment). The detection rules in the present disclosure are automatically generated according to logic based on information related to the actual execution of a file in the real runtime environment. Files running in the real runtime environment often perform malicious behavior, so the statistical information related to the actual execution of a file in the real runtime environment will include information related to the malicious behavior. This allows the generated detection rules to include more information related to the malicious behavior of the malicious file, or even all of the information related to the malicious behavior, thereby minimizing omissions of information related to the malicious behavior of the malicious file, thereby making the generated detection rules complete and highly generalizable. As can be seen, the detection rules generated based on the present disclosure have a high accuracy rate for detecting whether a file to be detected contains malicious behavior. After the detection rules are established, they can be put into use online, for example, to detect whether a file is malicious. FIG7 illustrates a file detection method disclosed herein, which is applied to electronic devices in a cloud computing scenario. Electronic devices in a cloud computing scenario may include physical devices or virtual devices. Physical devices include servers or terminals, and virtual devices may include virtual machines hosted on physical devices. Virtual machines may include ECSs, etc. The method may include: In step S401, obtaining operation information of the file to be detected. The operation information of the file to be detected includes multiple preset entities involved in the execution of the file to be detected in the execution environment. The multiple preset entities include at least the file name of the file to be detected and the name of the process used to process the file to be detected. In one embodiment, the detection rule may open an API (Application Programming Interface) of the detection rule to the outside world, so that when it is needed to detect whether the file to be detected is a malicious file, the detection rule may be called through the API of the detection rule to detect whether the file to be detected is a malicious file.Alternatively, in another embodiment, the cloud computing system includes a database for storing the operational information of each file in the cloud computing system. The operational information of each file is associated with its MD5. The database exposes its API to the public. When a detection rule is needed to detect whether a file is malicious, the MD5 of the file can be obtained first. Then, based on the MD5 of the file, the database API is called to retrieve the operational information of the file from the database. The file is then detected as malicious based on the operational information and the detection rule. In step S402, attribute information of multiple preset entities is obtained from the operational information of the file to be detected. The operational information of the file to be detected also includes the associations between the multiple preset entities involved in the execution of the file to be detected in the execution environment. In this manner, when obtaining attribute information of multiple entities of preset types in the execution information of the file to be detected, a knowledge graph of the file to be detected can be obtained based on the multiple entities of preset types and the associations between the entities of the preset types in the execution information of the file to be detected. A process name and a file name are determined in the knowledge graph of the file to be detected. The association between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name. A subgraph is extracted from the knowledge graph of the file to be detected, where the subgraph includes the determined process name and entities of the preset type within Q hops after the determined process name, where Q is less than or equal to V, and V is the number of node hops between the determined process name and the last entity of the preset type in the knowledge graph. Attribute information of each entity of the preset type in the subgraph is obtained. In step S403, based on the attribute information and pre-generated detection rules for detecting malicious files, whether the file to be detected is a malicious file is detected. The attribute information can be matched against pre-generated detection rules for detecting malicious files. If the attribute information matches the detection rules, the file to be detected can be determined to be malicious. Alternatively, if the attribute information does not match the detection rules, the file to be detected can be determined to be non-malicious. In one possible scenario, multiple pieces of execution information for a particular file may be retrieved from the database. In this case, this may be because the file has been executed on different electronic devices in the cloud computing system. Each electronic device generates execution information for the file when executing the file. The generated execution information is then associated with the file's MD5 and stored in the database.Thus, in one embodiment, there are multiple pieces of execution information for the file to be detected; the file names in the multiple pieces of execution information are all the file names of the file to be detected, but the process names in the multiple pieces of execution information are different. Thus, when obtaining attribute information of multiple preset types of entities in the execution information of the file to be detected, for any piece of execution information of the file to be detected, the attribute information of the multiple preset types of entities in the execution information is obtained to obtain the attribute information corresponding to the execution information. The same process is repeated for each piece of execution information of the file to be detected, thereby obtaining the attribute information corresponding to each piece of execution information of the file to be detected. Accordingly, when detecting whether a file to be detected is a malicious file based on the attribute information and pre-generated detection rules for detecting malicious files, the attribute information corresponding to any piece of operational information of the file to be detected is matched with the detection rule to obtain a detection result indicating whether the file to be detected is a malicious file. The same process is repeated for each piece of operational information corresponding to other pieces of operational information of the file to be detected. The number of attribute information corresponding to operational information that indicates the file to be detected is a malicious file is then counted. If the ratio between this number and the number of pieces of operational information of the file to be detected is greater than a preset threshold, the file to be detected is determined to be a malicious file. Alternatively, if the ratio between this number and the number of pieces of operational information of the file to be detected is less than or equal to the preset threshold, the file to be detected is determined to be not a malicious file. The preset threshold may include 0.5, 0.55, or 0.6, etc., and may be determined based on actual circumstances and is not limited in this disclosure. Furthermore, the method further includes: screening out falsely detected files from among the detected files that are not actually malicious files but are detected as such by the detection rules (this may be falsely detected files that are manually indicated as not actually malicious files but are detected as such by the detection rules, etc.); determining, in the detection rules, the logical expression involved in determining that the falsely detected files are malicious files (it is known which logical expression in the detection rules the falsely detected files match); and deleting the determined logical expression from the detection rules. Furthermore, manually constructing a new logical expression for the falsely detected files and inputting the new logical expression into the electronic device. The electronic device receives the manually input new logical expression and adds the new logical expression to the detection rules to update the detection rules.Among them, if the operating environment is a sandbox environment, the following defects 1-3 may exist: 1. The sandbox environment is an isolated operating environment. For security reasons, the sandbox environment is usually not connected to the external network. If the file to be detected is a malicious file, the malicious behavior in the file to be detected that requires an Internet connection to be triggered may not be triggered. That is, the file to be detected will not perform malicious behaviors that require an Internet connection to be triggered in the sandbox environment, resulting in the inability to capture the malicious behaviors of the file to be detected that require an Internet connection to be triggered in the sandbox environment. Since the malicious behaviors of the file to be detected that require an Internet connection to be triggered cannot be captured in the sandbox environment, if the file to be detected has no other malicious behaviors, a detection result indicating that the behavior of the file to be detected in the sandbox environment does not have malicious behaviors will be output, and then it will be determined that the file to be detected is not a malicious file but a safe file, which may cause the malicious file to fail to be detected, resulting in missed detection. It can be seen that the detection accuracy is low.
[0025] 2. Currently, various anti-sandbox environment methods exist. For example, if the file to be detected is malicious and has anti-sandbox environment methods, the file to be detected will detect whether the running environment of the file to be detected is a sandbox environment. If the running environment of the file to be detected is not a sandbox environment, the file to be detected may perform malicious behavior. Alternatively, if the running environment of the file to be detected is a sandbox environment, the file to be detected will usually not perform malicious behavior. It can be seen that if the running environment of the file to be detected is a sandbox environment, the malicious behavior of the file to be detected running in the sandbox environment will often fail to be triggered. That is, the file to be detected running in the sandbox environment will not perform malicious behavior, resulting in the inability to capture the malicious behavior of the file to be detected running in the sandbox environment. Since the malicious behavior of the file to be detected running in the sandbox environment cannot be captured, when the file to be detected has no other malicious behavior, the detection result that the behavior of the file to be detected in the sandbox environment does not have malicious behavior will be output, and then it will be determined that the file to be detected is not a malicious file, but a safe file, which may cause the malicious file to fail to be detected, resulting in missed detection. It can be seen that the detection accuracy is low.
[0026] 3. The sandbox environment is a simulated runtime environment, not a real one. If the file being tested is malicious, it may sometimes require that its runtime environment be a real one (not a simulated one, but rather one commonly used by the general public). For example, if the file being tested detects that its runtime environment is a real one, it will determine that the environment meets its requirements, and the file can only perform malicious actions in a real environment. Alternatively, if the file being tested detects that its runtime environment is not a real one, it will determine that the environment does not meet its requirements, and the file will generally not perform malicious actions in a non-real environment. Therefore, if the runtime environment of the file to be detected is a non-real runtime environment, the malicious behavior of the file to be detected running in the non-real runtime environment will often fail to be triggered. In other words, the file to be detected running in the non-real runtime environment will not perform malicious behavior, resulting in the inability to capture the malicious behavior of the file to be detected running in the non-real runtime environment. Since the malicious behavior of the file to be detected running in the non-real runtime environment cannot be captured, if the file to be detected does not have other malicious behavior, the detection result will be output that the behavior of the file to be detected in the sandbox environment does not have malicious behavior. In turn, the file to be detected will be determined to be a safe file rather than a malicious file. This may cause the malicious file to fail to be detected, resulting in missed detections, which can be seen as low detection accuracy. In one embodiment, the runtime environment in the present disclosure can include a real runtime environment, which enables the present disclosure to have the following beneficial effects 1-3:
[0027] 1. The present disclosure does not use a simulated operating environment, for example, it does not use a sandbox environment, but uses a real operating environment. The real operating environment is often connected to the Internet. The file to be detected runs in the real operating environment. If the file to be detected is a malicious file, the malicious behavior in the file to be detected that requires an Internet connection to be triggered will often be triggered. That is, the file to be detected will often execute malicious behavior that requires an Internet connection to be triggered in the real operating environment, so that the malicious behavior that requires an Internet connection to be triggered can be captured. Since the malicious behavior of the file to be detected in the real operating environment can be captured, a detection result that the file to be detected has malicious behavior in the real operating environment can often be obtained according to the detection rules, so that the file to be detected can be determined to be a malicious file, so that the malicious file can be detected and missed detection can be avoided. It can be seen that the detection accuracy can be improved. 2. If various anti-sandbox environment means currently exist, for example, if the file to be detected is a malicious file and there are anti-sandbox environment means, the file to be detected will detect whether the running environment of the file to be detected is a sandbox environment. If the running environment of the file to be detected is not a sandbox environment, the file to be detected may perform malicious behavior. Alternatively, if the running environment of the file to be detected is a sandbox environment, the file to be detected will often not perform malicious behavior. Since the present disclosure does not use a simulated runtime environment, for example, a sandbox environment, but rather a real runtime environment, that is, the file to be detected runs in a real runtime environment, the file to be detected will detect that the runtime environment in which the file to be detected is located is a real runtime environment, not a sandbox environment. Therefore, even if the file to be detected has anti-sandbox environment protection measures, since the real runtime environment is not a sandbox environment, the file to be detected running in the real runtime environment will often perform malicious behavior, which will often trigger the malicious behavior of the file to be detected running in the real runtime environment. In other words, the file to be detected running in the real runtime environment will often perform malicious behavior, which can be captured. Since the malicious behavior of the file to be detected running in the real runtime environment can be captured, a detection result indicating that the file to be detected has malicious behavior in the real runtime environment can often be obtained according to the detection rules. Therefore, the file to be detected can be determined to be a malicious file, so that the malicious file can be detected and missed detection can be avoided. As can be seen, the detection accuracy can be improved.
[0028] 3. Since the present disclosure does not use a simulated operating environment, for example, does not use a sandbox environment, but uses a real operating environment, the file to be detected runs in the real operating environment. If the file to be detected is a malicious file, then when the file to be detected detects that the operating environment where the file to be detected is located is a real operating environment (not a simulated operating environment, but an operating environment commonly used by the majority of users), the file to be detected may perform malicious behavior, so that the malicious behavior of the file to be detected running in the real operating environment is often triggered. In other words, the file to be detected running in the real operating environment often performs malicious behavior, so that the malicious behavior of the file to be detected running in the real operating environment can be captured. Since the malicious behavior of the file to be detected running in the real operating environment can be captured, a detection result indicating that the file to be detected has malicious behavior in the real operating environment can often be obtained according to the detection rules, so that the file to be detected can be determined to be a malicious file, so that the malicious file can be detected and missed detection can be avoided. As can be seen, the detection accuracy can be improved. It should be noted that, for simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the present disclosure is not limited by the order of the actions described, as certain steps may be performed in a different order or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are optional, and the actions involved are not necessarily required by the present disclosure. 8 , a block diagram of a device for establishing detection rules is shown, including: a first acquisition module 11, configured to acquire operation information of multiple files running in an operation environment, the operation information of the files including: multiple entities of preset types involved in the process of the files running in the operation environment, the multiple entities of preset types including at least: the file name and the name of the process for processing the file; a determination module 12, configured to determine, from the operation information of the multiple files, first operation information and second operation information, wherein the file name in the first operation information is the file name of a malicious file, and the file name in the second operation information is not the file name of a malicious file; a second acquisition module 13, configured to acquire attribute information of multiple entities of preset types in the first operation information, and to acquire attribute information of multiple entities of preset types in the second operation information; and a generation module 14, configured to generate a detection rule for detecting malicious files based on the attribute information of the multiple entities of preset types in the first operation information and the attribute information of the multiple entities of preset types in the second operation information.The generation module includes: a training unit configured to train a decision tree model based on attribute information of multiple preset entity types in the first operation information and attribute information of multiple preset entity types in the second operation information to obtain a binary tree structure; a first generation unit configured to generate detection rules for detecting malicious files based on the binary tree structure. The first generation unit is specifically configured to traverse the binary tree structure using a pre-order traversal method to generate detection rules for detecting malicious files. The operation environment is loaded on the electronic device; the first acquisition module includes: a first acquisition unit configured to obtain at least log information of the electronic device, the log information of the electronic device including log information generated by the electronic device during the execution of multiple files in the operation environment on the electronic device; and a second generation unit configured to generate operation information of the multiple files executed in the operation environment based on at least the log information of the electronic device. The second generation unit includes: a search subunit configured to search for multiple file names in the log information of the electronic device; a generation subunit configured to extract, for any one of the multiple file names found, at least one entity of a preset type associated with the file name from the log information of the electronic device, the at least one entity of the preset type including at least the name of a process processing the file corresponding to the file name; and to generate, based on at least the file name and the at least one entity of the preset type, execution information of the file corresponding to the file name running in the execution environment. The file execution information also includes: the associations between the multiple entities of the preset types involved in the execution of the file in the execution environment; and the determination module includes: a second acquisition unit configured to acquire, for any one of the file execution information, a knowledge graph of the file based on the multiple entities of the preset types and the associations between the multiple entities of the preset types in the file execution information; a detection unit configured to detect whether the file name in the knowledge graph is a malicious file name; a third acquisition unit configured to acquire, based on the knowledge graph, first execution information if the file name in the knowledge graph is a malicious file name; or a fourth acquisition unit configured to acquire, based on the knowledge graph, second execution information if the file name in the knowledge graph is not a malicious file name.Among them, the detection unit includes: a first determination sub-unit, which is used to determine the process name and the file name in the knowledge graph; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a first interception sub-unit, which is used to intercept a sub-graph in the knowledge graph, and the sub-graph includes the determined process name and an entity of a preset type within P hops after the determined process name, P is less than or equal to X, and X is the number of node hops between the determined process name and the entity of the preset type at the end of the knowledge graph; an input sub-unit, which is used to input the sub-graph into the trained graph neural network model, so that the trained graph neural network model processes the sub-graph to obtain a detection result of whether the determined file name in the sub-graph is the file name of a malicious file. Among them, the third acquisition unit includes: a second determination subunit, which is used to determine the process name of the malicious process in the knowledge graph; a selection subunit, which is used to select an entity of a preset type within N hops after the process name of the malicious process in the knowledge graph, where N is less than or equal to Y, and Y is the number of node hops between the process name of the malicious process and the entity of the preset type at the end of the knowledge graph; a third determination subunit, which is used to determine whether there is a file name of a malicious file in the entity of the preset type within N hops after the process name of the malicious process; a second interception subunit, which is used to intercept a subgraph in the knowledge graph when there is a file name of a malicious file in the entity of the preset type within N hops after the process name of the malicious process, the subgraph including the process name of the malicious process and an entity of the preset type within N hops after the process name of the malicious process, and the malicious process is used to process the malicious file; a first acquisition subunit, which is used to obtain first operation information according to the subgraph. Among them, the fourth acquisition unit includes: a fourth determination subunit, used to determine the process name and file name in the knowledge graph, and the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a third interception subunit, used to intercept a subgraph in the knowledge graph, the subgraph including the determined process name and an entity of a preset type within M hops after the determined process name, M is less than or equal to Z, and Z is the number of node hops between the determined process name and the entity of the preset type at the end of the knowledge graph; a second acquisition subunit, used to obtain second operation information according to the subgraph.9 , a block diagram of a file detection apparatus according to the present disclosure is shown, comprising: a third acquisition module 21 for acquiring running information of a file to be detected, where the running information of the file to be detected includes multiple entities of preset types involved in the running of the file to be detected in the running environment, where the multiple entities of preset types include at least the file name of the file to be detected and the name of the process for processing the file to be detected; a fourth acquisition module 22 for acquiring attribute information of the multiple entities of preset types from the running information of the file to be detected; and a detection module 23 for detecting whether the file to be detected is a malicious file based on the attribute information and pre-generated detection rules for detecting malicious files. Among them, the running information of the file to be detected also includes: the association relationship between multiple preset types of entities involved in the running process of the file to be detected in the running environment; the fourth acquisition module includes: a fifth acquisition unit, which is used to obtain the knowledge graph of the file to be detected based on the multiple preset types of entities in the running information of the file to be detected and the association relationship between the multiple preset types of entities; a first determination unit, which is used to determine the process name and the file name in the knowledge graph of the file to be detected; the association relationship between the determined process name and the determined file name includes that the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a capture unit, which is used to capture a subgraph in the knowledge graph of the file to be detected, wherein the subgraph includes the determined process name and the entities of the preset type within Q hops after the determined process name, where Q is less than or equal to V, and V is the number of node hops between the determined process name and the entity of the preset type at the end of the knowledge graph; a sixth acquisition unit, which is used to obtain attribute information of each entity of the preset type in the subgraph. Among them, there are multiple running information of the file to be detected; the file names in the multiple running information are all the file names of the file to be detected, and the process names in the multiple running information are different; accordingly, the fourth acquisition module includes: a seventh acquisition unit, which is used to obtain attribute information of multiple preset types of entities in any running information of the file to be detected, and obtain attribute information corresponding to the running information; accordingly, the detection module includes: a matching unit, which is used to match the attribute information corresponding to any running information of the file to be detected with the detection rule, and obtain a detection result indicating whether the attribute information corresponding to the running information is a malicious file; a statistical unit, which is used to count the number of attribute information corresponding to the running information indicating the detection result that the file to be detected is a malicious file; a second determination unit, which is used to determine that the file to be detected is a malicious file when the ratio between the number and the number of multiple running information of the file to be detected is greater than a preset threshold.The embodiments of the present disclosure further provide a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiments of the present disclosure. o The embodiments of the present disclosure provide one or more machine-readable media having instructions stored thereon. When executed by one or more processors, the electronic device executes one or more methods described in the above embodiments. In the embodiments of the present disclosure, the electronic device includes a server, a gateway, a sub-device, and the like, wherein the sub-device is an IoT device. The embodiments of the present disclosure can be implemented as an apparatus configured as desired using any appropriate hardware, firmware, software, or any combination thereof. The apparatus may include a server (cluster), a terminal device such as an IoT device, and other electronic devices. FIG10 schematically illustrates an exemplary apparatus 1300 that can be used to implement various embodiments of the present disclosure. o For one embodiment, FIG10 shows an exemplary apparatus 1300 having one or more processors 1302, a control module (chip set) 1304 coupled to at least one of the processor(s) 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304. O Processor 1302 may include one or more single-core or multi-core processors. Processor 1302 may include any combination of general-purpose processors or specialized processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, apparatus 1300 can function as a server device such as a gateway in the embodiments of the present disclosure. Apparatus 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 configured to execute instructions 1314 in conjunction with the one or more computer-readable media to implement a module and thereby perform actions in the present disclosure. oThe control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1302 and / or any suitable device or component in communication with the control module 1304. The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module. The memory 1306 may be used, for example, to load and store data and / or instructions 1314 for the apparatus 1300. O For one embodiment, memory 1306 may include any suitable volatile memory, such as a suitable DRAM (Dynamic Random Access Memory). In some embodiments, memory 1306 may include double data rate quad synchronous dynamic random access memory (DDR4 SDRAM). Control module 1304 may include one or more input / output controllers to provide an interface to NVM / storage device 1308 and (one or more) input / output devices 1310. For example, NVM / storage device 1308 may be used to store data and / or instructions 1314. O NVM / storage 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives). NVM / storage 1308 may include storage resources that are physically part of the device on which apparatus 1300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage 1308 may be accessible via a network via input / output device(s) 1310. Input / output device(s) 1310 may provide an interface for apparatus 1300 to communicate with any other suitable device and may include a communication component, a phonetic component, a sensor component, and the like. The network interface 1312 can provide an interface for the device 1300 to communicate through one or more networks. The device 1300 can wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing a wireless network based on a communication standard, such as WiFi.
[0029] Wireless communication is carried out using 2G (Wireless Fidelity), 2G (Second Generation), 3G (Third Generation), 4G (Fourth Generation), 5G (Fifth Generation), etc., or a combination of these.
[0030] At least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 1304. In one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers of the control module 1304 to form a system-in-package (SIP). In one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304. In one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304 to form a system-on-chip (SoC). In various embodiments, the apparatus 1300 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the apparatus 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker. Embodiments of the present disclosure provide an electronic device comprising: one or more processors; and one or more machine-readable media storing instructions, which, when executed by the one or more processors, enable the electronic device to perform one or more methods as described herein. Embodiments of the present disclosure provide a computer program product, which, when executed by the processor of the electronic device, enables the electronic device to perform one or more methods as described herein. The device embodiments are generally similar to the method embodiments, so the description is simplified. For relevant details, refer to the description of the method embodiments. The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referenced. The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal device to produce a machine, such that the instructions executed by the processor of the computer or other programmable information processing terminal device produce means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable information processing terminal device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable information processing terminal device, such that the computer or other programmable information processing terminal device executes a series of operational steps to produce a computer-implemented process, such that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. Although preferred embodiments of the present disclosure have been described, those skilled in the art, once aware of the basic inventive concepts, may make additional changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as encompassing the preferred embodiment and all variations and modifications falling within the scope of the disclosed embodiments. Finally, it should be noted that, herein, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, article, or terminal device comprising the element.The above describes in detail the method and apparatus for establishing detection rules and the method and apparatus for detecting files provided by the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The descriptions of the above embodiments are intended only to facilitate understanding of the method and core concepts of the present disclosure. Furthermore, those skilled in the art will appreciate that variations in the specific implementation methods and scope of application may occur based on the concepts of the present disclosure. Therefore, the contents of this specification should not be construed as limiting the present disclosure.
Claims
Claims 1. A method for establishing detection rules, wherein The method includes: obtaining the running information of multiple files running in a running environment, where the running information of the files includes: multiple preset types of entities involved in the process of the files running in the running environment, and the multiple preset types of entities at least include: the file names of the files and the process names for processing the files; determining first running information and second running information in the running information of the multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; obtaining the attribute information of the multiple preset types of entities in the first running information, and obtaining the attribute information of the multiple preset types of entities in the second running information; generating a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information.
2. The method according to claim 1, wherein The generating a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information includes: training a decision tree model according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information to obtain a binary tree structure; generating a detection rule for detecting malicious files according to the binary tree structure.
3. The method according to claim 2, wherein The generating a detection rule for detecting malicious files according to the binary tree structure includes: traversing the binary tree structure in a pre-order traversal manner to generate a detection rule for detecting malicious files.
4. The method according to any one of claims 1 to 3, wherein, The running environment is loaded on an electronic device; The obtaining the running information of multiple files running in a running environment includes: at least obtaining the log information of the electronic device, where the log information of the electronic device includes: the log information generated by the electronic device during the process of running multiple files on the running environment of the electronic device; at least generating the running information of the multiple files running in the running environment according to the log information of the electronic device.
5. The method according to claim 4, wherein The at least generating the running information of the multiple files running in the running environment according to the log information of the electronic device includes: searching for multiple file names in the log information of the electronic device; for any one of the multiple file names found, extracting at least one preset type of entity having an associated relationship with the file name in the log information of the electronic device, and the at least one preset type of entity at least includes: the process name for processing the file corresponding to the file name, and at least generating the running information of the file corresponding to the file name running in the running environment according to the file name and the at least one preset type of entity.
6. The method according to any one of claims 1 to 5, wherein The running information of the files further includes: the association relationship between multiple preset types of entities involved in the process of the files running in the running environment; the determining the first running information and the second running information in the running information of the multiple files includes: 27 For the running information of any file, according to multiple preset types of entities in the running information of the file and the association relationships between the multiple preset types of entities, obtain the knowledge graph of the file; detect whether the file name in the knowledge graph is the file name of a malicious file; in the case where the file name in the knowledge graph is the file name of a malicious file, obtain first running information according to the knowledge graph; or, in the case where the file name in the knowledge graph is not the file name of a malicious file, obtain second running information according to the knowledge graph.
7. The method according to claim 6, wherein The detecting whether the file name in the knowledge graph is the file name of a malicious file includes: determining a process name and a file name in the knowledge graph; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; intercept a subgraph in the knowledge graph, the subgraph includes the determined process name and preset type entities within P hops after the determined process name, P is less than or equal to X, and X is the number of node hops between the determined process name and the most terminal preset type entity in the knowledge graph; input the subgraph into a trained graph neural network model so that the trained graph neural network model processes the subgraph to obtain a detection result of whether the determined file name in the subgraph is the file name of a malicious file.
8. The method according to claim 6, wherein The obtaining first running information according to the knowledge graph includes: determining the process name of a malicious process in the knowledge graph; in the knowledge graph, select preset type entities within N hops after the process name of the malicious process, N is less than or equal to Y, and Y is the number of node hops between the process name of the malicious process and the most terminal preset type entity in the knowledge graph; determine whether there is a file name of a malicious file among the preset type entities within N hops after the process name of the malicious process; in the case where there is a file name of a malicious file among the preset type entities within N hops after the process name of the malicious process, intercept a subgraph in the knowledge graph, the subgraph includes the process name of the malicious process and preset type entities within N hops after the process name of the malicious process, and the malicious process is used to process the malicious file; obtain first running information according to the subgraph.
9. The method according to claim 6, wherein The obtaining second running information according to the knowledge graph includes: determining a process name and a file name in the knowledge graph, and the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; Intercept a sub-graph in the knowledge graph, where the sub-graph includes a determined process name and preset type entities within M hops after the determined process name, and M is less than or equal to Z, where Z is the number of node hops between the determined process name and the preset type entity at the end in the knowledge graph; obtain second running information according to the sub-graph.
10. A method for detecting a file, wherein The method includes: obtaining the running information of the file to be detected, where the running information of the file to be detected includes: multiple preset type entities involved in the process of the file to be detected running in the running environment, and the multiple preset type entities at least include: the file name of the file to be detected and the process name for processing the file to be detected; obtaining the attribute information of the multiple preset type entities in the running information of the file to be detected; detecting whether the file to be detected is a malicious file according to the attribute information and the detection rules for detecting malicious files generated in advance.
11. The method according to claim 10, wherein The running information of the file to be detected further includes: the association relationship between multiple preset type entities involved in the process of the file to be detected running in the running environment; the obtaining of the attribute information of the multiple preset type entities in the running information of the file to be detected includes: according to the multiple preset type entities in the running information of the file to be detected and the association relationship between the multiple preset type entities, obtaining the knowledge graph of the file to be detected; determining the process name and the file name in the knowledge graph of the file to be detected; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; intercepting a sub-graph in the knowledge graph of the file to be detected, where the sub-graph includes the determined process name and preset type entities within Q hops after the determined process name, and Q is less than or equal to V, where V is the number of node hops between the determined process name and the preset type entity at the end in the knowledge graph; obtaining the attribute information of each preset type entity in the sub-graph.
12. The method according to claim 10 or 11, wherein There are multiple pieces of running information of the file to be detected; the file names in the multiple pieces of running information are all the file name of the file to be detected, and the process names in the multiple pieces of running information are different; correspondingly, the obtaining of the attribute information of the multiple preset type entities in the running information of the file to be detected includes: for any piece of running information of the file to be detected, obtaining the attribute information of the multiple preset type entities in the running information to obtain the attribute information corresponding to the running information; Accordingly, detecting whether a file to be detected is a malicious file according to the attribute information and the pre-generated detection rules for detecting malicious files includes: for the attribute information corresponding to any running information of the file to be detected, matching the attribute information corresponding to the running information with the detection rules to obtain a detection result indicating whether the file to be detected is a malicious file by the attribute information corresponding to the running information; counting the number of attribute information corresponding to the running information of the detection result indicating that the file to be detected is a malicious file; and determining that the file to be detected is a malicious file when the ratio between the number and the number of multiple running information of the file to be detected is greater than a preset threshold.
13. A device for establishing detection rules, wherein The device includes: a first acquisition module, configured to acquire the running information of multiple files running in a running environment, where the running information of the files includes: multiple preset types of entities involved in the process of the files running in the running environment, and the multiple preset types of entities at least include: the file name of the file and the process name for processing the file; a determination module, configured to determine first running information and second running information from the running information of the multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; a second acquisition module, configured to acquire the attribute information of the multiple preset types of entities in the first running information, and acquire the attribute information of the multiple preset types of entities in the second running information; and a generation module, configured to generate detection rules for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information.
14. A device for detecting a file, wherein, The device includes: a third acquisition module, configured to acquire the running information of the file to be detected, where the running information of the file to be detected includes: multiple preset types of entities involved in the process of the file to be detected running in the running environment, and the multiple preset types of entities at least include: the file name of the file to be detected and the process name for processing the file to be detected; a fourth acquisition module, configured to acquire the attribute information of the multiple preset types of entities in the running information of the file to be detected; and a detection module, configured to detect whether the file to be detected is a malicious file according to the attribute information and the pre-generated detection rules for detecting malicious files.
15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the method described in any one of claims 1 to 12 is implemented.
16. A computer-readable storage medium, wherein, A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the method described in any one of claims 1 to 12 is implemented.
17. A computer program product, wherein, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is enabled to execute the method described in any one of claims 1 to 12.
Citation Information
Patent Citations
Android malicious application detection method and system based on multi-feature fusion
CN107180192A
Executable file detection method and device, equipment and storage medium
CN116305113A
Information processing method and server, and computer storage medium
US20170372069A1
System and method for automated machine-learning, zero-day malware detection
US20210256127A1
Cited By
AI large model content generation security protection method and system
CN121457642A