Method for establishing detection rule and method and device for detecting file

By generating detection rules based on decision tree model in the cloud computing system, and using entity attribute information in the file running environment, the problems of low accuracy and low update efficiency in the existing technology are solved, efficient and low-cost malicious file detection are achieved, and system security is enhanced.

CN120337212APending Publication Date: 2025-07-18HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410069347.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When detecting malicious files in cloud computing systems, the prior art has problems such as low detection accuracy, low update efficiency and high labor costs. Especially when faced with encryption, shelling and obfuscated files, the detection model is prone to missing features, resulting in missed detection and missed detection, affecting system security.

Method used

By obtaining the attribute information of multiple preset type entities of the file in the running environment, a detection rule based on the decision tree model is generated, and sample data is trained using the binary tree structure to generate detection rules for detecting malicious files, improving detection accuracy and reducing labor costs.

Benefits of technology

It improves the accuracy of detecting malicious files, reduces the time to regenerate detection rules, reduces labor costs, enhances the security of cloud computing systems, and avoids the reduction in detection accuracy caused by updating rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337212A_ABST
    Figure CN120337212A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for establishing a detection rule and a method and a device for detecting a file. Running information of the to-be-detected file is obtained, the running information of the to-be-detected file comprises multiple preset types of entities involved in the running process of the to-be-detected file in the real running environment, and the multiple preset types of entities at least comprise a file name of the to-be-detected file and a process name used for processing the to-be-detected file; obtaining attribute information of a plurality of preset types of entities in the operation information of the to-be-detected file; and detecting whether the to-be-detected file is the malicious file or not according to the attribute information and a pre-generated detection rule for detecting the malicious file, so that the detection accuracy of detecting whether the to-be-detected file is the malicious file or not can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular, to a method and device for establishing a detection rule and a method and device for detecting a file. Background Art

[0002] In a cloud computing system, the security of files is becoming increasingly important. If there are malicious files in the cloud computing system, it will pose a great threat to the security of the cloud computing system. Therefore, it is necessary to perform security detection on the files in the cloud computing system. If a malicious file is detected in the cloud computing system, the malicious file can be deleted in the cloud computing system to improve the security of the cloud computing system. Summary of the Invention

[0003] This application discloses a method and device for establishing a detection rule and a method and device for detecting a file.

[0004] In a first aspect, a method for establishing a detection rule is disclosed, including: obtaining the running information of a plurality of files running in a running environment, where the running information of the files includes a plurality of preset types of entities involved in the process of the files running in the running environment, and the plurality of preset types of entities include the file names of the files and the process names for processing the files; determining first running information and second running information from the running information of the plurality of files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; obtaining the attribute information of the plurality of preset types of entities in the first running information, and obtaining the attribute information of the plurality of preset types of entities in the second running information; generating a detection rule for detecting malicious files according to the attribute information of the plurality of preset types of entities in the first running information and the attribute information of the plurality of preset types of entities in the second running information.

[0005] In a second aspect, a method for detecting a file is disclosed, including: obtaining the running information of a file to be detected, where the running information of the file to be detected includes a plurality of preset types of entities involved in the process of the file to be detected running in a running environment, and the plurality of preset types of entities at least include the file name of the file to be detected and the process name for processing the file to be detected; obtaining the attribute information of the plurality of preset types of entities in the running information of the file to be detected; detecting whether the file to be detected is a malicious file according to the attribute information and a previously generated detection rule for detecting malicious files.

[0006] In a third aspect, a device for establishing a detection rule is shown, including: a first acquisition module, configured to acquire the running information of multiple files running in a running environment, where the running information of the file includes multiple preset types of entities involved in the process of the file running in the running environment, and the multiple preset types of entities include the file name of the file and the process name for processing the file; a determination module, configured to determine first running information and second running information from the running information of the multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; a second acquisition module, configured to acquire the attribute information of the multiple preset types of entities in the first running information, and acquire the attribute information of the multiple preset types of entities in the second running information; a generation module, configured to generate a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information.

[0007] In a fourth aspect, a device for detecting a file is shown, including: a third acquisition module, configured to acquire the running information of a file to be detected, where the running information of the file to be detected includes: multiple preset types of entities involved in the process of the file to be detected running in a running environment, and the multiple preset types of entities at least include: the file name of the file to be detected and the process name for processing the file to be detected; a fourth acquisition module, configured to acquire the attribute information of the multiple preset types of entities in the running information of the file to be detected; a detection module, configured to detect whether the file to be detected is a malicious file according to the attribute information and a pre-generated detection rule for detecting malicious files.

[0008] In a fifth aspect, an electronic device is shown, where the electronic device includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method shown in any of the foregoing aspects.

[0009] In a sixth aspect, a non-transitory computer-readable storage medium is shown, where when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the method shown in any of the foregoing aspects.

[0010] In a seventh aspect, a computer program product is shown, where when the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is enabled to execute the method shown in any of the foregoing aspects.

[0011] Compared with the prior art, the present application has the following advantages:

[0012] In this application, if it is necessary to further improve the detection accuracy of detecting whether there is malicious behavior in the behavior of the file to be detected based on the detection rule in the future, since the detection rule in this application is automatically generated logically according to the relevant information of the actual operation of the file in the running environment, the detection rule is often a white box, and the detection rule has interpretability, that is, usually the reason for detecting whether there is malicious behavior in the behavior of the file to be detected based on the detection rule can be determined. For example, usually the reason for not detecting the malicious behavior of the malicious file based on the detection rule can be determined, etc. In this way, for the misdetection situation, the detection rule can be updated without regenerating the detection rule, but new detection rules can be added on the basis of the original detection rule, or some of the original detection rules can be modified.

[0013] On the one hand, since it is possible to "not regenerate the detection rule", new detection rules can be added on the basis of the original detection rule or the original detection rule can be directly modified, so that the time for "regenerating the detection rule" can be saved and the update efficiency can be improved. On the other hand, it is not necessary to manually collect data for generating the detection rule again, thereby reducing the labor cost. On the other hand, since it is not necessary to regenerate the detection rule, if new detection rules are added on the basis of the original detection rule, there is no change to the original detection rule, which does not affect the original detection rule and will not reduce the detection accuracy of the original detection rule. Or, if some of the original detection rules are modified, the detection accuracy of the modified detection rule corresponding to these detection rules can be improved. Secondly, for the detection rules in the original detection rule other than these detection rules (the unmodified detection rules), since there is no modification or change to them, there is no impact on them and their detection accuracy will not be reduced.

[0014] After using the detection rule to detect whether there is malicious behavior in the behavior of the file to be detected for a long time, sometimes it may be adversarial to the detection rule, and then the detection accuracy of the detection rule may be reduced. For example, an attacker can send a large number of malicious files to test the detection rule. If it is found that the malicious behavior of some types of malicious files is not detected as malicious behavior by the detection rule, the attacker can determine that the malicious behavior of these types of malicious files can bypass the detection of the detection rule, and then specifically construct a large number of these types of malicious files, and make these types of malicious files execute malicious behaviors to attack the cloud computing system, and can resist the detection rule so as not to be detected by the detection rule for their malicious behavior, resulting in a reduction in the detection accuracy of the detection rule, and then affecting the security of the cloud computing system.

[0015] Further, since the detection accuracy of detecting whether there is a malicious behavior in the file to be detected based on the detection rule is reduced, it is often necessary to improve the detection accuracy of detecting whether there is a malicious behavior in the file to be detected based on the detection rule. For example, updating the detection rule, etc. However, as described above, the present application can improve the update efficiency, reduce the labor cost and will not reduce the detection accuracy of the original detection rule, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of the steps of a method for establishing a detection rule according to the present application.

[0017] Figure 2 is a schematic diagram of a binary tree structure according to the present application.

[0018] Figure 3 is a flowchart of the steps of a method for obtaining operation information according to the present application.

[0019] Figure 4 is a flowchart of the steps of a method for determining operation information according to the present application.

[0020] Figure 5 is a schematic diagram of a knowledge graph according to the present application.

[0021] Figure 6 is a schematic diagram of a knowledge graph according to the present application.

[0022] Figure 7 is a flowchart of the steps of a method for detecting a file according to the present application.

[0023] Figure 8 is a block diagram of the structure of a device for establishing a detection rule according to the present application.

[0024] Figure 9 is a block diagram of the structure of a device for detecting a file according to the present application.

[0025] Figure 10 is a block diagram of the structure of a device according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] When detecting whether a file to be detected is a malicious file, in one approach, it is possible to detect whether the file to be detected is a malicious file based on the static features of the file to be detected. For example, the file to be detected can be disassembled to obtain some inherent features (such as inherent attributes, etc.) of the file to be detected, and these are used as the static features of the file to be detected. Then, it is detected whether the static features of the file to be detected have the inherent features of malicious files. If the static features of the file to be detected have the inherent features of malicious files, it can be determined that the file to be detected is a malicious file. Or, if the static features of the file to be detected do not have the inherent features of malicious files, it can be determined that the file to be detected is not a malicious file but a safe file.

[0028] Among them, a large number of technicians can be arranged in advance to count the inherent features of malicious files on the market, and specifically determine which inherent features are the features that malicious files often have, so as to develop a detection model for the inherent features that malicious files often have. In this way, when detecting whether the static features of the file to be detected have the inherent features of malicious files, the inherent features of the file to be detected can be matched with the detection model. If the inherent features of the file to be detected match the detection model, it can be determined that the static features of the file to be detected have the inherent features of malicious files. Or, if the inherent features of the file to be detected do not match the detection model, it can be determined that the static features of the file to be detected do not have the inherent features of malicious files.

[0029] However, the inventor has found that the above approach has the following defects:

[0030] 1. The process of "counting the inherent features of malicious files on the market, specifically determining which inherent features are the features that malicious files often have, and developing a detection model for the inherent features that malicious files often have" often requires the participation of a large number of technicians, resulting in high labor costs.

[0031] 2. In order to avoid malicious files being detected, developers of malicious files often hide and process the inherent features of malicious files through encryption, shelling, and / or obfuscation, etc. (for example, turning the inherent features into garbled characters or random symbols, etc.).

[0032] This may result in not being able to count all the inherent features of malicious files on the market when "counting the inherent features of malicious files on the market", leading to omissions in the counted inherent features of malicious files, and further leading to omissions when "specifically determining which inherent features are the features that malicious files often have", and further leading to the developed detection model missing at least some of the inherent features of malicious files on the market, that is, the developed detection model does not cover at least some of the inherent features of malicious files on the market, that is, the developed detection model is incomplete and has low generalization ability.

[0033] In addition, when the file to be detected is a malicious file and the inherent features of the file to be detected are hidden through encryption, shelling, and / or obfuscation, etc., it may lead to the inability to extract at least some of the inherent features of the file to be detected. In the case where at least some of the inherent features of the file to be detected cannot be extracted, it is very likely that the detection model cannot be hit based on the above-mentioned method, which may result in the file to be detected not being detected as a malicious file, that is, it leads to missed detection.

[0034] Furthermore, for the files encrypted, shelled, and / or obfuscated mentioned above, although some of the inherent features of the file cannot be extracted based on the above-mentioned method. However, the inventor found that the file needs to use its inherent features during the running process. Thus, the file will first run the self-decrypting code in the file during runtime to automatically decrypt, de-shell, and / or de-obfuscate, so as to obtain the inherent features of the file, so that the inherent features of the file can be used during the subsequent running process. In view of this, another method is proposed, which can detect whether the file to be detected is a malicious file based on the dynamic detection method in the actual scenario of running the file.

[0035] For example, the file to be detected is delivered to a sandbox environment, the file to be detected is run in the sandbox environment, and the behavior of the file to be detected in the sandbox environment (such as the actions executed by the file to be detected in the sandbox environment, etc.) is recorded. Then, it can be detected whether there is malicious behavior in the behavior of the file to be detected in the sandbox environment. If there is malicious behavior in the behavior of the file to be detected in the sandbox environment, it can be determined that the file to be detected is a malicious file. Or, if there is no malicious behavior in the behavior of the file to be detected in the sandbox environment, it can be determined that the file to be detected is not a malicious file but a safe file.

[0036] In one example, the method based on dynamic detection may include the detection method based on an intelligent algorithm model, etc. For example, it is detected whether the file to be detected is a malicious file based on an intelligent algorithm model.

[0037] For example, a file can be delivered to a sandbox environment and run in the sandbox environment. Based on the running report of the sandbox environment, the process call sequence of the file in the sandbox environment (including the behavior of the file, etc.) can be extracted, and the process call sequence of the file in the sandbox environment can be vectorized to obtain vectorized features, and an intelligent algorithm model can be trained based on the vectorized features. After that, when detecting whether there is malicious behavior in the behavior of the file to be detected in the sandbox environment, the file to be detected can be delivered to the sandbox environment and run in the sandbox environment, and the process call sequence of the file to be detected in the sandbox environment can be recorded. Then, the process call sequence of the file to be detected in the sandbox environment is input into the trained intelligent algorithm model, so that the trained intelligent algorithm model processes the process call sequence of the file to be detected in the sandbox environment to obtain the detection result of whether there is malicious behavior in the behavior of the file to be detected in the sandbox environment.

[0038] However, the inventors found that the detection accuracy of detecting whether there is malicious behavior in the behavior of the file to be detected by the above another method is low.

[0039] Among them, sometimes it may be considered that it is the reason of the intelligent algorithm model itself that leads to the low detection accuracy of detecting whether there is malicious behavior in the behavior of the file to be detected. Therefore, it may be necessary to improve the detection accuracy of detecting whether there is malicious behavior in the behavior of the file to be detected based on the intelligent algorithm model. For example, update the intelligent algorithm model, etc.

[0040] However, the intelligent algorithm model is often a black box. Therefore, the intelligent algorithm model does not have interpretability, that is, usually the reason for detecting whether there is malicious behavior in the behavior of the file to be detected based on the intelligent algorithm model cannot be determined. For example, usually the reason for not detecting the malicious behavior of a malicious file based on the intelligent algorithm model cannot be determined, etc. Thus, for the misdetection / missing detection situation, updating the intelligent algorithm model often can only re-manually collect appropriate training data according to the missing detection situation and re-train the intelligent algorithm model using the re-collected training data.

[0041] However, on the one hand, the process of "re-manually collecting appropriate training data and re-training the intelligent algorithm model" often takes a long time, resulting in low update efficiency. On the other hand, re-manually collecting appropriate training data requires a large number of technical personnel to participate, resulting in high labor costs. On the other hand, the intelligent algorithm model re-trained using the re-collected training data may also introduce new misdetection / missing detection and other situations, and thus it is very likely that the detection accuracy of the re-trained intelligent algorithm model will be reduced, etc.

[0042] In addition, after using the intelligent algorithm model to detect whether there is malicious behavior in the behavior of the file to be detected for a long time, sometimes it may be adversarial to the intelligent algorithm model, thereby resulting in a decrease in the detection accuracy of the intelligent algorithm model.

[0043] For example, an attacker can deliver a large number of malicious files to test the intelligent algorithm model. If it is found that the malicious behaviors of certain types of malicious files are not detected as malicious behaviors by the intelligent algorithm model, the attacker can determine that the malicious behaviors of these certain types of malicious files can bypass the detection of the intelligent algorithm model. Then, a large number of these certain types of malicious files can be constructed specifically, and these certain types of malicious files can be made to execute malicious behaviors to attack the cloud computing system, and can resist the intelligent algorithm model so as not to be detected by the intelligent algorithm model for their malicious behaviors, resulting in a decrease in the detection accuracy rate of the intelligent algorithm model, and further affecting the security of the cloud computing system.

[0044] Furthermore, due to the defect of the low detection accuracy rate in detecting whether the behavior of the file to be detected has malicious behavior based on the above-mentioned another method, sometimes it may be considered that it is the reason of the intelligent algorithm model itself that leads to the low detection accuracy rate in detecting whether the behavior of the file to be detected has malicious behavior. Therefore, it may be necessary to improve the detection accuracy rate of detecting whether the behavior of the file to be detected has malicious behavior based on the intelligent algorithm model. For example, update the intelligent algorithm model, etc. However, as mentioned above, the update efficiency is low, the update labor cost is high, and the detection accuracy rate may also decrease after the update, etc.

[0045] Therefore, in order to solve the above problems, the solution of the present application is proposed.

[0046] Among them, before introducing the solution of the present application, the professional terms that may be involved in the present application are first explained.

[0047] Knowledge graph: A knowledge base called a semantic network, that is, a knowledge base with a directed graph structure. The knowledge graph includes at least two entities (which can also be called nodes), and there may be association relationships between different entities.

[0048] Binary file: A computer file format in which data is stored in binary form.

[0049] MD5, Message-Digest Algorithm 5, a widely used cryptographic hash function that can generate a 128-bit (16-byte) hash value to ensure the integrity and consistency of information transmission. The MD5 of a file is the signature or identifier of the file, and the MD5s of different files are different.

[0050] Sandbox: Simulates the host environment, has isolation security, allows malware to run inside, and records the running behavior information.

[0051] The decision tree model is a decision analysis method that, based on the known probabilities of various situations, evaluates project risks and determines its feasibility by constructing a decision tree to obtain the probability that the expected value of the net present value is greater than or equal to zero. A decision tree is a tree-like structure. Its decision branches are drawn in a graph that resembles the branches of a tree, so it is called a decision tree. A decision tree consists of a root node, internal nodes, and leaf nodes. Each decision tree has only one root node. Each internal node represents a test on an attribute. Each branch represents a test output. Each leaf node represents a category. The generation of a decision tree generally starts from the root node, selects the corresponding attribute, then selects the split point of the attribute corresponding to the node, and then splits the node according to the split point. The decision tree generates multiple child nodes by selecting features and corresponding split points. When the values in a certain node belong to only one category (or the variance is small), the child nodes are no longer further split.

[0052] Among them, referring to Figure 1 , a method for establishing a detection rule of the present application is shown. This method is applied to an electronic device in a cloud computing scenario. The electronic device in the cloud computing scenario may include: a physical device or a virtual device. The physical device includes a server or a terminal, etc. The virtual device may include a virtual machine installed on the physical device, etc. The virtual machine may include an ECS (Elastic Cloud Server), etc. Among them, this method may include:

[0053] In step S101, obtain the running information of multiple files running in the running environment; the running information of the files includes: multiple preset types of entities involved in the process of the files running in the running environment; the multiple preset types of entities at least include: the file name of the file and the process name used to process the file.

[0054] Furthermore, in one embodiment, in addition to including the process name of the process used to process the file and the file name of the file, the multiple preset types of entities may further include: the path of the file, the MD5 of the file (the MD5 of the file can be used as an identifier of the file, etc. The file names of different files may be the same, but the MD5s of different files are different. The MD5 of the file is used to uniquely identify the file), the identifier of the behavior executed by the file, the external port connected by the process, the external IP (Internet Protocol Address) address connected by the process, and the event name of the event created by the process (the event may include a login event, a payment event, or a call event, etc.).

[0055] The path of the file may include the storage path of the file in the electronic device. The identifier of the behavior executed by the file may include "reading file information", "accessing the registry", "invoking system functions", etc., and no further examples will be given here. The identifier of the behavior is used to indicate the type of the behavior, etc., and is used to uniquely identify the type of the behavior. The behavior executed by the file is a dynamic action, and the behavior executed by the file can be quantified through the identifier of the behavior. The identifier of the behavior executed by the file can be obtained according to the ATTCK matrix pre-written by the technician. The ATTCK matrix includes the mapping relationship between the description information of the behavior and the identifier of the behavior. In this way, the description information of the behavior executed by the file (such as obtained according to the behavior log, etc.) can be obtained, and then the identifier of the behavior executed by the file can be indexed in the ATTCK matrix according to the description information of the behavior executed by the file.

[0056] In this application, the running environment is loaded on the electronic device, and the file runs in the running environment. Thus, the specific running information of the file in the running environment is often stored in the relevant log information of the electronic device. For example, during the process of the file running in the running environment, the electronic device will collect the specific running information of the file in the running environment in real time and store it in the relevant log information of the electronic device. In this way, the running information of multiple files running in the running environment can be obtained according to the relevant log information of the electronic device. The files in this application may include binary files, etc.

[0057] Among them, for this step, reference can be specifically made to the embodiments shown later, and details will not be described here.

[0058] In step S102, determine the first running information and the second running information from the running information of multiple files; the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file.

[0059] The purpose of this step is to classify the running information of multiple files into two categories, one is the first running information and the other is the second running information. The file names in the first running information are entities of a preset type, and the file names in the second running information are entities of a preset type. In one embodiment, for the running information of any file, at least the file name of the file among the multiple entities of the preset type in the running information of the file may be the file name of a malicious file or a non-malicious file. If the file is a malicious file, the file name of the file is the file name of the malicious file, and then the first running information can be obtained according to the running information of the file. For example, the running information of the file can be used as the first running information, etc. Or, if the file is not a malicious file, the file name of the file is not the file name of the malicious file, and then the second running information can be obtained according to the running information of the file. For example, the running information of the file can be used as the second running information, etc. A malicious file is a file that poses a security risk to the cloud computing system. For example, malicious files include files that perform malicious behaviors, and the malicious behaviors performed by malicious files will pose a security risk to the cloud computing system. Or, in another embodiment, the first running information can be determined first from the running information of multiple files, and then the running information other than the first running information in the running information of multiple files can be used as the second running information.

[0060] Specifically, this step can be referred to the embodiments shown later and will not be elaborated here.

[0061] In step S103, obtain the attribute information of multiple entities of the preset type in the first running information, and obtain the attribute information of multiple entities of the preset type in the second running information.

[0062] In one embodiment, for any first running information, the entities of the preset type in the first running information may include: the file name of the file, the process name of the process for processing the file, the path of the file, the MD5 of the file, the identifier of the behavior executed by the file, the external port connected by the process, the external IP address connected by the process, and the event name of the event created by the process, etc.

[0063] For the entity "file name", the attribute information of "file name" may be the characters in the "file name". For example, if a file is "love.exe", where "love" is the file name, the file name "love" is an entity of the preset type, and the attribute information of the file name "love" may be "love".

[0064] For the entity "process name", keywords or commands can be set for the process in advance. For example, the keywords "keyword" can include: keyword = 'pty','spawn', 'pty.spawn', 'password','shadow', and 'passwd', etc. For another example, the commands "command" can include: 'cat', 'chmod', 'cp', 'curl','sqlite3', 'postgres', 'runuser', and 'ldapsearch', etc.

[0065] For the entity "process name", if the cmdline command line information of the process corresponding to the "process name" contains the command 'cat', then set the label of the "process name" as command_cat. If the cmdline command line information of the process corresponding to the "process name" contains the keyword password, then set the label of the "process name" as keyword_password. All the labels of the "process name" are used as the attribute information of the "process name".

[0066] For the entity "file path", the file path can be labeled. The purpose of labeling the file path is to map various paths to some fixed labels. For example, multiple different preset paths can be set in advance, and the keywords in different preset paths are not all the same. For the "file path", it can be determined which keyword of the preset path is included in the characters of the file path, and then the label of the file path can be obtained according to the determined keyword. The label of the file path is used as the attribute information of the file path. For example, the keywords of multiple preset paths can include: [' / dev / mem', ' / dev / tcp', ' / dev / udp', ' / dev / kmem', ' / dev / null', ' / inet / tcp']. Suppose the file path is ' / dev / mem / temp'. Since the file path ' / dev / mem / temp' contains the keyword " / dev / mem", the label of the file path ' / dev / mem / temp' can be 'path / dev / mem', and the format of the label can be path_{}, where {} is the keyword.

[0067] For the entity "file MD5", the attribute information of the "file MD5" can be the characters in the "MD5". For example, if the MD5 of a file is "XXXX...XXXX", where "XXXX...XXXX" is the MD5, and the file MD5 "XXXX...XXXX" is an entity of a preset type, then the attribute information of the file MD5 "XXXX...XXXX" can be "XXXX...XXXX".

[0068] For the entity "identifier of the behavior performed by a file", the attribute information of the "identifier of the behavior performed by a file" can be the characters in the "identifier". For example, if a file performs the behavior of accessing the registry, then the identifier of the behavior performed by the file is "access the registry", where "access the registry" is the identifier. The identifier of the behavior performed by the file, "access the registry", is an entity of a preset type, and the attribute information of the identifier of the behavior performed by the file, "access the registry", can be "access the registry".

[0069] For the entity "external port", the "external port" is an actual port. It is possible to determine the port segment in which the external port is located among multiple preset different port segments. Different port segments have different labels. It is possible to obtain the label of the determined port segment and then determine the label of the determined port segment as the attribute information of the external port.

[0070] For example, ports can be divided into three major categories. The first category: Well-known ports: (0 to 1023) are tightly bound to some services. Generally, the communication using these ports indicates the protocol of a certain service. (For example, port 80 for HTTP communication, port 21 is assigned to the FTP service, port 25 is assigned to the SMTP (Simple Mail Transfer Protocol) service, port 135 is assigned to the RPC (Remote Procedure Call) service, etc.). The second category: Registered ports (1024 to 49151): are loosely bound to some services. Many services are bound to these ports, and these ports are also used for many other purposes. (The system starts allocating dynamic ports from 1024). The third category: Dynamic / or private ports (49152 to 65535) generally should not be assigned to services. In fact, machines usually start allocating dynamic ports from 1024. However, there are exceptions: SUN's RPC ports start from 32768.

[0071] For the entity "external IP address", the "external IP address" is an actual IP address. It is possible to determine the address segment in which the external IP address is located among multiple preset different address segments. Different address segments have different labels. It is possible to obtain the label of the determined address segment and then determine the label of the determined address segment as the attribute information of the external IP address. For example, the label of one address segment is link_local (link-local address), the label of one address segment is loopback (loopback address), the label of one address segment is multicast (multicast), the label of one address segment is reserved (reserved), the label of one address segment is unspecified (unspecified), the label of one address segment is private (private), and the label of one address segment is public (public).

[0072] For the entity "event name", the attribute information of the "event name" can be the characters in the "event name". For example, an event is a login event, the event name of the login event is "LOGIN_EVENT", the event name "LOGIN_EVENT" is an entity of a preset type, and the attribute information of the event name "LOGIN_EVENT" can be "LOGIN_EVENT".

[0073] The same applies to each of the other first run information and second run information, which will not be elaborated here. Thus, the attribute information of multiple entities of the preset type in each first run information and the attribute information of multiple entities of the preset type in each second run information are obtained respectively.

[0074] In step S104, a detection rule for detecting malicious files is generated based on the attribute information of multiple entities of the preset type in the first run information and the attribute information of multiple entities of the preset type in the second run information.

[0075] In one embodiment, this step can be implemented through the following process, including:

[0076] 1041. Train the decision tree model based on the attribute information of multiple entities of the preset type in the first run information and the attribute information of multiple entities of the preset type in the second run information to obtain a binary tree structure.

[0077] For any first run information, since the file name in the multiple entities of the preset type in this first run information is the file name of a malicious file, negative sample training data can be obtained according to the attribute information of each entity of the preset type in this first run information. For example, the attribute information of each entity of the preset type in this first run information is combined into a sequence in a specific order and used as negative sample training data. The specific order can be the order between the entities of the preset type, etc. The specific order can be set in advance, for example, specified in advance by technicians, etc. The same applies to each of the other first run information. Thus, several negative sample training data can be obtained for as many first run information as there are. For any second run information, since the file name in the multiple entities of the preset type in this second run information is not the file name of a malicious file, positive sample training data can be obtained according to the attribute information of each entity of the preset type in this second run information. For example, the attribute information of each entity of the preset type in this second run information is combined into a sequence in a specific order and used as positive sample training data. The same applies to each of the other second run information. Thus, several positive sample training data can be obtained for as many second run information as there are. Then, the decision tree model can be trained based on the positive sample training data and the negative sample training data to obtain a binary tree structure.

[0078] Each node in the binary tree structure represents an attribute information respectively, and different nodes represent different attribute information. Any attribute information is located in the negative sample training data, or in the positive sample training data, or in both the negative sample training data and the positive sample training data.

[0079] In an example, the obtained binary tree structure for training can be referred to Figure 2 as shown. Figure 2 It includes nodes A to K. Nodes A to K represent an attribute information respectively, and nodes A to K represent different attribute information respectively. In the binary tree structure, the leaf nodes will be attached with the result of whether they hit the negative sample training data and / or the result of whether they hit the positive sample training data.

[0080] In Figure 2 it, node A is connected to node B, node A is connected to node C, node B is connected to node D, node B is connected to node E, node D is connected to node G, node D is connected to node H, node H is connected to node K, node E is connected to node I, node E is connected to node J, and node C is connected to node F. Node A is the root node, and nodes G, K, I, J, and F are leaf nodes. The leaf node G only hits all the negative sample training data, the leaf node K only hits all the negative sample training data, the leaf node I only hits all the negative sample training data, the leaf node J hits both the negative sample training data and the positive sample training data, and the node F only hits all the positive sample training data.

[0081] 1042. Generate a detection rule for detecting malicious files according to the binary tree structure.

[0082] For example, in an embodiment, the binary tree structure can be traversed in a preorder traversal manner to generate a detection rule for detecting malicious files. For example, starting from the leaf nodes of the binary tree, the binary tree structure is traversed in a preorder traversal manner to generate a detection rule for detecting malicious files. The detection rule has a logical expression.

[0083] The detection rule has attribute information, and the attribute information is at least part of the attribute information of multiple preset types of entities in the first running information and the attribute information of multiple preset types of entities in the second running information. For example, it is part of or all of the attribute information of multiple preset types of entities in the first running information and the attribute information of multiple preset types of entities in the second running information. At least part of the attribute information is connected by logical symbols, and the logical symbols include "logical AND", "logical OR", and "logical NOT", etc.

[0084] In one embodiment, "logical AND" or "logical NOT" logical symbols can be respectively set for each node in the binary tree structure to obtain a binary tree structure with logical symbols. For example, for any node in the binary tree structure, if the node is the left branch of its parent node, it means that the node representing the left branch does not contain the attribute information represented by the node. Therefore, the logical symbol "logical NOT" can be added to the left side of the node. Or, if the node is the right branch of its parent node, it means that the node representing the right branch contains the attribute information represented by the node. Therefore, the logical symbol "logical AND" can be added to the left side of the node. The same applies to each of the other nodes in the binary tree structure, thus obtaining a binary tree structure with logical symbols. All nodes other than the non-root nodes in the binary tree structure with logical symbols have the logical symbols "logical AND" or "logical NOT".

[0085] Or, in another embodiment, before respectively setting the "logical AND" or "logical NOT" logical symbols for each node in the binary tree structure, the logical symbol "left any match" can be respectively set for each node in the binary tree structure. For example, the logical symbol "left any match" can be added to the left side of each node to improve the generalization of the subsequently generated detection rules. Then, the "logical AND" or "logical NOT" logical symbols are respectively set for each node in the binary tree structure that has been set with the logical symbol "left any match". For example, the "logical AND" or "logical NOT" logical symbols are added to the left side of each node in the binary tree structure that has been set with the logical symbol "left any match" to obtain a binary tree structure with logical symbols.

[0086] After that, among all the leaf nodes in the binary tree structure with logical symbols, the leaf nodes that only hit all the negative sample training data can be selected. Then, starting from any one of the selected leaf nodes, the binary tree structure is traversed in pre-order (recursively traversing the left node first and then the right node). During the traversal process, a logical expression is continuously accumulated until after traversing the root node and all the selected leaf nodes, the final detection rule is generated.

[0087] Among them, when establishing a logical expression between a node and its parent node, a logical expression connected by the logical symbol "logical AND" between the node and its parent node can be established. And, when establishing a logical expression between a node and its sibling node (the node and its sibling node have the same parent node), a logical expression connected by the logical symbol "logical OR" between the node and its sibling node can be established. After traversing the root node and all the selected leaf nodes from the leaf nodes, a logical expression containing the logical symbols "logical AND", "logical NOT", and "logical OR" is generated and used as the detection rule.

[0088] For example, logical symbols of "logical AND" or "logical NOT" are respectively set for each node in the binary tree structure in the foregoing manner, and a binary tree structure with logical symbols is obtained. Other nodes except the non-root nodes in the binary tree structure with logical symbols all have logical symbols of "logical AND" or "logical NOT".

[0089] The layer where the root node in the binary tree structure with logical symbols is located is the first layer, the Nth layer is the layer where the node farthest from the root node is located, and the leaf nodes may be located in the Nth layer, may be located in the N - 1th layer, may be located in the N - 2th layer, etc., where N is greater than or equal to 2 or 3, etc.

[0090] For example, one of the leaf nodes in the binary tree structure with logical symbols is N - 1 - A (N - 1 is the layer number where the leaf node is located, A is a number of the leaf node in the N - 1th layer, and the numbers of different nodes in the same layer are different), and the leaf node N - 1 - A is the leaf node that only hits all the negative sample training data.

[0091] In this way, traversal can be performed from the leaf node to the root node direction. For example, traversing to the parent node N - 2 - A of the leaf node N - A - 1 (N - 2 is the layer number where the node N - 2 - A is located, A is a number of the node N - 2 - A in the N - 2th layer).

[0092] In the present application, the relationship between two nodes with a parent - child relationship can be a "logical AND" relationship.

[0093] In this way, a logical expression 1 connected by the logical symbol "logical AND" can be established between the leaf node N - 1 - A and the parent node N - 2 - A of the leaf node N - 1 - A: [(leaf node N - 1 - A) && (node N - 2 - A)]. "&&" represents the logical symbol "logical AND".

[0094] In the branch of the child nodes of the node N - 2 - A, if in addition to the branch of the leaf node N - 1 - A, there are other branches and the leaf nodes in the other branches do not have leaf nodes that only hit all the negative sample training data, or there are no other branches, then traversal can continue from the node N - 2 - A in the direction of the root node.

[0095] Or, in the branch of the child nodes of the node N - 2 - A, if in addition to the branch of the leaf node N - 1 - A, there are other branches, and the leaf nodes in the other branches have leaf nodes that only hit all the negative sample training data, then a logical expression can be established between the node N - 2 - A and the nodes in the other branches. Then traversal continues from the node N - 2 - A in the direction of the root node.

[0096] For example, nodes in other branches include node N-1-B (N-1 is the layer number where leaf node N-1-B is located, and B is a number of leaf node N-1-B in the N-1th layer) and leaf node N-A (N is the layer number where leaf node N-A is located, and A is a number of leaf node N-A in the Nth layer). Leaf node N-A is a leaf node that only hits all negative sample training data. Node N-1-B is the parent node of leaf node N-A, and node N-1-B is the child node of node N-2-A.

[0097] In this way, a logical expression 2 connected by the logical symbol "logical AND" can be established among node N-2-A, node N-1-B, and node N-A: [(node N-2-A) && (node N-1-B) && (node N-A)].

[0098] Among them, there is a sibling relationship between leaf node N-1-A and node N-1-B. The relationship between two nodes with a sibling relationship can be a "logical OR" relationship.

[0099] After logical expressions are respectively established for all branches of node N-2-A that have leaf nodes that only hit all negative sample training data, the logical expressions established for all branches of node N-2-A that have leaf nodes that only hit all negative sample training data can be connected by the logical symbol "logical OR" to obtain logical expression N-2-A.

[0100] For example, from the perspective of node N-2-A, a logical expression N-2-A connected by the logical symbol "logical OR" can be established between logical expression 1 and logical expression 2 (the merged logical expression can be named after the highest-level node in the merged logical expression): [(logical expression 1) || (logical expression 2)].

[0101] In an example, logical expression N-2-A can be: [(leaf node N-1-A) && (node N-2-A)] || [(node N-2-A) && (node N-1-B) && (node N-A)]. "||" represents the logical symbol "logical OR".

[0102] After that, traversal can continue from node N-2-A in the direction towards the root node. For example, traversing to the parent node N-3-A of node N-2-A (N-3 is the layer number where node N-3-A is located, and A is a number of node N-3-A in the N-3th layer).

[0103] In this application, the relationship between two nodes with a parent-child relationship can be a "logical AND" relationship.

[0104] In this way, a logical expression 3 connected by the logical symbol "logical AND" can be established between the logical expression N-2-A and the parent node N-3-A of the node N-2-A: [(logical expression N-2-A) && (node N-3-A)].

[0105] In the branches of the child nodes of the node N-3-A, if there are other branches in addition to the branch of the leaf node N-2-A and the leaf nodes in the other branches do not have leaf nodes that only hit all the negative sample training data, or there are no other branches, then the traversal can continue from the node N-3-A in the direction of the root node.

[0106] Alternatively, in the branches of the child nodes of the node N-3-A, if there are other branches in addition to the branch of the leaf node N-2-A and the leaf nodes in the other branches have leaf nodes that only hit all the negative sample training data, then a logical expression can be established between the node N-3-A and the nodes in the other branches. Then, the traversal continues from the node N-3-A in the direction of the root node.

[0107] For example, the nodes in the other branches include the node N-2-B (N-2 is the layer number where the leaf node N-2-B is located, and B is a number of the leaf node N-2-B in the N-2nd layer) and the leaf node N-1-C (N-1 is the layer number where the leaf node N-1-C is located, and C is a number of the leaf node N-1-C in the N-1st layer). The leaf node N-1-C is a leaf node that only hits all the negative sample training data. The node N-2-B is the parent node of the leaf node N-1-C, and the node N-2-B is a child node of the node N-3-A.

[0108] In this way, a logical expression 4 connected by the logical symbol "logical AND" can be established among the node N-3-A, the node N-2-B, and the node N-1-C: [(node N-3-A) && (node N-2-B) && (node N-1-C)].

[0109] Among them, the leaf nodes N-2-A and N-2-B are sibling nodes. The relationship between two sibling nodes can be a "logical OR" relationship.

[0110] After logical expressions are respectively established for all the branches of the leaf nodes that the node N-3-A has and that only hit all the negative sample training data, the logical expressions established for all the branches of the leaf nodes that the node N-3-A has and that only hit all the negative sample training data can be connected by the logical symbol "logical OR" to obtain the logical expression N-3-A.

[0111] For example, from the perspective of node N-3-A, a logical expression N-3-A (the merged logical expression is named after the top-level node in the merged logical expression) connected by the logical symbol "logical or" between logical expression 3 and logical expression 4 can be established: [(logical expression 3) || (logical expression 4)].

[0112] In one example, the logical expression N-3-A can be: [(logical expression N-2-A) && (node N-3-A)] || [(node N-3-A) && (node N-2-B) && (node N-1-C)].

[0113] And so on, until the root node in the binary tree structure with logical symbols is traversed, and until all branches of the leaf nodes related to the root node that only hit all negative sample training data are included in the logical expression, then the finally obtained logical expression can be used as the detection rule.

[0114] Sometimes, the binary tree structure is complex, resulting in a large number of nodes (attribute information) in the detection rule generated in the above manner, which in turn makes the generated detection rule complex and inconvenient for subsequent updating of the detection rule. Thus, in another embodiment of the present application, the generated detection rule can be split into multiple sub-detection rules, and the number of nodes (attribute information) in each sub-detection rule is less than or equal to a preset value. The preset value can include 4, 5, or 6, etc., and can be determined according to the actual situation, and the present application does not limit this. Or, during the process of generating the detection rule, when performing a pre-order recursive traversal of the binary tree structure with logical symbols, if a parent node is encountered, it is stored, and when a sibling node is encountered, it can first be determined whether the number of nodes in the generated detection rule is greater than or equal to the preset value. If it is greater than or equal to the preset value, then when the sibling node is encountered, it is divided into multiple branches and traversed pre-order upward simultaneously. If it does not exceed, it is connected to the branch of the sibling node through the logical symbol "logical or".

[0115] In another embodiment of the present application, the operating environment is loaded on an electronic device. Refer to Figure 3 , step S101 includes:

[0116] In step S201, at least the log information of the electronic device is obtained. The log information of the electronic device includes: the log information generated by the electronic device during the process of running multiple files on the operating environment of the electronic device.

[0117] The log information of the electronic device includes at least one of the following: the log information of the processes of the electronic device, the log information of the network of the electronic device, the log information of the system of the electronic device, and the log information of the storage of the electronic device, etc. Of course, it can be understood that according to the actual situation, the log information of the electronic device can also include other types of log information of the electronic device, which is not limited in this application.

[0118] In one embodiment, when a file is run in the operating environment of the electronic device, the file is often run through the processes in the electronic device. The electronic device will automatically obtain the specific relevant logs of the processes in the electronic device and record them in the log information of the processes in the electronic device. In this way, the log information of the processes in the electronic device can be obtained, which is convenient for obtaining the running information of multiple files running in the operating environment subsequently.

[0119] In another embodiment, when a file is run in the operating environment (such as a real operating environment, etc.) of the electronic device, in some scenarios, the electronic device will perform network interactions with the outside for the file. The electronic device will automatically obtain the specific relevant information of the network interactions with the outside and record them in the log information of the network of the electronic device. In this way, the log information of the network of the electronic device can be obtained, which is convenient for obtaining the running information of multiple files running in the operating environment subsequently.

[0120] In yet another embodiment, when a file is run in the operating environment of the electronic device, the operating environment will also automatically obtain the specific relevant logs of the system of the electronic device (including the relevant information of running the file in the operating environment of the electronic device) and record them in the log information of the system (operating environment) of the electronic device. In this way, the log information of the system of the electronic device can be obtained, which is convenient for obtaining the running information of multiple files running in the operating environment subsequently.

[0121] Furthermore, intelligence information can also be obtained. The intelligence information involves malicious files statistically external and other malicious entities having an associated relationship with the malicious files, etc., which is convenient for obtaining the running information of multiple files running in the operating environment subsequently by combining with the log information.

[0122] Furthermore, for the entity "process name", the cmdline command line information of the process corresponding to the "process name" can be extracted, and implicit relationship mining is performed on the command line to extract more entities related to the process corresponding to the "process name" from the cmdline command line information, such as more IP addresses connected by the process, file names of more files processed by the process, paths of more files processed by the process, system commands called by the process, and preset types of entities such as keywords. So as to obtain the running information of multiple files running in the operating environment subsequently by combining with the log information and / or the intelligence information.

[0123] In step S202, generate the running information of multiple files running in the running environment at least according to the log information of the electronic device.

[0124] In one embodiment, this step can be implemented through the following process, including:

[0125] 2021. Search for multiple file names in the log information of the electronic device.

[0126] The multiple file names found in the log information of the electronic device can all be regarded as entities of a preset type.

[0127] In one embodiment, the log information of the electronic device involves the relevant information of multiple files that have run in the running environment. The file names can be searched in the log information of the electronic device by means of keywords and regarded as entities of a preset type, or the file names can be searched in the log information of the electronic device by means of semantic analysis and regarded as entities of a preset type. This application does not limit the way of searching for file names in the log information.

[0128] Among them, the file names can be searched separately in each log information such as the log information of the processes of the electronic device, the log information of the network of the electronic device, and the log information of the system of the electronic device, and regarded as entities of a preset type.

[0129] Further, for any one of the multiple file names found, the following processes 2022-2023 can be executed.

[0130] 2022. Extract at least one entity of a preset type having an association relationship with the file name from the log information of the electronic device. The at least one entity of a preset type at least includes: the process name of the process that processes the file corresponding to the file name.

[0131] The process names of different processes in the electronic device are different. Although the file name also belongs to the entity of a preset type, the at least one entity of a preset type having an association relationship with the file name may not include the file name. The preset type of the entity to be searched can be statistically set by technicians in advance, that is, which entities are entities of a preset type can be statistically set by technicians in advance.

[0132] Further, at least one entity of a preset type may at least include at least one of the following: the path of a file, the MD5 of a file, the identifier of the behavior executed by the file, the external port to which a process is connected, the external IP address to which a process is connected, and the event name of the event created by a process, etc. Among them, the association relationship may be reflected as follows: the file corresponding to the file name has the path of the file corresponding to the file name, the file corresponding to the file name has the MD5 of the file corresponding to the file name, the MD5 of the file corresponding to the file name has the identifier of the behavior executed by the file corresponding to the file name, the process corresponding to the process name has the external port to which the process corresponding to the process name is connected, the process corresponding to the process name has the external IP address to which the process corresponding to the process name is connected, and, the process corresponding to the process name has the event name of the event created by the process corresponding to the process name, etc.

[0133] In the present application, the association relationship between entities of a preset type can be reflected in the log information of the electronic device. At least one entity of a preset type associated with the file name can be found in the log information of the electronic device by means of keywords, or at least one entity of a preset type associated with the file name can be found in the log information of the electronic device by means of semantic analysis. At least one entity of a preset type at least includes the process name of the process that processes the file corresponding to the file name. The present application does not limit the method of finding at least one entity of a preset type associated with the file name in the log information.

[0134] 2023. At least based on the file name and at least one entity of a preset type, generate the running information of the file corresponding to the file name running in the running environment.

[0135] In one example, the file name and at least one entity of a preset type can be combined to obtain the running information of the file corresponding to the file name running in the running environment.

[0136] In another embodiment, the running information of the file corresponding to the file name running in the running environment can be generated according to the file name, at least one entity of a preset type, and the association relationship between the file name and at least one entity of a preset type. In one example, the file name, at least one entity of a preset type, and the association relationship between the file and at least one entity of a preset type can be combined to obtain the running information of the file corresponding to the file name running in the running environment. If there are more than two entities of a preset type, the combined content may further include: the association relationship between more than two entities of a preset type.

[0137] In one embodiment, for a computer, the entities in the file running information can be stored in the form of a data structure, etc., and the association relationship between the entities in the file running information can be stored in the form of a data structure, etc.

[0138] Alternatively, in another embodiment, if it is necessary to display the running information of a file, a knowledge graph can be generated. The knowledge graph includes nodes and relationships. The nodes are entities in the running information of the file, and the relationships are association relationships between the entities in the running information of the file for people to view.

[0139] In this application, when determining the first running information and the second running information from the running information of multiple files in step S102, in another embodiment of this application, for the running information of any file, the file name can be searched for among multiple preset types of entities in the running information of the file, and it is determined whether the found file name is the file name of a malicious file. If the found file name is the file name of a malicious file, the first running information can be obtained according to the running information of the file. For example, the running information of the file is determined as the first running information. Alternatively, if the found file name is not the file name of a malicious file, the second running information can be obtained according to the running information of the file. For example, the running information of the file is determined as the second running information. The same applies to the running information of each other file.

[0140] Among them, in some embodiments, the running information of the file may further include: the association relationship between multiple preset types of entities involved in the process of the file running in the running environment. For example, the association relationship between the process name and the file name includes: the process corresponding to the process name is used to process the file corresponding to the file name.

[0141] During the process of the file running in the running environment, it is the process in the electronic device that processes (loads / starts / deletes / sends / modifies, etc.) the file. Thus, the file has a file name, the process has a process name, and the association relationship between the process name and the file name can include: the process corresponding to the process name is used to process the file corresponding to the file name. Thus, the association relationship between entities can include: one entity owns another entity, or, one entity performs an action on another entity (such as one entity processes another entity, etc.), or, one entity belongs to another entity, etc. This application does not limit the types of association relationships between entities.

[0142] When the multiple entities of the preset types further include the path of a file, the MD5 of the file, the identifier of the behavior executed by the file, the external port to which the process connects, the external IP address to which the process connects, and the event name of the event created by the process, the association relationships between these multiple entities of the preset types may include at least one of the following: The file corresponding to the file name has the path of the file corresponding to the file name, the file corresponding to the file name has the MD5 of the file corresponding to the file name, the MD5 of the file corresponding to the file name has the identifier of the behavior executed by the file corresponding to the file name, the process corresponding to the process name has the external port to which the process corresponding to the process name connects, the process corresponding to the process name has the external IP address to which the process corresponding to the process name connects, and the process corresponding to the process name has created the event name of the event corresponding to the process name, etc.

[0143] Specifically, referring to Figure 4 , step S102 includes:

[0144] In step S301, for the running information of any file, according to the multiple entities of the preset types and the association relationships between the multiple entities of the preset types in the running information of the file, the knowledge graph of the file is obtained.

[0145] For example, assume that the multiple entities in the running information of the file corresponding to the file name include: the file name, the path of the file corresponding to the file name, the MD5 of the file corresponding to the file name, the identifier of the behavior executed by the file corresponding to the file name, the process name of the process that processes the file corresponding to the file name, the external port to which the process that processes the file corresponding to the file name connects, the external IP address to which the process that processes the file corresponding to the file name connects, and the login event created by the process that processes the file corresponding to the file name, etc. The association relationships between these multiple entities of the preset types may include: The file corresponding to the file name has the path of the file corresponding to the file name, the file corresponding to the file name has the MD5 of the file corresponding to the file name, the MD5 of the file corresponding to the file name has the identifier of the behavior executed by the file corresponding to the file name, the process corresponding to the process name has the external port to which the process corresponding to the process name connects, the process corresponding to the process name has the external IP address to which the process corresponding to the process name connects, and the process corresponding to the process name has created the event name of the event corresponding to the process name, etc. Thus, the knowledge graph of the file corresponding to the file name generated can be referred to Figure 5 as shown.

[0146] In step S302, it is detected whether the file name in the knowledge graph of the file is the file name of a malicious file.

[0147] The file name in the knowledge graph of the file is an entity of the preset type.

[0148] In one embodiment, the knowledge graph of the file can be traversed. For example, based on the association relationships between multiple preset types of entities in the running information of the file, the entity serving as the starting node in the knowledge graph of the file can be determined, and then the knowledge graph of the file can be traversed starting from the entity serving as the starting node.

[0149] In this application, the knowledge spectrum diagram shown can be improved with the help of external security detection results. Figure 5 The external security detection results can be the detection results manually detected by technicians, or can be the detection results detected using other automated detection tools, etc., and this application does not limit this. The external security detection results can indicate which entities are malicious entities.

[0150] For example, assume that the external security detection results indicate that Figure 5 the file name in the knowledge graph shown is the file name of a malicious file, indicating that Figure 5 the process name in the knowledge graph shown is the process name of a malicious process, indicating that Figure 5 the login event in the knowledge graph shown is a malicious event. Then, a malicious result ABNORMAL (as a node) can be added to the Figure 5 knowledge spectrum diagram shown. The association relationship between the process name and the malicious result ABNORMAL is that the process corresponding to the process name has the malicious result ABNORMAL. The association relationship between the file name and the malicious result ABNORMAL is that the file corresponding to the file name has the malicious result ABNORMAL. The association relationship between the login event and the malicious result ABNORMAL is that the login event has the malicious result ABNORMAL, so as to obtain Figure 6 the knowledge graph shown.

[0151] Among them, in the Figure 5 knowledge graph of the file shown in 5 or 6, the circular nodes represent entities, and the arrowed lines connect two entities, indicating that there is an association relationship between the two entities. The text on the arrow represents the type of the association relationship between the two entities. For the two entities connected by the arrowed line, in the association relationship between the entity pointed to by the arrow of the line and the other entity, the entity pointed to by the arrow of the line is passive relative to the other entity, and the other entity is active relative to the entity pointed to by the arrow of the line. That is, the association relationship starts from the active entity and contacts the passive entity. It can be seen that the order of the active entities often comes before the order of the passive entities.

[0152] In this way, the entity serving as the starting node in the knowledge graph of the file can be regarded as: the entity not pointed to by the arrow. For example, in Figure 5In the knowledge graph shown in FIG. 5 or FIG. 6, the entity not pointed to by an arrow is the process name. Thus, the knowledge graph of this file can be traversed starting from the process name. Among them, it is possible to traverse Figure 6 the knowledge graph shown, and during the traversal of Figure 6 the knowledge graph shown, whenever an entity is traversed in the knowledge graph, it is determined whether the entity is a file name. If the entity is a file name, it can be determined whether the file name is the file name of a malicious file. For example, it can be determined whether there is an association relationship between the file name and the malicious result ABNORMAL. If there is an association relationship between the file name and the malicious result ABNORMAL, it can be determined that the file name is the file name of a malicious file. Or, if there is no association relationship between the file name and the malicious result ABNORMAL, it can be determined that the file name is not the file name of a malicious file.

[0153] In the case where the file name is the file name of a malicious file, step S303 can be executed.

[0154] However, the external security detection result is also generated based on known malicious entities. Thus, by using the external security detection result to detect whether the file name in the knowledge graph is the file name of a malicious file, only the file names of known malicious files can be determined to be the file names of malicious files, and the external security detection result does not involve currently unknown malicious entities, and thus does not involve the file names of unknown malicious files. Thus, in the case where the file name in the knowledge graph is the file name of an unknown malicious file, it is impossible to determine that the file name in the knowledge graph is the file name of a malicious file based on the external security detection result, which may lead to errors.

[0155] Thus, in order to avoid errors, in another embodiment, during the traversal of Figure 5 the knowledge graph shown in FIG. 5 or FIG. 6, in the case where an entity that is a file name is traversed, the knowledge graph can be input into a trained graph neural network model so that the trained graph neural network model processes the knowledge graph to obtain a detection result on whether the file name in the knowledge graph is the file name of a malicious file, and outputs the detection result.

[0156] The graph neural network model can be trained in advance. For example, a positive sample knowledge graph and a negative sample knowledge graph are obtained. The file names in the positive sample knowledge graph are not the file names of malicious files, and the file names in the negative sample knowledge graph are the file names of malicious files. The initialized graph neural network model is trained based on the positive sample knowledge graph and the negative sample knowledge graph until the parameters in the model converge to obtain a trained graph neural network model. Through the intelligence and learning ability of the graph neural network model, the discrimination accuracy of determining whether the file name in the knowledge graph is the file name of a malicious file can be improved.

[0157] However, sometimes, the knowledge graph of the document includes many entities of the preset type, and the associations between the entities of the preset type are very complex, resulting in a large amount of data in the knowledge graph of the document. For example, Figure 5 and 6 The knowledge graph shown is only an exemplary example. In actual situations, the knowledge graph of a file is often complex, and the knowledge graph includes many entities of preset types. For example, the knowledge graph of a file includes multiple processes, one process starts another process, another process starts another process, another process starts another process, and then another process starts the file. In addition, in addition to starting another process, the process also starts other entities, for example, other events are created, and other IP addresses may be involved. In addition to starting another process, another process also starts other entities, for example, other events are created, and other IP addresses may be involved. In addition to starting another process, another process also starts other entities, for example, other events are created, and other IP addresses may be involved. In addition to starting another process, another process also starts other entities, for example, other events are created, and other IP addresses may be involved. In addition to starting the file, another process also starts other entities, for example, other events are created, and other IP addresses may be involved. As a result, the knowledge graph of the file is very complex, which in turn leads to a large amount of data in the knowledge graph of the file. Therefore, if the knowledge graph of the file is directly input into the trained graph neural network model, it may take a long time for the trained graph neural network model to process the knowledge graph of the file, and the large amount of data may cause memory overflow. Therefore, in order to avoid the above problems, in another embodiment, when detecting whether the file name in the entity of the preset type in the knowledge graph of the file corresponding to the file name is the file name of a malicious file, the following process can be referred to:

[0158] 3021. Determine a process name and a file name in the knowledge graph of the file; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name.

[0159] The process name and file name determined in the knowledge graph of the file can be regarded as entities of preset types.

[0160] 3022. A subgraph is captured from the knowledge graph of the file, where the subgraph includes a determined process name and an entity of a preset type within P hops after the determined process name.

[0161] P is less than or equal to X, where X is the number of node hops between the determined process name and the entity of the preset type at the end of the knowledge graph of the file. The number of node hops between two entities is the sum of the number of layers where the entities between the two entities are located and the value "1", that is, how many hops it takes for a node to reach another node. For example, P can include 3, 4, 5 or 6, etc., depending on the actual situation, and this application does not limit this.

[0162] The subgraph has a determined process name and an entity of a preset type within P hops after the determined process name in the knowledge graph of the file, but does not have an entity before the determined process name in the knowledge graph of the file.

[0163] The subgraph may reflect the association relationship between a determined process name and an entity of a preset type within P hops after the determined process name, and may reflect the association relationship between entities of a preset type within P hops after the determined process name, and, in the subgraph, the entity of the preset type within P hops after the determined process name has a determined file name.

[0164] 3023. Input the subgraph into the trained graph neural network model so that the trained graph neural network model processes the subgraph to obtain a detection result of whether the determined file name in the subgraph is the file name of a malicious file.

[0165] In the subgraph input to the trained graph neural network model, there is a determined process name and an entity of a preset type within P hops after the determined process name, but there is no entity before the determined process name in the knowledge graph of the file. In the subgraph, there is a determined process name and a determined file name, and in the knowledge graph of the file, the entity farther from the determined process name and the determined file name has a more distant association with the determined process name and the determined file name, and the entity closer to the determined process name and the determined file name has a closer association with the determined process name and the determined file name.

[0166] Generally, the process corresponding to a certain process name is used to process the file corresponding to a certain file name, and the process corresponding to the certain process name is proactive to the file corresponding to the certain file name. Therefore, in the knowledge graph of the file, the entity before the certain process name is often more distant from the file corresponding to the certain file name.

[0167] Therefore, the entities of the preset type within P jumps can focus on the two key pieces of information, namely the determined process name and the determined file name, in the knowledge graph of the file, and can also focus on other entities that are closely related to the association relationship between the determined process name and the determined file name, discard entities that are farther from the determined process name and the determined file name, and discard entities that are before the determined process name, and reduce the use of entities that are more distantly related to the association relationship between the determined process name and the determined file name. As a result, the key information can be focused. Without affecting the accuracy of the detection result of whether the determined file name is the file name of a malicious file by the pre-trained graph neural network model, the time consumption of the processing process can be reduced, and the amount of data to be processed can be reduced, so as to avoid memory overflow as much as possible.

[0168] In the case where the file name in the knowledge graph of the file is the file name of a malicious file, in step S303, the first running information is obtained according to the knowledge graph of the file.

[0169] The entity with the file name of the preset type in the knowledge graph of the file.

[0170] In one embodiment, when obtaining the first running information according to the knowledge graph of the file, the knowledge graph of the file can be used as the first running information.

[0171] However, sometimes, there are many entities of the preset type included in the knowledge graph of the file, and the association relationships between many entities of the preset type are very complex, resulting in a large amount of data in the knowledge graph of the file. For example, Figure 5The knowledge graph shown in FIG. 5 or FIG. 6 is merely an exemplary example. In actual situations, the knowledge graph of a file is often more complex. The knowledge graph of a file will include many entities of preset types. For example, the knowledge graph of a file includes multiple processes. One process starts another process, the other process starts another process, and this other process starts yet another process, and yet another process starts still another process, and only then does this still another process start the file. In addition, in addition to starting another process, this one process also starts other entities. For example, it also creates other events, etc. In addition, it may also involve other IP addresses, etc. Another process, in addition to starting another process, also starts other entities. For example, it also creates other events, etc. In addition, it may also involve other IP addresses, etc. Yet another process, in addition to starting still another process, also starts other entities. For example, it also creates other events, etc. In addition, it may also involve other IP addresses, etc. Still another process, in addition to starting the file, also starts other entities. For example, it also creates other events, etc. In addition, it may also involve other IP addresses, etc. This leads to the knowledge graph of the file being very complex, and further leads to a large amount of data in the knowledge graph of the file. Therefore, if the knowledge graph of the file is directly used as the first running information, it may lead to a large amount of data in the first running information. When using the first running information to train a decision tree model later, it may lead to a long calculation time for the first running information during the training process, and due to the large amount of data, it may lead to an out-of-memory situation. Therefore, in order to avoid the above problems, in another embodiment, when obtaining the first running information according to the knowledge graph of the file, the following process can be referred to:

[0172] 3031. Determine the process name of the malicious process in the knowledge graph of the file.

[0173] 3032. In the knowledge graph of the file, select the entities of the preset type within N hops after the process name of the malicious process.

[0174] N is less than or equal to Y, where Y is the number of node hops between the process name of the malicious process and the outermost entity of the preset type in the knowledge graph of the file. For example, N can include 3, 4, 5, or 6, etc., and can be specifically determined according to the actual situation. This application does not limit this.

[0175] 3033. Determine whether the file name of the malicious file exists among the entities of the preset type within N hops after the process name of the malicious process.

[0176] 3034. When the file name of the malicious file exists among the entities of the preset type within N hops after the process name of the malicious process, intercept a subgraph in the knowledge graph of the file. The subgraph includes the process name of the malicious process and the entities of the preset type within N hops after the process name of the malicious process. The subgraph has the file name of the malicious file, and the malicious process is used to process the malicious file.

[0177] The subgraph has the process name of the malicious process and an entity of a preset type within N hops after the process name of the malicious process in the knowledge graph, but does not have an entity before the process name of the malicious process in the knowledge graph.

[0178] The subgraph can reflect the association relationship between the process name of the malicious process and the entity of the preset type located within N hops after the process name of the malicious process, and can reflect the association relationship between the entities of the preset type located within N hops after the process name of the malicious process, and, in the subgraph, the entity of the preset type located within N hops after the process name of the malicious process has the file name of the malicious file, wherein the malicious process is used to process the malicious file, and the malicious file is the file corresponding to the file name, that is, there is an association relationship between the process name of the malicious process and the file name of the malicious file, and the malicious process corresponding to the process name is used to process the malicious file corresponding to the file name.

[0179] 3035. Obtain first operation information according to the subgraph.

[0180] For example, the sub-graph is used as the first operation information, etc.

[0181] The first running information contains the process name of the malicious process and an entity of a preset type within N hops after the process name of the malicious process, but does not contain an entity before the process name of the malicious process in the knowledge graph.

[0182] The first operation information may reflect the association between the process name of the malicious process and the entity of the preset type located within N hops after the process name of the malicious process, and the association between the entities of the preset type located within N hops after the process name of the malicious process may be reflected, and, in the first operation information, the entity of the preset type located within N hops after the process name of the malicious process has the file name of the malicious file, wherein the malicious process is used to process the malicious file. The malicious file is the file corresponding to the file name, that is, the process name of the malicious process has an association with the file name of the malicious file, and the malicious process corresponding to the process name is used to process the malicious file corresponding to the file name.

[0183] The first running information contains the process name of a malicious process and the file name of a malicious file processed by the malicious process. In the knowledge graph, entities that are farther away from the process name of the malicious process and the file name of the malicious file have a more distant association with the process name of the malicious process and the file name of the malicious file, and entities that are closer to the process name of the malicious process and the file name of the malicious file have a more close association with the process name of the malicious process and the file name of the malicious file.

[0184] Normally, the process that handles malicious files can often be regarded as a malicious process, and the malicious process is proactive towards the malicious file. Therefore, in the knowledge graph, the entity before the process name of the malicious process is often more distant from the malicious file. Therefore, the preset type of entity within N hops can focus on the two key information of the process name of the malicious process and the file name of the malicious file in the knowledge graph, and can focus on other entities that are closely related to the process name of the malicious process and the file name of the malicious file, etc., discarding entities that are more distant from the process name of the malicious process and the file name of the malicious file, and discarding entities that are located before the process name of the malicious process, reducing the use of entities that are more distantly related to the process name of the malicious process and the file name of the malicious file, so as to focus on key information, reduce the calculation time without affecting the detection accuracy of the detection rules generated subsequently, and reduce the amount of data calculated to avoid memory overflow as much as possible.

[0185] It should be noted that if the process of S3021 to 3023 is executed in step S302, and the detection result that the file name in the sub-graph is the file name of a malicious file is obtained, then step 3035 can be directly executed in this step, and steps 3031 to 3034 can no longer be executed, thereby saving time and computing resources.

[0186] When the file name in the knowledge graph of the file is not the file name of a malicious file, in step S304, second operation information is obtained according to the knowledge graph of the file.

[0187] The file name in the knowledge graph of this file is an entity of a preset type.

[0188] In one embodiment, when obtaining the second operation information based on the knowledge graph of the file, the knowledge graph of the file may be used as the second operation information.

[0189] However, sometimes, the knowledge graph of the document includes many entities of the preset type, and the associations between the entities of the preset type are very complex, resulting in a large amount of data in the knowledge graph of the document. For example, Figure 5The knowledge graph shown in Figure 5 or 6 is merely an exemplary example. In actual situations, the knowledge graph of a file is often relatively complex. The knowledge graph of a file will include many preset types of entities. For example, the knowledge graph of a file includes multiple processes, one process starts another process, the other process starts another process, and this other process starts yet another process, and yet another process starts still another process, and only then does this still another process start the file. Additionally, in addition to starting another process, this one process also starts other entities, such as creating other events, etc. Additionally, other IP addresses, etc. may also be involved. Another process, in addition to starting another process, also starts other entities, such as creating other events, etc. Additionally, other IP addresses, etc. may also be involved. Another process, in addition to starting yet another process, also starts other entities, such as creating other events, etc. Additionally, other IP addresses, etc. may also be involved. Still another process, in addition to starting the file, also starts other entities, such as creating other events, etc. Additionally, other IP addresses, etc. may also be involved. This leads to the knowledge graph of the file being very complex, and further leads to a large amount of data in the knowledge graph of the file. Therefore, if the knowledge graph of the file is directly used as the second running information, it may lead to a large amount of data in the second running information. When using the second running information to train the decision tree model later, it may lead to a long calculation time for the second running information during the training process, and due to the large amount of data, it may lead to memory overflow. Therefore, to avoid the above problems, in another embodiment, when obtaining the second running information according to the knowledge graph of the file, the following process can be referred to:

[0190] 3041. Determine the process name and file name in the knowledge graph of the file. The association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name.

[0191] Both the process name and the file name determined in the knowledge graph of the file can be regarded as preset types of entities.

[0192] 3042. Intercept a sub-graph in the knowledge graph of the file. The sub-graph includes the determined process name and preset types of entities within M hops after the determined process name.

[0193] M is less than or equal to Z, where Z is the number of node hops between the determined process name and the most terminal preset type of entity in the knowledge graph; for example, M can include 3, 4, 5, or 6, etc., and can be specifically determined according to the actual situation, and this application does not limit this. M can be greater than or equal to the aforementioned N.

[0194] The subgraph has a determined process name and an entity of a preset type within M hops after the determined process name in the knowledge graph of the file, but does not have an entity before the determined process name in the knowledge graph of the file.

[0195] The subgraph may reflect the association between the determined process name and the entity of the preset type within M hops after the determined process name, and the association between the entities of the preset type within M hops after the determined process name. In the subgraph, the entity of the preset type within M hops after the determined process name has a determined file name. The determined process name is the process name of a non-malicious process, the determined file name is the file name of a non-malicious file, and the non-malicious process is used to process the non-malicious file, that is, there is an association between the process name of the non-malicious process and the file name of the non-malicious file, and the non-malicious process corresponding to the process name is used to process the non-malicious file corresponding to the file name.

[0196] 3043. Obtain second operation information according to the subgraph.

[0197] For example, the sub-graph is used as the second operation information, etc.

[0198] The second running information has the process name of the non-malicious process and an entity of a preset type within M hops after the process name of the non-malicious process, but does not have an entity before the process name of the non-malicious process in the knowledge graph.

[0199] The second operation information may reflect the association between the process name of the non-malicious process and the entity of the preset type located within M hops after the process name of the non-malicious process, and the association between the entity of the preset type located within M hops after the process name of the non-malicious process may be reflected, and in the second operation information, the entity of the preset type located within M hops after the process name of the non-malicious process has the file name of the non-malicious file, wherein the non-malicious process is used to process the non-malicious file. The second operation information has the process name of the non-malicious process and the file name of the non-malicious file processed by the non-malicious process. In the knowledge graph, the entity that is farther away from the process name of the non-malicious process and the file name of the non-malicious file has a more distant association with the process name of the non-malicious process and the file name of the non-malicious file, and the entity that is closer to the process name of the non-malicious process and the file name of the non-malicious file has a closer association with the process name of the non-malicious process and the file name of the non-malicious file.

[0200] Under normal circumstances, the processes that handle non-malicious files can often be regarded as non-malicious processes. Non-malicious processes are proactive towards non-malicious files. Therefore, in the knowledge graph of such files, the entities before the process names of non-malicious processes often have a more distant relationship with non-malicious files. Thus, the entities of the preset type within M hops can focus on the two key pieces of information, namely, the process names of non-malicious processes and the file names of non-malicious files, in the knowledge graph, and can also focus on other entities that have a close association with the process names of non-malicious processes and the file names of non-malicious files. It discards the entities that are farther away from the process names of non-malicious processes and the file names of non-malicious files, and discards the entities before the process names of non-malicious processes, reducing the use of entities that have a more distant association with the process names of non-malicious processes and the file names of non-malicious files. As a result, it can focus on the key information, reduce the computing time consumption, reduce the amount of data to be computed, and as much as possible avoid memory overflow and the like without affecting the detection accuracy of the subsequent generated detection rules.

[0201] It should be noted that if the processes of S3021 to 3023 are executed in step S302 and the detection result that the file name in the subgraph is not the file name of a malicious file is obtained, then in this step, step 3044 can be directly executed, and steps 3041 to 3043 do not need to be executed anymore, thus saving time and computing resources.

[0202] Among them, if the running environment is a sandbox environment, the following defects 1-3 may exist:

[0203] 1. The sandbox environment is an isolated running environment. For security considerations, the sandbox environment usually does not connect to the external network. If the file is a malicious file, it may cause the malicious behaviors that need to connect to the network to trigger cannot be triggered. That is to say, the file will not execute the malicious behaviors that need to connect to the network in the sandbox environment, resulting in the malicious behaviors that need to connect to the network of the file cannot be captured in the sandbox environment. Since the malicious behaviors that need to connect to the network of the file cannot be captured in the sandbox environment, it will cause the detection rules not to involve the relevant information of the malicious behaviors that cannot be captured, so that the generated detection rules miss the relevant information of some malicious behaviors of the malicious file, and further cause the generated detection rules to be incomplete and have low generalization. It can be seen that the detection accuracy of detecting whether the behavior of the file is a malicious behavior based on the detection rules generated by this application is low.

[0204] 2. Currently, there are various means to counter sandbox environments. For example, if a file is a malicious file and has means to counter sandbox environments, the file will detect whether the running environment where the file is located is a sandbox environment. If the running environment where the file is located is not a sandbox environment, the file may execute malicious behavior. Or, if the running environment where the file is located is a sandbox environment, the file usually will not execute malicious behavior. It can be seen that if the running environment where the file is located is a sandbox environment, it often leads to the inability to trigger the malicious behavior of the file running in the sandbox environment, that is, the file running in the sandbox environment will not execute malicious behavior, resulting in the inability to capture the malicious behavior of the file running in the sandbox environment. Since the malicious behavior of the file running in the sandbox environment cannot be captured, it will cause the detection rules not to cover the relevant information of the malicious behavior that cannot be captured. In this way, the generated detection rules will miss some relevant information about the malicious behavior of malicious files, and then lead to the generated detection rules being incomplete and having low generalization. It can be seen that the detection accuracy of detecting whether the behavior of a file is malicious based on the detection rules generated by this application is low.

[0205] 3. A sandbox environment is a simulated running environment, not a real running environment. Sometimes, if a file is a malicious file, the file has requirements on whether the running environment where it is located is a real running environment (not a simulated running environment, which is the running environment commonly used by the majority of users). For example, when the file detects that the running environment where the file is located is a real running environment, it will determine that the running environment where the file is located meets the file's requirements for the running environment, and the file may execute malicious behavior in the real running environment. Or, when the file detects that the running environment where the file is located is not a real running environment, it will determine that the running environment where the file is located does not meet the file's requirements for the running environment, and the file usually does not execute malicious behavior in a non-real running environment, etc.

[0206] Therefore, if the running environment where the file is located is a non-real running environment, it often leads to the inability to trigger the malicious behavior of the file running in the non-real running environment, that is, the file running in the non-real running environment will not execute malicious behavior, resulting in the inability to capture the malicious behavior of the file running in the non-real running environment. Since the malicious behavior of the file running in the non-real running environment cannot be captured, it will cause the detection rules not to cover the relevant information of the malicious behavior that cannot be captured. In this way, the generated detection rules will miss some relevant information about the malicious behavior of malicious files, and then lead to the generated detection rules being incomplete and having low generalization. It can be seen that the detection accuracy of detecting whether the behavior of a file is malicious based on the detection rules generated by this application is low.

[0207] To this end, in one embodiment, the operating environment in the present application may include a real operating environment. The real operating environment is not an isolated operating environment. The real operating environment is an operating environment capable of interacting with the outside world. The real operating environment is not a simulated operating environment. For example, the real operating environment is not a sandbox environment. The real operating environment may include an operating system installed on a real electronic device, etc. The operating system may include operating systems used by a large number of manufacturers in the market, such as Windows operating system, Android operating system, or Linux operating system, etc.

[0208] It can be seen that the detection rules of the present application are decoupled from the simulated operating environment (such as a sandbox environment). The detection rules in the present application are automatically generated according to the relevant information of the file actually running in the real operating environment according to logic. The file runs in the real operating environment. Files running in the real operating environment often perform malicious behaviors, so that the relevant information of the file actually running in the real operating environment will include the relevant information of malicious behaviors. In this way, the generated detection rules can cover more or even all the relevant information of the malicious behaviors of malicious files, and can avoid missing the relevant information of the malicious behaviors of malicious files as much as possible. Furthermore, the generated detection rules are complete and have high generalization. It can be seen that the detection accuracy of detecting whether the behavior of the file to be detected is a malicious behavior based on the detection rules generated by the present application is high.

[0209] After establishing the detection rules, the detection rules can be put into use online. For example, use the detection rules to detect whether a file is a malicious file. Among them, referring to Figure 7 , a method for detecting files in the present application is shown. This method is applied to an electronic device in a cloud computing scenario. The electronic device in the cloud computing scenario may include: a physical device or a virtual device. The physical device includes a server or a terminal, etc. The virtual device may include a virtual machine installed on the physical device, etc. The virtual machine may include an ECS, etc. Among them, this method may include:

[0210] In step S401, obtain the running information of the file to be detected. The running information of the file to be detected includes: multiple preset types of entities involved in the process of the file to be detected running in the operating environment. The multiple preset types of entities at least include: the file name of the file to be detected and the process name used to process the file to be detected.

[0211] In one embodiment, the detection rules may open the API (Application Programming Interface) of the detection rules to the outside world, so that in the case of needing to detect whether the file to be detected is a malicious file later, the detection rules can be called through the API of the detection rules to detect whether the file to be detected is a malicious file.

[0212] Alternatively, in another embodiment, the cloud computing system has a database for storing the running information of each file in the cloud computing system, and the running information of each file is respectively associated with the MD5 of each file. The database exposes the API of the database. When it is necessary to use the detection rule to detect whether a certain file is a malicious file, the MD5 of the certain file can be obtained first, and then the API of the database can be called according to the DM5 of the certain file to retrieve the running information of the certain file from the database, and then whether the certain file is a malicious file can be detected according to the running information of the certain file and the detection rule.

[0213] In step S402, the attribute information of multiple preset types of entities in the running information of the file to be detected is obtained.

[0214] The running information of the file to be detected further includes: the association relationship between multiple preset types of entities involved in the process of the file to be detected running in the running environment. Thus, when obtaining the attribute information of multiple preset types of entities in the running information of the file to be detected, the knowledge graph of the file to be detected can be obtained according to the multiple preset types of entities in the running information of the file to be detected and the association relationship between the multiple preset types of entities; the process name and the file name are determined in the knowledge graph of the file to be detected; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a sub-graph is intercepted in the knowledge graph of the file to be detected, and the sub-graph includes the determined process name and the preset types of entities within Q hops after the determined process name, where Q is less than or equal to V, and V is the number of node hops between the determined process name and the outermost preset type of entity in the knowledge graph; the attribute information of each preset type of entity in the sub-graph is obtained.

[0215] In step S403, according to the attribute information and the detection rule generated in advance for detecting malicious files, it is detected whether the file to be detected is a malicious file.

[0216] The attribute information can be matched with the detection rule generated in advance for detecting malicious files. If the attribute information hits the detection rule, it can be determined that the file to be detected is a malicious file, or if the attribute information does not hit the detection rule, it can be determined that the file to be detected is not a malicious file.

[0217] In a possible case, there may be multiple pieces of running information of a certain file retrieved from the database. In this case, it may be because the certain file has been run separately on different electronic devices in the cloud computing system. Each electronic device generates the running information of the certain file when running the certain file, associates the generated running information of the certain file with the MD5 of the certain file, and stores it in the database.

[0218] Thus, in one embodiment, there are multiple pieces of running information of the file to be detected; the file names in the multiple pieces of running information are all the file name of the file to be detected, and the process names in the multiple pieces of running information are different.

[0219] Thus, when obtaining the attribute information of multiple preset types of entities in the running information of the file to be detected, for any piece of running information of the file to be detected, obtain the attribute information of the multiple preset types of entities in the running information to obtain the attribute information corresponding to the running information; for each of the other pieces of running information of the file to be detected, do the same, so as to obtain the attribute information corresponding to each piece of running information of the file to be detected respectively.

[0220] Correspondingly, when detecting whether the file to be detected is a malicious file according to the attribute information and the detection rules generated in advance for detecting malicious files, for the attribute information corresponding to any piece of running information of the file to be detected, match the attribute information corresponding to the running information with the detection rules to obtain the detection result indicating whether the file to be detected is a malicious file corresponding to the attribute information of the running information; for the attribute information corresponding to each of the other pieces of running information of the file to be detected, do the same, and then count the number of the attribute information corresponding to the running information whose detection result indicates that the file to be detected is a malicious file; in the case where the ratio between the number and the number of the multiple pieces of running information of the file to be detected is greater than a preset threshold, determine that the file to be detected is a malicious file; or, in the case where the ratio between the number and the number of the multiple pieces of running information of the file to be detected is less than or equal to the preset threshold, determine that the file to be detected is not a malicious file. The preset threshold may include 0.5, 0.55, 0.6, etc., and can be determined according to the actual situation specifically, and the present application does not limit this.

[0221] Furthermore, the method further includes: screening out misdetected files that are actually not malicious files but are detected as malicious files by the detection rules in the detected files (it can be misdetected files manually indicated as actually not malicious files but detected as malicious files by the detection rules, etc.); determining the logical expression involved in detecting the misdetected files as malicious files in the detection rules (it can be known which logical expression in the detection rules the misdetected files are determined to be malicious files because they hit); deleting the determined logical expression in the detection rules.

[0222] Further, a new logical expression is manually constructed for the misdetected file, and the new logical expression is input into the electronic device. The electronic device receives the new logical expression input by the human and adds the new logical expression to the detection rule to update the detection rule.

[0223] Among them, if the running environment is a sandbox environment, there may be the following defects 1-3:

[0224] 1. The sandbox environment is an isolated running environment. For security reasons, the sandbox environment usually does not connect to the external network. If the file to be detected is a malicious file, it may cause the malicious behavior that needs to connect to the network to be triggered in the file to be detected not to be triggered. That is, the file to be detected will not execute the malicious behavior that needs to connect to the network in the sandbox environment, resulting in the malicious behavior of the file to be detected that needs to connect to the network not being captured in the sandbox environment. Since the malicious behavior of the file to be detected that needs to connect to the network cannot be captured in the sandbox environment, in the case where the file to be detected has no other malicious behavior, the detection result that the behavior of the file to be detected in the sandbox environment has no malicious behavior will be output, and then it will be determined that the file to be detected is not a malicious file but a safe file, which may lead to malicious files not being detected, resulting in missed detections. It can be seen that the detection accuracy rate is low.

[0225] 2. Currently, there are various anti-sandbox environment means. For example, if the file to be detected is a malicious file and has anti-sandbox environment means, the file to be detected will detect whether the running environment where the file to be detected is located is a sandbox environment. If the running environment where the file to be detected is located is not a sandbox environment, the file to be detected may execute malicious behavior. Or, if the running environment where the file to be detected is located is a sandbox environment, the file to be detected usually will not execute malicious behavior.

[0226] It can be seen that if the running environment where the file to be detected is located is a sandbox environment, it will often cause the malicious behavior of the file to be detected running in the sandbox environment not to be triggered. That is, it will cause the file to be detected running in the sandbox environment not to execute malicious behavior, resulting in the malicious behavior of the file to be detected running in the sandbox environment not being captured. Since the malicious behavior of the file to be detected running in the sandbox environment cannot be captured, in the case where the file to be detected has no other malicious behavior, the detection result that the behavior of the file to be detected in the sandbox environment has no malicious behavior will be output, and then it will be determined that the file to be detected is not a malicious file but a safe file, which may lead to malicious files not being detected, resulting in missed detections. It can be seen that the detection accuracy rate is low.

[0227] 3. The sandbox environment is a simulated operating environment, not a real one. If the file to be detected is a malicious file, sometimes the file to be detected has requirements for whether the operating environment it is in is a real operating environment (not a simulated one, but the operating environment commonly used by the majority of users). For example, when the file to be detected detects that the operating environment where it is located is a real operating environment, it will determine that the operating environment where the file to be detected is located meets the requirements of the file to be detected for the operating environment, and the file to be detected may perform malicious behaviors only in the real operating environment. Or, when the file to be detected detects that the operating environment where it is located is not a real operating environment, it will determine that the operating environment where the file to be detected is located does not meet the requirements of the file to be detected for the operating environment, and the file to be detected often does not perform malicious behaviors in a non-real operating environment, etc.

[0228] Therefore, if the operating environment where the file to be detected is located is a non-real operating environment, it often leads to the inability to trigger the malicious behaviors of the file to be detected running in the non-real operating environment. That is, the file to be detected running in the non-real operating environment does not perform malicious behaviors, resulting in the inability to capture the malicious behaviors of the file to be detected running in the non-real operating environment. Since the malicious behaviors of the file to be detected running in the non-real operating environment cannot be captured, in the case where the file to be detected has no other malicious behaviors, the detection result that the behaviors of the file to be detected in the sandbox environment are not malicious will be output. Furthermore, it will be determined that the file to be detected is not a malicious file but a safe file, which may lead to the failure to detect malicious files and cause missed detections. Obviously, the detection accuracy is low.

[0229] In one embodiment, the operating environment in the present application may include a real operating environment, so that the present application has the following beneficial effects 1-3:

[0230] 1. The present application does not use a simulated operating environment. For example, it does not use a sandbox environment, but uses a real operating environment. The real operating environment often accesses the external network. When the file to be detected runs in the real operating environment, if the file to be detected is a malicious file, the malicious behaviors that need to be triggered by networking in the file to be detected often will be triggered. That is, the file to be detected often performs malicious behaviors that need to be triggered by networking in the real operating environment, enabling the malicious behaviors that need to be triggered by networking to be captured. Since the malicious behaviors of the file to be detected in the real operating environment can be captured, thus, according to the detection rules, the detection result that the file to be detected has malicious behaviors in the real operating environment can often be obtained, so that the malicious file can be detected and missed detections can be avoided. Obviously, the detection accuracy can be improved.

[0231] 2. If there are various means to counter sandbox environments currently, for example, if the file to be detected is a malicious file and has means to counter sandbox environments, then the file to be detected will check whether the running environment where the file to be detected is located is a sandbox environment. If the running environment where the file to be detected is located is not a sandbox environment, then the file to be detected may execute malicious behaviors. Or, if the running environment where the file to be detected is located is a sandbox environment, then the file to be detected usually will not execute malicious behaviors.

[0232] Since this application does not use a simulated running environment, for example, does not use a sandbox environment, but uses a real running environment, that is, the file to be detected runs in a real running environment. Therefore, the file to be detected will detect that the running environment where the file to be detected is located is a real running environment, rather than a sandbox environment. Thus, even if the file to be detected has means to counter sandbox environments, because the real running environment is not a sandbox environment, the file to be detected running in the real running environment usually will execute malicious behaviors, making the malicious behaviors of the file to be detected running in the real running environment tend to be triggered. That is, the file to be detected running in the real running environment usually will execute malicious behaviors, making the malicious behaviors of the file to be detected running in the real running environment able to be captured. Since the malicious behaviors of the file to be detected running in the real running environment can be captured, in this way, according to the detection rules, usually the detection result that the file to be detected has malicious behaviors in the real running environment can be obtained, so that it can be determined that the file to be detected is a malicious file, enabling malicious files to be detected and avoiding missed detections. It can be seen that the detection accuracy can be improved.

[0233] 3. Since this application does not use a simulated running environment, for example, does not use a sandbox environment, but uses a real running environment, and the file to be detected runs in the real running environment. If the file to be detected is a malicious file, then when the file to be detected detects that the running environment where the file to be detected is located is a real running environment (not a simulated running environment, but the running environment commonly used by the majority of users), the file to be detected may execute malicious behaviors, making the malicious behaviors of the file to be detected running in the real running environment tend to be triggered. That is, the file to be detected running in the real running environment usually will execute malicious behaviors, making the malicious behaviors of the file to be detected running in the real running environment able to be captured. Since the malicious behaviors of the file to be detected running in the real running environment can be captured, in this way, according to the detection rules, usually the detection result that the file to be detected has malicious behaviors in the real running environment can be obtained, so that it can be determined that the file to be detected is a malicious file, enabling malicious files to be detected and avoiding missed detections. It can be seen that the detection accuracy can be improved.

[0234] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily essential to this application.

[0235] Referring to Figure 8 , a structural block diagram of a device for establishing a detection rule is shown, including: a first acquisition module 11, configured to acquire the running information of multiple files running in a running environment, where the running information of the files includes: multiple preset types of entities involved in the process of the files running in the running environment, and the multiple preset types of entities at least include: the file name of the file and the process name for processing the file; a determination module 12, configured to determine first running information and second running information from the running information of the multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; a second acquisition module 13, configured to acquire the attribute information of the multiple preset types of entities in the first running information, and acquire the attribute information of the multiple preset types of entities in the second running information; a generation module 14, configured to generate a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information.

[0236] Among them, the generation module includes: a training unit, configured to train a decision tree model according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information to obtain a binary tree structure; a first generation unit, configured to generate a detection rule for detecting malicious files according to the binary tree structure.

[0237] Among them, the first generation unit is specifically configured to traverse the binary tree structure in a pre-order traversal manner to generate a detection rule for detecting malicious files.

[0238] Among them, the running environment is loaded on an electronic device; the first acquisition module includes: a first acquisition unit, configured to acquire at least the log information of the electronic device, where the log information of the electronic device includes: the log information generated by the electronic device during the process of running multiple files in the running environment on the electronic device; a second generation unit, configured to generate the running information of the multiple files running in the running environment at least according to the log information of the electronic device.

[0239] Among them, the second generation unit includes: a search subunit, configured to search for multiple file names in the log information of the electronic device; a generation subunit, configured to, for any one of the multiple file names found, extract at least one preset type of entity associated with the file name from the log information of the electronic device, and the at least one preset type of entity at least includes: the process name for processing the file corresponding to the file name, and generate the running information of the file corresponding to the file name running in the running environment at least based on the file name and the at least one preset type of entity.

[0240] Among them, the running information of the file further includes: the association relationship between multiple preset types of entities involved in the running process of the file in the running environment; the determination module includes: a second acquisition unit, configured to, for the running information of any one file, acquire the knowledge graph of the file according to the multiple preset types of entities in the running information of the file and the association relationship between the multiple preset types of entities; a detection unit, configured to detect whether the file name in the knowledge graph is the file name of a malicious file; a third acquisition unit, configured to, when the file name in the knowledge graph is the file name of a malicious file, acquire the first running information according to the knowledge graph; or a fourth acquisition unit, configured to, when the file name in the knowledge graph is not the file name of a malicious file, acquire the second running information according to the knowledge graph.

[0241] Among them, the detection unit includes: a first determination subunit, configured to determine the process name and the file name in the knowledge graph; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a first interception subunit, configured to intercept a subgraph in the knowledge graph, and the subgraph includes the determined process name and the preset type of entities within P hops after the determined process name, where P is less than or equal to X, and X is the number of node hops between the determined process name and the outermost preset type of entity in the knowledge graph; an input subunit, configured to input the subgraph into the trained graph neural network model, so that the trained graph neural network model processes the subgraph to obtain the detection result of whether the determined file name in the subgraph is the file name of a malicious file.

[0242] Among them, the third acquisition unit includes: a second determination subunit, configured to determine the process name of a malicious process in the knowledge graph; a selection subunit, configured to select, in the knowledge graph, preset type entities within N hops after the process name of the malicious process, where N is less than or equal to Y, and Y is the number of node hops between the process name of the malicious process and the preset type entity at the end of the knowledge graph; a third determination subunit, configured to determine whether the file name of a malicious file exists among the preset type entities within N hops after the process name of the malicious process; a second interception subunit, configured to intercept a subgraph in the knowledge graph when the file name of a malicious file exists among the preset type entities within N hops after the process name of the malicious process, where the subgraph includes the process name of the malicious process and the preset type entities within N hops after the process name of the malicious process, and the malicious process is used to process the malicious file; a first acquisition subunit, configured to obtain first running information according to the subgraph.

[0243] Among them, the fourth acquisition unit includes: a fourth determination subunit, configured to determine a process name and a file name in the knowledge graph, and the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a third interception subunit, configured to intercept a subgraph in the knowledge graph, where the subgraph includes the determined process name and the preset type entities within M hops after the determined process name, and M is less than or equal to Z, and Z is the number of node hops between the determined process name and the preset type entity at the end of the knowledge graph; a second acquisition subunit, configured to obtain second running information according to the subgraph.

[0244] Referring to Figure 9 , a structural block diagram of a device for detecting a file according to the present application is shown, including: a third acquisition module 21, configured to acquire running information of a file to be detected, where the running information of the file to be detected includes multiple preset type entities involved in the process of the file to be detected running in a running environment, and the multiple preset type entities at least include: the file name of the file to be detected and the process name of the process used to process the file to be detected; a fourth acquisition module 22, configured to acquire attribute information of the multiple preset type entities in the running information of the file to be detected; a detection module 23, configured to detect whether the file to be detected is a malicious file according to the attribute information and a detection rule for detecting a malicious file generated in advance.

[0245] Among them, the running information of the file to be detected further includes: the association relationships between multiple preset types of entities involved in the running process of the file to be detected in the running environment; the fourth acquisition module includes: a fifth acquisition unit, configured to acquire the knowledge graph of the file to be detected according to the multiple preset types of entities and the association relationships between the multiple preset types of entities in the running information of the file to be detected; a first determination unit, configured to determine the process name and the file name in the knowledge graph of the file to be detected; the association relationship between the determined process name and the determined file name includes that the process corresponding to the determined process name is used to process the file corresponding to the determined file name; a truncation unit, configured to truncate a sub-graph in the knowledge graph of the file to be detected, where the sub-graph includes the determined process name and the preset type of entities within Q hops after the determined process name, and Q is less than or equal to V, and V is the number of node hops between the determined process name and the outermost preset type of entity in the knowledge graph; a sixth acquisition unit, configured to acquire the attribute information of each preset type of entity in the sub-graph.

[0246] Among them, there are multiple pieces of running information of the file to be detected; the file names in the multiple pieces of running information are all the file name of the file to be detected, and the process names in the multiple pieces of running information are different; correspondingly, the fourth acquisition module includes: a seventh acquisition unit, configured to, for any piece of running information of the file to be detected, acquire the attribute information of the multiple preset types of entities in the running information to obtain the attribute information corresponding to the running information; correspondingly, the detection module includes: a matching unit, configured to, for the attribute information corresponding to any piece of running information of the file to be detected, match the attribute information corresponding to the running information with the detection rule to obtain the detection result indicating whether the file to be detected is a malicious file; a statistics unit, configured to count the number of the attribute information corresponding to the running information of the detection result indicating that the file to be detected is a malicious file; a second determination unit, configured to determine that the file to be detected is a malicious file when the ratio between the number and the number of the multiple pieces of running information of the file to be detected is greater than a preset threshold.

[0247] The embodiment of the present application further provides a non-volatile readable storage medium, in which one or more modules (programs) are stored, and when the one or more modules are applied to a device, the device can be caused to execute the instructions (instructions) of each method step in the embodiment of the present application.

[0248] The embodiment of the present application provides one or more machine-readable media, on which instructions are stored, and when executed by one or more processors, cause an electronic device to execute the methods as described in one or more of the above embodiments. In the embodiment of the present application, the electronic device includes a server, a gateway, a sub-device, etc., and the sub-device is a device such as an Internet of Things device.

[0249] Embodiments of the present disclosure can be implemented as a device configured as desired using any suitable hardware, firmware, software, or any combination thereof. The device may include electronic devices such as servers (clusters), terminal devices such as IoT devices, etc.

[0250] Figure 10 Exemplary device 1300 that can be used to implement various embodiments in the present application is schematically shown. For one embodiment, Figure 10 Exemplary device 1300 is shown, which has one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the (one or more) processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM, Non-Volatile Memory) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0251] Processor 1302 may include one or more single-core or multi-core processors. Processor 1302 may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1300 is capable of acting as a server device such as a gateway in the embodiments of the present application.

[0252] Device 1300 may include one or more computer-readable media (such as memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 combined with the one or more computer-readable media and configured to execute the instructions 1314 to implement modules to perform actions in the present disclosure.

[0253] Control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the (one or more) processors 1302 and / or any suitable device or component communicating with control module 1304. Control module 1304 may include a memory controller module to provide an interface to memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0254] Memory 1306 may be used, for example, to load and store data and / or instructions 1314 for device 1300. For one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1306 may include double data rate four synchronous dynamic random access memory (DDR4 SDRAM).

[0255] The control module 1304 may include one or more input / output controllers to provide an interface to the NVM / storage device 1308 and the input / output device(s) 1310. For example, the NVM / storage device 1308 may be used to store data and / or instructions 1314. The NVM / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more optical disc (CD) drives, and / or one or more digital versatile disc (DVD) drives). The NVM / storage device 1308 may include storage resources that are physically part of the device on which the device 1300 is mounted, or it may be accessible by the device without being part of the device. For example, the NVM / storage device 1308 may be accessed via the input / output device(s) 1310 over a network. The input / output device(s) 1310 may provide an interface for the device 1300 to communicate with any other suitable device, and the input / output device 1310 may include a communication component, a pinyin component, a sensor component, etc. The network interface 1312 may provide an interface for the device 1300 to communicate over one or more networks, and the device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0256] At least one of the processor(s) 1302 may be logically encapsulated with one or more controllers (e.g., a memory controller module) of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be logically encapsulated with one or more controllers of the control module 1304 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 1302 may be logically integrated on the same die with one or more controllers of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be logically integrated on the same die with one or more controllers of the control module 1304 to form a system-on-chip (SoC).

[0257] In various embodiments, the device 1300 may be, but is not limited to, a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.) and other terminal devices. In various embodiments, the device 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.

[0258] Embodiments of the present application provide an electronic device, including: one or more processors; and one or more machine-readable media storing instructions thereon, which, when executed by one or more processors, cause the electronic device to execute the methods as described in one or more of the present application.

[0259] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.

[0260] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable information processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable information processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto the computer or other programmable information processing terminal devices, such that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal devices provide for implementing the processFigure 1 steps of one process or multiple processes and / or blocks Figure 1 Steps of the functions specified in one block or multiple blocks. Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0261] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element. The above has introduced in detail the method and device for establishing a detection rule and the method and device for detecting a document provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for establishing a detection rule, characterized in that The method includes: Obtaining the running information of multiple files running in a running environment, where the running information of the files includes: multiple preset types of entities involved in the process of the files running in the running environment, and the multiple preset types of entities at least include: the file names of the files and the process names for processing the files; Determining first running information and second running information from the running information of the multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; Obtaining the attribute information of the multiple preset types of entities in the first running information, and obtaining the attribute information of the multiple preset types of entities in the second running information; Generating a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information.

2. The method according to claim 1, characterized in that The generating a detection rule for detecting malicious files according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information includes: Training a decision tree model according to the attribute information of the multiple preset types of entities in the first running information and the attribute information of the multiple preset types of entities in the second running information to obtain a binary tree structure; Generating a detection rule for detecting malicious files according to the binary tree structure.

3. The method according to claim 2, characterized in that, The generating a detection rule for detecting malicious files according to the binary tree structure includes: Traversing the binary tree structure in a pre-order traversal manner to generate a detection rule for detecting malicious files.

4. The method according to claim 1, wherein The running environment is loaded on an electronic device; The obtaining the running information of multiple files running in a running environment includes: At least obtaining the log information of the electronic device, where the log information of the electronic device includes: the log information generated by the electronic device during the process of running multiple files on the running environment of the electronic device; At least generating the running information of the multiple files running in the running environment according to the log information of the electronic device.

5. The method according to claim 4, characterized in that The at least generating the running information of the multiple files running in the running environment according to the log information of the electronic device includes: Searching for multiple file names in the log information of the electronic device; For any one of the multiple file names found, extracting at least one preset type of entity having an associated relationship with the file name from the log information of the electronic device, where the at least one preset type of entity at least includes: the process name for processing the file corresponding to the file name, and at least generating the running information of the file corresponding to the file name running in the running environment according to the file name and the at least one preset type of entity.

6. The method according to claim 1, characterized in that, The running information of the files further includes: the association relationship between multiple preset types of entities involved in the process of the files running in the running environment; The determining first running information and second running information from the running information of the multiple files includes: For the running information of any one file, obtaining the knowledge graph of the file according to the multiple preset types of entities in the running information of the file and the association relationship between the multiple preset types of entities. Detect whether the file name in the knowledge graph is the file name of a malicious file; In the case where the file name in the knowledge graph is the file name of a malicious file, obtain first running information according to the knowledge graph; Or, In the case where the file name in the knowledge graph is not the file name of a malicious file, obtain second running information according to the knowledge graph.

7. The method according to claim 6, wherein The detection of whether the file name in the knowledge graph is the file name of a malicious file includes: Determine the process name and file name in the knowledge graph; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; Intercept a subgraph in the knowledge graph, where the subgraph includes the determined process name and preset type entities within P hops after the determined process name, P is less than or equal to X, and X is the number of node hops between the determined process name and the outermost preset type entity in the knowledge graph; Input the subgraph into a trained graph neural network model so that the trained graph neural network model processes the subgraph to obtain the detection result of whether the determined file name in the subgraph is the file name of a malicious file.

8. The method according to claim 6, characterized in that, The obtaining of the first running information according to the knowledge graph includes: Determine the process name of the malicious process in the knowledge graph; In the knowledge graph, select preset type entities within N hops after the process name of the malicious process, N is less than or equal to Y, and Y is the number of node hops between the process name of the malicious process and the outermost preset type entity in the knowledge graph; Determine whether the file name of a malicious file exists among the preset type entities within N hops after the process name of the malicious process; In the case where the file name of a malicious file exists among the preset type entities within N hops after the process name of the malicious process, intercept a subgraph in the knowledge graph, where the subgraph includes the process name of the malicious process and the preset type entities within N hops after the process name of the malicious process, and the malicious process is used to process the malicious file; Obtain the first running information according to the subgraph.

9. The method according to claim 6, wherein The obtaining of the second running information according to the knowledge graph includes: Determine the process name and file name in the knowledge graph, and the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; Intercept a subgraph in the knowledge graph, where the subgraph includes the determined process name and preset type entities within M hops after the determined process name, M is less than or equal to Z, and Z is the number of node hops between the determined process name and the outermost preset type entity in the knowledge graph; Obtain the second running information according to the subgraph.

10. A method for detecting a file, characterized in that, The method includes: Obtain the running information of the file to be detected, where the running information of the file to be detected includes: multiple preset type entities involved in the process of the file to be detected running in the running environment, and the multiple preset type entities at least include: the file name of the file to be detected and the process name used to process the file to be detected; Obtain the attribute information of the multiple preset type entities in the running information of the file to be detected; Detect whether the file to be detected is a malicious file according to the attribute information and the pre-generated detection rules for detecting malicious files.

11. The method according to claim 10, wherein, The running information of the file to be detected further includes: the association relationships between multiple preset types of entities involved in the running process of the file to be detected in the running environment; The obtaining the attribute information of multiple preset types of entities in the running information of the file to be detected includes: According to multiple preset types of entities in the running information of the file to be detected and the association relationships between the multiple preset types of entities, obtain the knowledge graph of the file to be detected; Determine the process name and file name in the knowledge graph of the file to be detected; the association relationship between the determined process name and the determined file name includes: the process corresponding to the determined process name is used to process the file corresponding to the determined file name; Intercept a sub-graph in the knowledge graph of the file to be detected, where the sub-graph includes the determined process name and preset types of entities within Q hops after the determined process name, and Q is less than or equal to V, where V is the number of node hops between the determined process name and the outermost preset type of entity in the knowledge graph; Obtain the attribute information of each preset type of entity in the sub-graph.

12. The method according to claim 10, wherein There are multiple pieces of running information of the file to be detected; the file names in the multiple pieces of running information are all the file name of the file to be detected, and the process names in the multiple pieces of running information are different; Correspondingly, the obtaining the attribute information of multiple preset types of entities in the running information of the file to be detected includes: For any piece of running information of the file to be detected, obtain the attribute information of multiple preset types of entities in the running information to obtain the attribute information corresponding to the running information; Correspondingly, the detecting whether the file to be detected is a malicious file according to the attribute information and the pre-generated detection rules for detecting malicious files includes: For the attribute information corresponding to any piece of running information of the file to be detected, match the attribute information corresponding to the running information with the detection rules to obtain the detection result indicating whether the file to be detected is a malicious file of the attribute information corresponding to the running information; Count the number of the attribute information corresponding to the running information of the detection result indicating that the file to be detected is a malicious file; When the ratio between the number and the number of multiple pieces of running information of the file to be detected is greater than a preset threshold, determine that the file to be detected is a malicious file.

13. An apparatus for establishing a detection rule, characterized in that The device includes: A first obtaining module, configured to obtain the running information of multiple files running in the running environment, where the running information of the file includes: multiple preset types of entities involved in the running process of the file in the running environment, and the multiple preset types of entities at least include: the file name of the file and the process name for processing the file; A determining module, configured to determine first running information and second running information in the running information of multiple files, where the file name in the first running information is the file name of a malicious file, and the file name in the second running information is not the file name of a malicious file; A second obtaining module, configured to obtain the attribute information of multiple preset types of entities in the first running information, and obtain the attribute information of multiple preset types of entities in the second running information; A generation module, configured to generate a detection rule for detecting malicious files according to the attribute information of multiple preset types of entities in the first running information and the attribute information of multiple preset types of entities in the second running information.

14. A device for detecting a file, characterized in that, The device includes: A third acquisition module, configured to acquire the running information of the file to be detected, where the running information of the file to be detected includes: multiple preset types of entities involved in the process of the file to be detected running in the running environment, and the multiple preset types of entities at least include: the file name of the file to be detected and the process name for processing the file to be detected; A fourth acquisition module, configured to acquire the attribute information of multiple preset types of entities in the running information of the file to be detected; A detection module, configured to detect whether the file to be detected is a malicious file according to the attribute information and the detection rule for detecting malicious files generated in advance.

15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 12 is implemented.

16. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 12 is implemented.