Supply Chain Attack Detection Method, Device and Related Equipment

The method addresses the challenge of detecting supplier chain attacks by combining static and dynamic scanning to identify malicious installations in target packages, improving detection accuracy and reducing false positives.

CN114547603BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011341419.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-25
Publication Date
2025-07-15
Estimated Expiration
2040-11-25

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively detect and prevent supply chain attacks, resulting in numerous crises in enterprise data security.

Method used

By determining the malicious installation package in the target installation package, using the target static characteristics for static scanning and malicious code detection hook function for dynamic scanning, combining static and dynamic scanning methods to accurately detect malicious installation packages.

Benefits of technology

It improves the accuracy of the detection effect, reduces the false alarm rate of the detection results, and effectively prevents supply chain attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547603B_ABST
    Figure CN114547603B_ABST
Patent Text Reader

Abstract

The present disclosure provides a supply chain attack detection method, apparatus, electronic device, and computer-readable storage medium. The method includes: obtaining a target installation package; scanning the target installation package according to target static features in a target behavior scenario to obtain first candidate malicious codes from the target installation package; determining suspicious codes from the first candidate malicious codes according to malicious code baseline features, where the malicious code baseline features are known malicious code features; determining a suspicious installation package where the suspicious codes are located from the target installation package; injecting a malicious code detection hook function into the suspicious installation package and simulating the operation of the suspicious installation package; determining first malicious codes from second candidate malicious codes according to the malicious code baseline features; and determining a first malicious installation package where the first malicious codes are located from the suspicious installation package. The technical solution provided by the embodiments of the present disclosure can achieve the detection of supply chain attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer network technology, and in particular to a supply chain attack detection method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Supply chain attacks, also known as value chain attacks or third-party attacks, occur when an attacker infiltrates internal systems through external channels that have access to corporate systems and data, such as open source dependency libraries. This attack method has greatly changed the typical corporate attack method in the past few years, because the number of suppliers and service providers with access to sensitive data is greater than ever before. Supply chain attacks bring unprecedented high risks. Due to the emergence of new attack methods, the public's awareness of the threat continues to increase, and regulators have also strengthened their supervision. At the same time, attackers have more resources and tools than ever before. In this context, corporate security will face many crises.

[0003] Therefore, a detection method for supply chain attacks is very important to maintain enterprise data security.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure. Summary of the invention

[0005] The embodiments of the present disclosure provide a supply chain attack detection method and device, an electronic device, and a computer-readable storage medium, which can determine malicious installation packages in a target installation package to detect supply chain attacks.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.

[0007] The disclosed embodiment proposes a supply chain attack detection method, which includes: obtaining a target installation package; scanning the target installation package according to the target static characteristics under the target behavior scenario, and obtaining a first candidate malicious code from the target installation package; determining suspicious code in the first candidate malicious code according to the malicious code baseline characteristics, wherein the malicious code baseline characteristics are known malicious code characteristics; determining the suspicious installation package where the suspicious code is located from the target installation package; injecting a malicious code detection hook function into the suspicious installation package, and simulating the operation of the suspicious installation package; during the simulated operation of the suspicious installation package, obtaining a second candidate malicious code from the suspicious installation package through the malicious code detection hook function; determining a first malicious code in the second candidate malicious code according to the malicious code baseline characteristics; and determining a first malicious installation package where the first malicious code is located.

[0008] The disclosed embodiment provides a supply chain attack detection device, which includes: an installation package acquisition module, a first candidate malicious code determination module, a suspicious code determination module, a suspicious installation package determination module, a simulation running module, a second candidate malicious code determination module, a first malicious code determination module, and a first malicious installation package determination module.

[0009] Among them, the installation package acquisition module can be configured to acquire a target installation package. The first candidate malicious code determination module can be configured to scan the target installation package according to the target static features in a target behavior scenario, and acquire the first candidate malicious code from the target installation package. The suspicious code determination module can be configured to determine suspicious codes from the first candidate malicious codes according to the malicious code baseline features, where the malicious code baseline features are known malicious code features. The suspicious installation package determination module can be configured to determine the suspicious installation package where the suspicious codes are located from the target installation package. The simulation running module can be configured to inject a malicious code detection hook function into the suspicious installation package and simulate the running of the suspicious installation package. The second candidate malicious code determination module can be configured to acquire the second candidate malicious code from the suspicious installation package through the malicious code detection hook function during the simulation running of the suspicious installation package. The first malicious code determination module can be configured to determine the first malicious code from the second candidate malicious codes according to the malicious code baseline features. The first malicious installation package determination module can be configured to determine the first malicious installation package where the first malicious code is located.

[0010] In some embodiments, the target behavior scenario includes the behavior scenario of a target malicious operation, and the target static features include suspicious malicious operation features; among them, the first candidate malicious code determination module can include: an installation package scanning sub-module, a suspicious malicious operation code determination sub-module, and a first candidate malicious code determination sub-module.

[0011] Among them, the installation package scanning sub-module can be configured to scan the target installation package; the suspicious malicious operation code determination sub-module can be configured to determine the code including the suspicious malicious operation features in the target installation package as the suspicious malicious operation code; the first candidate malicious code determination sub-module can be configured to determine the first candidate malicious code according to the suspicious malicious operation code.

[0012] In some embodiments, the target behavior scenario includes an external connection behavior scenario, and the target static features include target external connection features; among them, the first candidate malicious code determination module can include: a first scanning sub-module, an external connection code determination sub-module, and a first candidate malicious code determination sub-module.

[0013] Among them, the first scanning sub-module can be configured to scan the target installation package. The external connection code determination sub-module can be configured to determine the code including the target external connection feature in the target installation package as the external connection code. The first candidate malicious code determination sub-module can be configured to determine the first candidate malicious code according to the external connection code.

[0014] In some embodiments, the first candidate malicious code determination sub-module can include: a target scenario parameter acquisition unit and a first candidate malicious code determination unit.

[0015] Among them, the target scenario parameter acquisition unit can be configured to acquire the target scenario parameters of the target behavior scenario. The first candidate malicious code determination unit can be configured to determine the first candidate malicious code according to the target scenario parameters and the external connection code.

[0016] In some embodiments, the target installation package can include target files; among them, the first candidate malicious code determination unit can include: a first value processing sub-unit, a second value processing sub-unit, and a third value processing sub-unit.

[0017] Among them, the first value processing sub-unit can be configured to, if the target scenario parameter is the first value, determine the external connection code as the first candidate malicious code. The second value processing sub-unit can be configured to, if the target scenario parameter is the second value, use the external connection code and the code on the upper and lower target scenario parameter lines thereof as the first candidate malicious code. The third value processing sub-unit can be configured to, if the target scenario code is the third value, determine the first candidate malicious code according to the target file where the external connection code is located.

[0018] In some embodiments, the target behavior scenario includes an obfuscation behavior scenario, and the target static feature includes a target obfuscation feature; among them, the first candidate malicious code determination module can include: a second scanning sub-module, an obfuscated code determination sub-module, and a second candidate malicious code determination sub-module.

[0019] Among them, the second scanning sub-module can be configured to scan the target installation package. The obfuscated code determination sub-module can be configured to determine the code including the target obfuscation feature in the target installation package as the obfuscated code. The second candidate malicious code determination sub-module can be configured to determine the first candidate malicious code according to the obfuscated code.

[0020] In some embodiments, the target behavior scenario includes a system command execution scenario, and the target static feature includes a target system command execution feature; wherein, the first candidate malicious code determination module may include: a third scanning sub-module, a system command execution code determination sub-module, and a third candidate malicious code determination sub-module.

[0021] Among them, the third scanning sub-module may be configured to scan the target installation package. The system command execution code determination sub-module may be configured to determine, in the target installation package, the code including the target system command execution feature as the system command execution code. The third candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the system command execution code.

[0022] In some embodiments, the system command execution code determination sub-module may include: a target language determination unit, a target official library determination unit, a similar system command execution feature determination unit, and a target system command execution feature determination unit.

[0023] Among them, the target language determination unit may be configured to determine the target language of the target installation package. The target official library determination unit may be configured to link the target official library of the target language. The similar system command execution feature determination unit may be configured to supplement the target system command execution feature according to the target official library. The target system command execution feature determination unit may be configured to determine, in the target installation package, the code of the target system command execution feature according to the supplemented target system command execution feature.

[0024] In some embodiments, the target static features include a target black website feature, a target black IP feature, and a target black domain name feature; wherein, the first candidate malicious code determination module may include: a fourth scanning sub-module, a black address code determination sub-module, and a third candidate malicious code determination sub-module.

[0025] Among them, the fourth scanning sub-module may be configured to scan the target installation package. The black address code determination sub-module may be configured to determine, in the target installation package, the code including the target black website feature, the target black IP feature, and the target black domain name feature as the black address code. The third candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the black address code.

[0026] In some embodiments, the first candidate malicious code includes a first target candidate malicious code and a second target candidate malicious code, and the first target candidate malicious code and the second target candidate malicious code correspond to the same target static feature; wherein, the first candidate malicious code determination module may further include: a first hash value acquisition sub-module, a second hash value acquisition sub-module, a target distance determination sub-module, a target distance threshold acquisition sub-module, and a duplicate removal processing sub-module.

[0027] Among them, the first hash value acquisition sub-module may be configured to perform hash encoding processing on the first target candidate malicious code to generate a first hash value of the first target candidate malicious code. The second hash value acquisition sub-module may be configured to perform hash encoding processing on the second target candidate malicious code to generate a second hash value of the second target candidate malicious code. The target distance determination sub-module may be configured to determine a target distance between the first hash value and the second hash value. The target distance threshold acquisition sub-module may be configured to acquire a target scenario parameter of the target behavior scenario and determine a target distance threshold according to the target scenario parameter. The duplicate removal processing sub-module may be configured to perform duplicate removal processing on the first target candidate malicious code and the second target candidate malicious code if the target distance between the first hash value and the second hash value is less than the target distance threshold.

[0028] In some embodiments, the malicious code detection hook function includes a low-level malicious code detection hook function; wherein, the simulation running module may include: a script sandbox running sub-module, and the second candidate malicious code determination module may include a low-level malicious code detection sub-module.

[0029] Among them, the script sandbox running sub-module may be configured to simulate running the suspicious installation package in a target script sandbox and inject the low-level malicious code detection hook function into the suspicious installation package. The low-level malicious code detection sub-module may be configured to determine whether the suspicious code is the second candidate malicious code according to the target static feature during the simulation running process of the suspicious installation package through the low-level malicious code detection hook function. In some embodiments, the suspicious code determination module may include: a second installation package acquisition sub-module and a display sub-module.

[0030] Among them, the second installation package acquisition sub-module can be configured to determine a second malicious code and a second malicious installation package where the second malicious code is located in the first candidate malicious codes according to the malicious code baseline feature. The display sub-module can be configured to, in response to an installation request of a target object for the second malicious installation package, display a target warning work order according to the second malicious code and the second malicious installation package, where the target warning work order includes the second malicious code and a behavior scenario code segment where the second malicious code is located.

[0031] An embodiment of the present disclosure provides an electronic device, which includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the supply chain attack detection method described in any one of the above.

[0032] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the supply chain attack detection method described in any one of the above.

[0033] An embodiment of the present disclosure provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above supply chain attack detection method.

[0034] The supply chain attack detection method, device, electronic device and computer-readable storage medium provided by the embodiments of the present disclosure, on the one hand, perform static scanning on a target installation package through target static features in a target behavior scenario to determine a suspicious installation package from the target installation package; on the other hand, perform dynamic scanning on the suspicious installation package through a malicious code detection hook function to determine a malicious installation package from the suspicious installation package. By combining static scanning and dynamic scanning, the above method can accurately detect malicious installation packages in the target installation package, improve the accuracy of the detection effect, and effectively reduce the false alarm rate of the detection result.

[0035] It should be understood that the above general description and subsequent detailed description are only exemplary and cannot limit the present disclosure. Description of the Drawings

[0036] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. The accompanying drawings described below are only some embodiments of the present disclosure. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 A schematic diagram showing an exemplary system architecture of a supply chain attack method or a supply chain attack device applied to an embodiment of the present disclosure.

[0038] Figure 2 A schematic structural diagram of a computer system applied to a supply chain attack device shown according to an exemplary embodiment.

[0039] Figure 3 A flowchart of a supply chain attack detection method shown according to an exemplary embodiment.

[0040] Figure 4 A schematic diagram of an operation log shown according to an exemplary embodiment.

[0041] Figure 5 A schematic diagram of a code snippet shown according to an exemplary embodiment.

[0042] Figure 6 A schematic diagram of an early warning work order shown according to an exemplary embodiment.

[0043] Figure 7 is Figure 3 A flowchart of step S2 in an exemplary embodiment in

[0044] Figure 8 is Figure 7 A flowchart of step S212 in an exemplary embodiment in

[0045] Figure 9 is Figure 3 A flowchart of step S2 in an exemplary embodiment in

[0046] Figure 10 is Figure 3 A flowchart of step S2 in an exemplary embodiment in

[0047] Figure 11 is Figure 3 A flowchart of step S2 in an exemplary embodiment in

[0048] Figure 12 is Figure 3 A flowchart of step S2 in an exemplary embodiment in

[0049] Figure 13 Yes Figure 3 It is a flowchart of steps S5 and S6 in an exemplary embodiment.

[0050] Figure 14 It is a schematic structural diagram of a supply chain attack detection shown according to an exemplary embodiment.

[0051] Figure 15 It is a schematic diagram of an application scenario of a supply to attack detection shown according to an exemplary embodiment.

[0052] Figure 16 It is a block diagram of a supply chain attack detection device shown according to an exemplary embodiment. Detailed implementation manners

[0053] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted.

[0054] The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will realize that one or more of the specific details can be omitted in practicing the technical solutions of this disclosure, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this disclosure.

[0055] The accompanying drawings are only schematic illustrations of this disclosure, and identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted. Some of the block diagrams shown in the drawings do not necessarily have to correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0056] The flowcharts shown in the accompanying drawings are only illustrative, and do not necessarily include all the contents and steps, nor do they necessarily have to be executed in the order described. For example, some steps can be decomposed, and some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0057] In this specification, the terms "a", "one", "the", "said", and "at least one" are used to indicate the presence of one or more elements / components / etc.; the terms "comprising", "including", and "having" are used to mean an open-ended inclusion, indicating that there may be additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", "third", etc. are only used as labels and do not limit the quantity of their objects.

[0058] The exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0059] Figure 1 FIG. shows a schematic diagram of an exemplary system architecture to which the supply chain attack detection method or supply chain attack detection device according to the embodiments of the present disclosure can be applied.

[0060] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0061] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Among them, the terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, wearable devices, virtual reality devices, smart homes, etc.

[0062] The server 105 may be a server providing various services, such as a background management server that supports the operations performed by users using the terminal devices 101, 102, 103. The background management server may analyze and process data such as requests received, and feedback the processing results to the terminal devices.

[0063] The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, etc. The present disclosure does not limit this.

[0064] Server 105 can, for example, obtain a target installation package; Server 105 can, for example, scan the target installation package according to the target static features in the target behavior scenario, and obtain the first candidate malicious code from the target installation package; Server 105 can, for example, determine the suspicious code among the first candidate malicious codes according to the malicious code baseline features, where the malicious code baseline features are known malicious code features; Server 105 can, for example, determine the suspicious installation package where the suspicious code is located from the target installation package; Server 105 can, for example, inject a malicious code detection hook function into the suspicious installation package and simulate running the suspicious installation package; Server 105 can, for example, during the simulation running process of the suspicious installation package, obtain the second candidate malicious code from the suspicious installation package through the malicious code detection hook function; Server 105 can, for example, determine the first malicious code among the second candidate malicious codes according to the malicious code baseline features; Server 105 can, for example, determine the first malicious installation package where the first malicious code is located.

[0065] It should be understood that Figure 1 the number of terminal devices, networks, and servers in

[0066] The following refers to Figure 2 which shows a schematic structural diagram of a computer system 200 of a terminal device suitable for implementing the embodiments of the present application. Figure 2 The terminal device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0067] As Figure 2 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage section 208 into the random access memory (RAM) 203. In the RAM 203, various programs and data required for the operation of the system 200 are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other through a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.

[0068] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as required. A removable medium 211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 210 as required so that a computer program read therefrom is installed into the storage section 208 as required.

[0069] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 209, and / or installed from the removable medium 211. When the computer program is executed by a central processing unit (CPU) 201, the above-described functions defined in the system of the present application are executed.

[0070] It should be noted that the computer-readable storage medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and this computer-readable storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0071] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0072] The modules and / or sub-modules and / or units and / or sub-units involved in the embodiments of the present application can be implemented in software or in hardware. The described modules and / or sub-modules and / or units and / or sub-units can also be provided in a processor. For example, it can be described as: a processor includes a sending unit, an obtaining unit, a determining unit, and a first processing unit. Among them, the names of these modules and / or sub-modules and / or units and / or sub-units do not constitute a limitation to the modules and / or sub-modules and / or units and / or sub-units themselves in some cases.

[0073] As another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium can be included in the device described in the above embodiments; it can also exist alone without being assembled into the device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the device, the functions that the device can implement include: obtaining a target installation package; scanning the target installation package according to the target static features in the target behavior scenario, and obtaining a first candidate malicious code from the target installation package; determining a suspicious code from the first candidate malicious code according to the malicious code baseline features, where the malicious code baseline features are known malicious code features; determining a suspicious installation package where the suspicious code is located from the target installation package; injecting a malicious code detection hook function into the suspicious installation package and simulating the operation of the suspicious installation package; during the simulated operation of the suspicious installation package, obtaining a second candidate malicious code from the suspicious installation package through the malicious code detection hook function; determining a first malicious code from the second candidate malicious code according to the malicious code baseline features; and determining a first malicious installation package where the first malicious code is located from the suspicious installation package.

[0074] Figure 3 is a flowchart of a supply chain attack detection method shown according to an exemplary embodiment. The method provided by the embodiments of the present disclosure can be executed by any electronic device with computing and processing capabilities. For example, the method can be executed by the server or terminal device in the above Figure 1 embodiments, or can be jointly executed by the server and the terminal device. In the following embodiments, the server is used as an example of the execution subject for illustration, but the present disclosure is not limited thereto.

[0075] Refer to Figure 3 , the supply chain attack detection method provided by the embodiments of the present disclosure may include the following steps.

[0076] In step S1, obtain a target installation package.

[0077] In some embodiments, the target installation package may be an installation package written in a target scripting language, which may be Python (a cross-platform computer programming language), JavaScript (a lightweight, interpreted or just-in-time compiled high-level programming language with function priority, JS), PHP (Hypertext Preprocessor, a general-purpose open-source scripting language), etc., and the present disclosure does not limit this.

[0078] The embodiments of the present disclosure will be described by taking an installation package written in the Python scripting language as an example, but the present disclosure is not limited thereto.

[0079] It can be understood that the obtained target installation package in this embodiment may be one or more, and the present disclosure does not limit this.

[0080] In step S2, the target installation package is scanned according to the target static features in the target behavior scenario, and the first candidate malicious code is obtained from the target installation package.

[0081] In some embodiments, the target behavior scenario may be the behavior scenario when a malicious installation package performs a target malicious operation. The target malicious operation may refer to a malicious attack operation implemented by the malicious installation package through behaviors such as external connection, obfuscation, and system command invocation, or may also be an operation of communicating with the outside through IOC (Indicator of Compromise) in the script to achieve a malicious attack. Therefore, the behavior scenario of the target malicious operation may at least include an external connection behavior scenario, an obfuscation behavior scenario, a system command execution scenario, etc., and the present disclosure does not limit this.

[0082] In some embodiments, in the external connection behavior scenario, the malicious installation package may connect to the outside through the external connection behavior to achieve a malicious operation; in the obfuscation behavior scenario, the malicious installation package may obfuscate and package parameters, code, etc. in the installation package through the obfuscation behavior to achieve a malicious operation, such as converting a string parameter to a hexadecimal parameter, splitting a parameter into a collection of multiple parameters (such as splitting name into na+me), etc.; in the system command execution scenario, the malicious installation package may also call the system command through the system command execution behavior to achieve a malicious operation, for example, in the Python language, the system command is called through the "os.system" code; in the IOC recognition behavior scenario, the malicious installation package may also communicate with the outside through the IOC (Indicator of Compromise) in the script, for example, communicate with the outside through the black address information in the installation package to achieve a malicious operation.

[0083] In some embodiments, the target static feature may refer to a static feature that can reflect the target behavior in the target behavior scenario. In the behavior scenario of a target malicious operation, the target static feature may refer to a suspicious malicious operation feature that can implement the malicious operation, and this suspicious malicious operation feature may be a common coding feature or a pre-specified coding feature in the process of implementing the malicious operation. The target static feature may include high-risk features such as external connection feature, obfuscation feature, system command execution feature, etc., which are relatively common in the process of implementing malicious operations. For example, assume that a malicious installation package implements a malicious operation through an external connection behavior, then the target external connection feature that implements the external connection behavior is the target static feature; assume that a malicious installation package implements a malicious operation through an obfuscation behavior, then the target obfuscation feature that implements the obfuscation behavior is the target static feature; assume that a malicious installation package implements a malicious operation through a system call behavior, then the target system command execution feature that implements the system call behavior can be the target static feature. For example, in the Python language, executing "urllib2.request (a request function)" can implement the behavior of connecting to the outside, so the feature "urllib2.request" can be the target external connection feature in the external connection behavior scenario; for another example, in the Python language, executing "urlencode (a coding function)" can encode parameters to implement the obfuscation behavior, so "urlencode" can be the target obfuscation feature in the obfuscation behavior scenario; for yet another example, in the Python language, executing "os.syetem (a system command call function)" can implement the behavior of calling system commands, so "os.syetem" can be the target system command execution feature in the system command execution scenario.

[0084] The above-mentioned target static features can be obtained by analyzing the malicious operations of known malicious installation packages. For example, a real simulation test can be conducted on a malicious installation package in a special attack scenario such as a supply chain attack to obtain a list of high-risk functions. The list of high-risk functions can be high-risk functions corresponding to system command execution, external connection, file writing behavior, or coding functions (such as base64, urlencode, etc.). Such high-risk functions or code segments corresponding to the high-risk functions (code segments that call the high-risk function or code segments called by the high-risk function) can be used as the above-mentioned target static features.

[0085] In some embodiments, if the target behavior scenario is the behavior scenario of a target malicious operation and the target static feature includes a suspicious feature of the malicious operation, obtaining the first candidate malicious code from the target installation package may include: performing a static scan on the target installation package; determining, in the target installation package, the code including the suspicious feature of the malicious operation as the suspicious code of the malicious operation; and then determining the first candidate malicious code according to the candidate suspicious code of the malicious operation. For example, taking the code segment including the suspicious code of the malicious operation as the first candidate malicious code, etc.

[0086] In step S3, determine the suspicious code from the first candidate malicious code according to the baseline feature of the malicious code, where the baseline feature of the malicious code is a known malicious code feature.

[0087] Among them, the baseline feature of the malicious code may be a known malicious code segment obtained by analyzing known malicious installation packages.

[0088] It can be understood that the above first candidate malicious code may include malicious code or may include normal code. To determine whether the first candidate malicious code is malicious code, it is necessary to compare the first candidate malicious code with the baseline feature of the malicious code and the baseline feature of the non-malicious code. In some embodiments, the first candidate malicious code may be compared with the baseline feature of the malicious code, and the first candidate malicious code with a successful comparison is taken as the second malicious code.

[0089] In some embodiments, if there is a second malicious code in the first candidate malicious code, the installation package where the second malicious code is located is taken as the second malicious installation package. In some embodiments, when the target object issues an installation request for the second malicious installation package, a target warning work order as shown in Figure 4 may be displayed to the target object according to the second malicious code and the second malicious installation package. The target warning work order can not only display the package name, version number, upload time, upload information, etc. of the malicious installation package, but also display the code segment of the behavior scenario where the second malicious code is located.

[0090] Among them, the code segment of the behavior scenario where the second malicious code is located may refer to the code segment of the second malicious code and the Cn lines of code above and below the second malicious code, where Cn is a positive integer greater than or equal to 0.

[0091] In some embodiments, multiple baseline features of non-malicious codes may be determined in advance according to known white samples. For example, determining known white codes from the white samples as the baseline features of the non-malicious codes (generally, the white codes may be known codes without malicious behavior).

[0092] In some embodiments, the non-malicious code baseline features can be compared with the first candidate malicious code, and the code with a successful comparison can be used as non-malicious code.

[0093] In some embodiments, code snippets in the first candidate malicious code that are neither malicious code nor non-malicious code can be used as suspicious code for further judgment of the suspicious code.

[0094] In step S4, the suspicious installation package where the suspicious code is located is determined from the target installation package.

[0095] In step S5, a malicious code detection hook function is injected into the suspicious installation package, and the suspicious installation package is simulated to run.

[0096] In step S6, during the simulated running of the suspicious installation package, the second candidate malicious code is obtained from the suspicious installation package through the malicious code detection hook function.

[0097] In some embodiments, a target script sandbox can be customized in advance at the script layer according to the basic image of the script, and a malicious code detection hook function can be customized in the target script sandbox. The malicious code detection hook function can include both a code-level hook function and a low-level API (Application Programming Interface) level hook function, and the present disclosure does not limit this.

[0098] In some embodiments, when the suspicious installation package is run to the code line including the target static feature through the malicious code detection hook function, the target running log as shown in Figure 5 is output simultaneously. The target running log will identify the code line where the target static feature is located (as shown in Figure 5 , the target running log identifies that there are target static features at line 35 and line 32). Then, according to the target running log, the code where the target static feature is located in the suspicious installation package, that is, the second candidate malicious code, can be determined.

[0099] In some embodiments, since some malicious features cannot be determined only through static scanning, such as malicious strings in hexadecimal representation after being packaged and obfuscated, etc., but the obfuscated malicious string will inevitably reveal its prototype during the running process. Therefore, it can be determined whether the suspicious installation package includes the second candidate malicious code through the malicious code detection hook function during the running process of the suspicious installation package.

[0100] Performing simulation runs through a script sandbox may include the following processes: Import the package name and version number of the suspicious installation package into a message queue. A main control program retrieves the target parameters and distributes the target parameters to each worker process; Call the interface API of the target script sandbox image docker, combine with the target parameters generated by the main control program, and start a container according to the script sandbox image to simulate the installation process of the suspicious installation package. Then, crop the log file recorded at the hook points during the installation process to extract sensitive function code snippets as the second candidate malicious code. By customizing sensitive behavior hook points through the above process, it can effectively solve the problems such as obfuscation and camouflage in malicious code that are difficult to cover by static detection.

[0101] In step S7, determine the first malicious code in the second candidate malicious code according to the malicious code baseline feature.

[0102] In some embodiments, after determining the second candidate malicious code, it is also possible to continue to compare the malicious code baseline feature with the first candidate malicious code, and use the code with successful comparison as the malicious code (hereinafter referred to as the first malicious code).

[0103] In step S8, determine the first malicious installation package where the first malicious code is located from the suspicious installation package.

[0104] In some embodiments, if the first malicious code exists in the second candidate malicious code, then use the installation package where the first malicious code is located as the first malicious installation package. In some embodiments, when the target object issues an installation request for the first malicious installation package, it is possible to display the target warning work order as shown in Figure 4 to the target object according to the first malicious code and the first malicious installation package. The target warning work order may include the behavior scenario code segment where the first malicious code is located as shown in Figure 4

[0105] Among them, the behavior scenario code segment where the first malicious code is located may refer to the code segment of the first malicious code and the Cn lines of code above and below the first malicious code, where Cn is a positive integer greater than or equal to 0.

[0106] The technical solution provided in this embodiment, on the one hand, performs a static scan of the target installation package through the target static features in the target behavior scenario to determine the suspicious installation package from the target installation package; on the other hand, performs a dynamic scan of the suspicious installation package through the malicious code detection hook function to determine the malicious installation package from the suspicious installation package. The above method can accurately detect the malicious installation package in the target installation package through the combination of static scan and dynamic scan.

[0107] Figure 7 Yes Figure 3 ​Flowchart of step S2 in an exemplary embodiment.

[0108] In some embodiments, the target behavior scenario may include an external connection behavior scenario, and the target static feature may include a target external connection feature.

[0109] In some embodiments, the target external connection feature may refer to the code that can implement an external connection behavior (i.e., a behavior of connecting to the outside) in the target installation package, such as "urllib2.request" in the Python language.

[0110] Reference Figure 7 , the above step S2 may include the following process.

[0111] In step S211, scan the target installation package.

[0112] In step S212, determine the code including the target external connection feature in the target installation package as the external connection code.

[0113] In step S213, determine the first candidate malicious code according to the external connection code.

[0114] In some embodiments, the target scenario parameter Cn of the target behavior scenario may be obtained, and the first candidate malicious code may be determined according to the target scenario parameter and the external connection code.

[0115] For example, it can be set that if the target scenario parameter is the first value (for example, set to 0), then the external connection code itself is determined as the first candidate malicious code; it can be set that if the target scenario parameter is the second value (for example, set to N, where N is a positive integer greater than or equal to 1), then the external connection code and the N lines of code above and below the external connection code are used as the first candidate malicious code; it can be set that if the target scenario code is the third value (for example, set to -1), then all the code in the target file where the external connection code is located can be determined as the first candidate malicious code. It can be understood that the target installation package can be composed of multiple target files, and each target file is composed of multiple lines of code. It should be noted that different target scenario parameters will affect the content of the hit code segments. The larger the scenario parameter is set, the more comprehensive the code segment content is, the more accurate the detection result is, and the lower the false alarm rate of the detection; the smaller the scenario parameter is, the smaller the code segment content is, and the relatively lower the operation and processing cost is. In actual operation, technicians can adjust the target scenario parameter according to the actual application scenario. For example, in a scenario with a higher security detection level, the target scenario parameter can be set larger to expand the content of the hit code segments; in a situation with poorer detection resources, the target scenario parameter can be set smaller to minimize the operation cost as much as possible. In some embodiments, the value of N can be set according to the number of the determined first candidate malicious codes, and the number of the determined first candidate malicious codes is positively correlated with N, that is, the more the number of the determined first candidate malicious codes, the higher the risk level of the installation package, then the suspicious features of malicious operations in the code (such as external connection code, system command execution code, etc.) and the more lines of code above and below the suspicious features of malicious operations can be used as the first candidate malicious codes for detection, so as to improve the accuracy of detection and flexibly configure the detection resources.

[0116] The technical solution provided in this embodiment determines the first candidate malicious code in the target installation package through the target external connection feature, so as to detect the supply chain attack according to the first candidate malicious code.

[0117] Figure 8 Yes Figure 7 It is the flowchart of step S212 in an exemplary embodiment.

[0118] In some embodiments, since the target programmer has limited understanding of the programming language, the target programmer may also have limited knowledge of the code that can implement the external connection behavior in the target language. Therefore, the following method can be used to supplement the target external connection feature.

[0119] In step S2121, determine the target language of the target installation package.

[0120] In some embodiments, the target language for writing the target installation package can be determined. The target language can be, for example, the Python language, the JavaScript language, the PHP language, etc., and the present disclosure does not limit this.

[0121] In step S2122, link the target official library of the target language.

[0122] In step S2123, supplement the target external connection feature according to the target official library.

[0123] In some embodiments, the target official library of the target language can be linked to determine, from the target official library, the same-type external connection features that are of the same type as the target external connection feature, such as other external connection features that may implement external connection behaviors set in the target official library, to supplement the target external connection feature.

[0124] In step S2124, determine the code of the target system command execution feature in the target installation package according to the supplemented target system command execution feature.

[0125] The technical solution provided in this embodiment can achieve a supplement to the target external connection feature by linking the target official library of the target language, so as to more accurately obtain the first candidate malicious code from the target installation package.

[0126] Figure 9 Yes Figure 3 It is the flowchart of step S2 in an exemplary embodiment.

[0127] In some embodiments, the target behavior scenario can include an obfuscation behavior scenario, and the target static feature can include a target obfuscation feature.

[0128] In some embodiments, the target obfuscation feature can refer to code that can obfuscate and package parameters, etc. (such as encoding parameters and displaying the parameters through hexadecimal codes) within the target installation package, so that the parameters are not presented in their original form, such as "urlencode" in the Python language.

[0129] Reference Figure 9 , the above step S2 can include the following process.

[0130] In step S221, scan the target installation package.

[0131] In step S222, determine the code including the target obfuscation feature in the target installation package as the obfuscated code.

[0132] In step S223, determine the first candidate malicious code according to the obfuscated code.

[0133] In some embodiments, the target scenario parameter Cn of the target behavior scenario can be obtained, and the first candidate malicious code can be determined according to the target scenario parameter and the obfuscation code.

[0134] For example, it can be set that if the target scenario parameter is the first value (for example, set to 0), then the obfuscation code itself is determined as the first candidate malicious code; it can be set that if the target scenario parameter is the second value (for example, set to N, where N is a positive integer greater than or equal to 1), then the obfuscation code and the N lines of code above and below the obfuscation code are used as the first candidate malicious code; it can be set that if the target scenario code is the third value (for example, set to -1), then all the code in the target file where the obfuscation code is located can be determined as the first candidate malicious code. It can be understood that the target installation package can be composed of multiple target files, and each target file is composed of multiple lines of code.

[0135] In some embodiments, since the target programmer has limited understanding of the target language for writing the target installation package, the target programmer may also have limited knowledge of the code in the target language that can implement the obfuscation behavior. Therefore, the following method can be used to supplement the target obfuscation features: determine the target language of the target installation package; link the target official library of the target language; supplement the target obfuscation features according to the target official library.

[0136] The technical solution provided in this embodiment determines the first candidate malicious code in the target installation package through the target obfuscation features, so as to detect supply chain attacks based on the first candidate malicious code.

[0137] Figure 10 Yes Figure 3 It is the flowchart of step S2 in an exemplary embodiment.

[0138] In some embodiments, the target behavior scenario may include a system command execution scenario, and the target static feature may include a target system command execution feature.

[0139] In some embodiments, the target system command execution feature may refer to the code that can call system commands within the target installation package. For example, the "os.system" code in the python language.

[0140] Reference Figure 10 As described above, step S2 may include the following process.

[0141] In step S231, the target installation package is scanned.

[0142] In step S232, the code including the target system command execution feature is determined in the target installation package as the system command execution code.

[0143] In step S233, determine the first candidate malicious code according to the system command execution code.

[0144] In some embodiments, the target scenario parameter Cn of the target behavior scenario can be obtained, and the first candidate malicious code can be determined according to the target scenario parameter Cn and the system command execution code.

[0145] For example, it can be set that if the target scenario parameter is the first value (for example, set to 0), then the system command execution code itself is determined as the first candidate malicious code; it can be set that if the target scenario parameter is the second value (for example, set to N, where N is a positive integer greater than or equal to 1), then the system command execution code and the N lines of code above and below the system command execution code are used as the first candidate malicious code; it can be set that if the target scenario code is the third value (for example, set to -1), then all the code in the target file where the system command execution code is located can be determined as the first candidate malicious code. It can be understood that the target installation package can be composed of multiple target files, and each target file is composed of multiple lines of code.

[0146] In some embodiments, due to the limited understanding of the programming language by the target programmer, the target programmer may also have limited knowledge of the code in the target language that can implement the system command call behavior. Therefore, the following method can be used to supplement the system command execution code: determine the target language of the target installation package; link the target official library of the target language; supplement the target system command execution characteristics according to the target official library. Finally, the code of the target system command execution characteristics can be determined in the target installation package according to the supplemented target system command execution characteristics.

[0147] The technical solution provided in this embodiment determines the first candidate malicious code in the target installation package through the target system command execution characteristics, so as to detect supply chain attacks according to the first candidate malicious code.

[0148] Figure 11 Yes Figure 3 It is the flowchart of step S2 in an exemplary embodiment.

[0149] In some embodiments, not only can the first candidate malicious code be determined according to the external connection behavior, obfuscation behavior, and system command execution behavior, but also the first candidate malicious code can be determined in the target installation package according to the IOC, where the IOC identifier can include target black website characteristics, target black IP characteristics, target black domain name characteristics, etc.

[0150] Reference Figure 11 The above step S2 may include the following process.

[0151] In step S241, scan the target installation package.

[0152] In step S242, determine the code including the target black website feature, target black IP feature, and target black domain name feature in the target installation package as the black address code.

[0153] In some embodiments, the code including the target black website feature, target black IP feature, and target black domain name feature can be determined from the target installation package through static scanning as the black address code.

[0154] In step S243, determine the first candidate malicious code according to the black address code.

[0155] In some embodiments, the target scenario parameter Cn of the target behavior scenario can be obtained, and the first candidate malicious code can be determined according to the target scenario parameter and the black address code.

[0156] For example, it can be set that if the target scenario parameter is the first value (for example, set to 0), then the black address code itself is determined as the first candidate malicious code; it can be set that if the target scenario parameter is the second value (for example, set to N, where N is a positive integer greater than or equal to 1), then the black address code and the N lines of code above and below the black address code are used as the first candidate malicious code; it can be set that if the target scenario code is the third value (for example, set to -1), then all the code in the target file where the black address code is located is determined as the first candidate malicious code. It can be understood that the target installation package can be composed of multiple target files, and each target file is composed of multiple lines of code.

[0157] The technical solution provided in this embodiment determines the first candidate malicious code in the target installation package through the IOC identifier, so as to detect supply chain attacks according to the first candidate malicious code.

[0158] Figure 12 Yes Figure 3 It is the flowchart of step S2 in an exemplary embodiment.

[0159] In some embodiments, since the same target static feature may hit multiple code fragments, and there may be duplicates among the multiple hit code fragments, it is necessary to deduplicate the multiple code fragments hit by the same target static feature. For example, the above multiple code fragments can be deduplicated through the simhash (text deduplication) algorithm.

[0160] Reference Figure 12 , the process of deduplication through the simhash algorithm can include the following process.

[0161] Assume that the first candidate malicious code includes a first target candidate malicious code and a second target candidate malicious code, where both the first target candidate malicious code and the second target candidate malicious code are hit by the same target static feature.

[0162] In step S251, perform hash encoding processing on the first target candidate malicious code to generate a first hash value of the first target candidate malicious code.

[0163] In some embodiments, the first target candidate malicious code can be tokenized to obtain effective feature vectors; then, hash encoding processing is performed on the effective feature vectors of the first target candidate malicious code to obtain hash values of each token; finally, the hash values of each token are weighted according to the proportion of each token in the first target candidate malicious code to generate a first hash value of the first target candidate malicious code.

[0164] In step S252, perform hash encoding processing on the second target candidate malicious code to generate a second hash value of the second target candidate malicious code.

[0165] In some embodiments, the second target candidate malicious code can be tokenized to obtain effective feature vectors; then, hash encoding processing is performed on the effective feature vectors of the second target candidate malicious code to obtain hash values of each token; finally, the hash values of each token are weighted according to the proportion of each token in the second target candidate malicious code to generate a second hash value of the second target candidate malicious code.

[0166] In step S253, determine the target distance between the first hash value and the second hash value.

[0167] In some embodiments, dimensionality reduction processing is performed on the first hash value (for example, values greater than 0 are set to 1, and values less than or equal to 0 are set to 0) to obtain the simhash value of the first target candidate malicious code.

[0168] In some embodiments, dimensionality reduction processing is performed on the second hash value (for example, values greater than 0 are set to 1, and values less than or equal to 0 are set to 0) to obtain the simhash value of the second target candidate malicious code.

[0169] In step S254, obtain the target scenario parameters of the target behavior scenario and determine the target distance threshold according to the target scenario parameters.

[0170] In some embodiments, the target distance threshold can be determined according to the target scenario parameters of the target behavior scenario. For example, when the target scenario parameters are small, it means that the fragment of the first candidate malicious code is small, and at this time, it is required that the target distance threshold should also be small accordingly; for example, when the target scenario is large, it means that the fragment of the first candidate malicious code is large, and at this time, it is required that the target distance threshold should also be large accordingly. By using the target distance threshold determined according to the target scenario parameters above, potential risk misreports can be avoided.

[0171] In step S255, if the target distance between the first hash value and the second hash value is less than the target distance threshold, then duplicate removal processing is performed on the first target candidate malicious code and the second target candidate malicious code.

[0172] In some embodiments, the target distance between the simhash value of the first target candidate malicious code and the simhash value of the second target candidate malicious code can be calculated. If the target distance between the simhash value of the first target candidate malicious code and the simhash value of the second target candidate malicious code is less than the target distance threshold, it can be considered that the first target candidate malicious code and the second candidate malicious code are duplicates, and then duplicate removal processing needs to be performed on the first target candidate malicious code and the second target candidate malicious code.

[0173] The above target distance can refer to Hamming distance, Euclidean distance, or Hamming distance, etc., and the present disclosure does not limit this.

[0174] The technical solution provided in this embodiment can perform duplicate removal processing on duplicate first candidate malicious codes through hash encoding, so as to reduce the calculation amount and improve the efficiency of supply chain attack detection.

[0175] Figure 13 Yes Figure 3 It is the flowchart of steps S5 and S6 in an exemplary embodiment.

[0176] In some embodiments, the malicious code detection hook function can include both a code-level hook function and a low-level malicious code detection hook function, and the present disclosure does not limit this.

[0177] Reference Figure 13 , the above step S5 can include the following process.

[0178] In step S51, the suspicious installation package injected with the low-level malicious code detection hook function is simulated and run in the target script sandbox.

[0179] Among them, the low-level malicious code detection hook function is a function for code detection of library functions running at the low level.

[0180] In some embodiments, the target script sandbox can be pre-customized to customize the underlying malicious code detection hook function in the target script sandbox.

[0181] Reference Figure 13 , the above step S6 may include the following process.

[0182] In step S61, through the underlying malicious code detection hook function, determine whether the suspicious code is the second candidate malicious code according to the target static feature.

[0183] The following will explain how a function completes the underlying call in combination with specific examples.

[0184] For example, for the process of calling a certain link through the requst.get() function, its underlying call can be divided into the following five steps:

[0185] 1. Top-level call: request.get(“http: / / www.baidu.com”) / / Initiate a network request to an external server.

[0186] 2. One level down: request.get(‘get’, url, params=params, **kwarges) / / Supplement parameters for the network request

[0187] 3. Two levels down: resp = self.send(prep, **send_kwarge) / / Send the request with parameters

[0188] 4. Three levels down: r = adapter.send(request, **kwarge) / / Call the send of the adapter at the second level to send the request

[0189] 5. Four levels down: rep = conn.urlopen( / / Call the urlopen method of conn at the fourth level to send the request

[0190] method = request.method,

[0191] url = url,

[0192] body = request.body,

[0193] headers = request.headers,

[0194] redirect = False,

[0195] assert_same_host = False,

[0196] preload_content = False,

[0197] decode_content = False,

[0198] retries = self.max_retries,

[0199] timeout = timeout)

[0200] urllib3.connection.HTTPConnection

[0201] http.client.HTTPConnection.request / / Call the'request' method of the 'HTTPConnection' class in the 'http.client' block

[0202] After N - layer call tracking, finally call the'request' method of the 'HTTPConnection' class in the 'http.client' block. This method is the call of the library function provided by the official Python library, that is, the so - called low - level API call.

[0203] According to the above idea, the library functions in the official Python library can be hooked.

[0204] The technical solution provided in this embodiment can more accurately determine whether the suspicious code is the second - candidate malicious code by means of the low - level malicious code detection hook function and dynamically running the suspicious installation package.

[0205] Figure 14 It is a schematic structural diagram of a supply - chain attack detection shown according to an exemplary embodiment.

[0206] Such as Figure 14As shown in the figure, the supply chain detection attack method may include the following steps: 1. Obtain the target installation package from the official source, and perform a static scan on the target installation package through a static feature extraction set generated in advance from the target static features; 2. Obtain the first candidate suspicious code from the target installation package; 3. Transmit the first candidate suspicious code to the analysis engine; 4. Analyze the first candidate suspicious code through the non-malicious code baseline features to determine the non-malicious code, and store the non-malicious code in the baseline DB (DoggaByte, data storage unit) to update the non-malicious code baseline features; 5. Analyze the first candidate suspicious code through the malicious code baseline features to determine the second malicious code and its behavior scenario segments (i.e., black results); 6. Issue a work order warning according to the second malicious code and its behavior scenario segments; 7. Put the target installation package corresponding to the code that is neither non-malicious code nor malicious code in the first candidate malicious code into the script sandbox for running and detection; 8. Obtain the detection log returned by the script sandbox running, and determine the second candidate malicious code according to the detection log; 9. Return the second candidate malicious code to the analysis engine so that the analysis engine can perform analysis and judgment according to the malicious code baseline features.

[0207] Figure 15 It is a schematic diagram of an application scenario for supply attack detection shown according to an exemplary embodiment.

[0208] The technical solution provided in this embodiment can be applied to the security protection within each company. This disclosure can be used for the security construction of internal source application scenarios. The following will take the deployment of internal software of a certain company as an example for explanation, and the specific process is as follows.

[0209] 1. The security side obtains the third-party package (i.e., the target installation package) from the official source, and performs a scan according to the technical solution provided in the embodiment of this disclosure for supply chain attack detection.

[0210] 2. Determine the malicious code (including the first malicious code and the second malicious code) according to the scan result, and regard the file where the malicious code is located as a risk file to generate a risk file list according to the risk file.

[0211] 3. Synchronize the installation package of the official source with the mirror source.

[0212] 4. The user sends a request to the local source to install the target third-party package.

[0213] 5. In response to the user's installation request for the target third-party package, the mirror source calls the risk file list (determined by the malicious code baseline features) to generate an interception rule to intercept the third-party package, so that the user can determine whether to install the target third-party package according to the interception result.

[0214] 6. Store the user's download log through the log storage unit.

[0215] 7. Security auditors conduct security audits on users' download records.

[0216] In the above process, by deploying a mirror proxy source within the company and directly synchronizing with major official source repositories in real time, through internal process management, it is stipulated that all internal downloads are completed through the internal mirror proxy source. At the same time, the security side conducts real-time scans on third-party packages (i.e., target installation packages) of major synchronized target official source repositories, synchronizes the discovered risks to the internal mirror proxy source, provides the ability to intercept risks during the process through the perception of risks in advance (scanning official repositories), and provides traceability and auditing capabilities in cooperation with the third-party package download logs of the internal mirror source.

[0217] Figure 16 It is a block diagram of a supply chain attack detection device shown according to an exemplary embodiment. Refer to Figure 16 , the supply chain attack detection device 1600 provided by the embodiments of the present disclosure may include: an installation package acquisition module 1601, a first candidate malicious code determination module 1602, a suspicious code determination module 1603, a suspicious installation package determination module 1604, a simulation running module 1605, a second candidate malicious code determination module 1606, a first malicious code determination module 1607, and a first malicious installation package determination module 1608.

[0218] Among them, the installation package acquisition module 1601 may be configured to acquire a target installation package. The first candidate malicious code determination module 1602 may be configured to scan the target installation package according to the target static features in the target behavior scenario and obtain the first candidate malicious code from the target installation package. The suspicious code determination module 1603 may be configured to determine suspicious codes in the first candidate malicious codes according to the malicious code baseline features, and the malicious code baseline features are known malicious code features. The suspicious installation package determination module 1604 may be configured to determine the suspicious installation package where the suspicious code is located from the target installation package. The simulation running module 1605 may be configured to inject a malicious code detection hook function into the suspicious installation package and simulate the running of the suspicious installation package. The second candidate malicious code determination module 1606 is configured to obtain the second candidate malicious code from the suspicious installation package through the malicious code detection hook function during the simulation running of the suspicious installation package. The first malicious code determination module 1607 may be configured to determine the first malicious code in the second candidate malicious codes according to the malicious code baseline features. The first malicious installation package determination module 1608 may be configured to determine the first malicious installation package where the first malicious code is located from the suspicious installation package.

[0219] In some embodiments, the target behavior scenario includes the behavior scenario of a target malicious operation, and the target static feature includes the suspicious feature of the malicious operation; wherein, the first candidate malicious code determination module may include: an installation package scanning sub-module, a suspicious malicious operation code determination sub-module, and a first candidate malicious code determination sub-module.

[0220] Wherein, the installation package scanning sub-module may be configured to scan the target installation package; the suspicious malicious operation code determination sub-module may be configured to determine, in the target installation package, the code including the suspicious feature of the malicious operation as the suspicious malicious operation code; and the first candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the suspicious malicious operation code.

[0221] In some embodiments, the target behavior scenario includes an external connection behavior scenario, and the target static feature includes a target external connection feature; wherein, the first candidate malicious code determination module 1602 may include: a first scanning sub-module, an external connection code determination sub-module, and a first candidate malicious code determination sub-module.

[0222] Wherein, the first scanning sub-module may be configured to scan the target installation package. The external connection code determination sub-module may be configured to determine, in the target installation package, the code including the target external connection feature as the external connection code. The first candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the external connection code.

[0223] In some embodiments, the first candidate malicious code determination sub-module may include: a target scenario parameter acquisition unit and a first candidate malicious code determination unit.

[0224] Wherein, the target scenario parameter acquisition unit may be configured to acquire the target scenario parameters of the target behavior scenario. The first candidate malicious code determination unit may be configured to determine the first candidate malicious code according to the target scenario parameters and the external connection code.

[0225] In some embodiments, the target installation package may include a target file; wherein, the first candidate malicious code determination unit may include: a first value processing sub-unit, a second value processing sub-unit, and a third value processing sub-unit.

[0226] Among them, the first value processing subunit may be configured to, if the target scenario parameter is the first value, determine the external connection code as the first candidate malicious code. The second value processing subunit may be configured to, if the target scenario parameter is the second value, use the external connection code and the codes on the upper and lower lines of the target scenario parameter as the first candidate malicious code. The third value processing subunit may be configured to, if the target scenario code is the third value, determine the first candidate malicious code according to the target file where the external connection code is located.

[0227] In some embodiments, the target behavior scenario includes an obfuscation behavior scenario, and the target static feature includes a target obfuscation feature; among them, the first candidate malicious code determination module 1602 may include: a second scanning sub-module, an obfuscated code determination sub-module, and a second candidate malicious code determination sub-module.

[0228] Among them, the second scanning sub-module may be configured to scan the target installation package. The obfuscated code determination sub-module may be configured to determine, in the target installation package, the code including the target obfuscation feature as the obfuscated code. The second candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the obfuscated code.

[0229] In some embodiments, the target behavior scenario includes a system command execution scenario, and the target static feature includes a target system command execution feature; among them, the first candidate malicious code determination module 1602 may include: a third scanning sub-module, a system command execution code determination sub-module, and a third candidate malicious code determination sub-module.

[0230] Among them, the third scanning sub-module may be configured to scan the target installation package. The system command execution code determination sub-module may be configured to determine, in the target installation package, the code including the target system command execution feature as the system command execution code. The third candidate malicious code determination sub-module may be configured to determine the first candidate malicious code according to the system command execution code.

[0231] In some embodiments, the system command execution code determination sub-module may include: a target language determination unit, a target official library determination unit, a similar system command execution feature determination unit, and a target system command execution feature determination unit.

[0232] Among them, the target language determination unit can be configured to determine the target language of the target installation package. The target official library determination unit can be configured to link the target official library of the target language. The same type of system command execution feature determination unit can be configured to supplement the target system command execution feature according to the target official library. The target system command execution feature determination unit can be configured to determine the code of the target system command execution feature in the target installation package according to the supplemented target system command execution feature.

[0233] In some embodiments, the target static features include target black website features, target black IP features, and target black domain features; among them, the first candidate malicious code determination module 1602 may include: a fourth scanning sub-module, a black address code determination sub-module, and a third candidate malicious code determination sub-module.

[0234] Among them, the fourth scanning sub-module can be configured to scan the target installation package. The black address code determination sub-module can be configured to determine the code including the target black website feature, target black IP feature, and target black domain feature in the target installation package as the black address code. The third candidate malicious code determination sub-module can be configured to determine the first candidate malicious code according to the black address code.

[0235] In some embodiments, the first candidate malicious code includes a first target candidate malicious code and a second target candidate malicious code, and the first target candidate malicious code and the second target candidate malicious code correspond to the same target static feature; among them, the first candidate malicious code determination module 1602 may further include: a first hash value acquisition sub-module, a second hash value acquisition sub-module, a target distance determination sub-module, a target distance threshold acquisition sub-module, and a duplicate removal processing sub-module.

[0236] Among them, the first hash value acquisition sub-module can be configured to perform hash encoding processing on the first target candidate malicious code to generate the first hash value of the first target candidate malicious code. The second hash value acquisition sub-module can be configured to perform hash encoding processing on the second target candidate malicious code to generate the second hash value of the second target candidate malicious code. The target distance determination sub-module can be configured to determine the target distance between the first hash value and the second hash value. The target distance threshold acquisition sub-module can be configured to acquire the target scenario parameter of the target behavior scenario and determine the target distance threshold according to the target scenario parameter. The duplicate removal processing sub-module can be configured to perform duplicate removal processing on the first target candidate malicious code and the second target candidate malicious code if the target distance between the first hash value and the second hash value is less than the target distance threshold.

[0237] In some embodiments, the malicious code detection hook function includes a low-level malicious code detection hook function; wherein, the simulation running module 1605 may include: a script sandbox running sub-module, and the second candidate malicious code determination module may include a low-level malicious code detection sub-module.

[0238] Wherein, the script sandbox operation sub-module may be configured to simulate running the suspicious installation package in a target script sandbox and inject the low-level malicious code detection hook function into the suspicious installation package. The low-level malicious code detection sub-module may be configured to determine whether the suspicious code is the second candidate malicious code according to the target static feature during the simulation running process of the suspicious installation package through the low-level malicious code detection hook function. In some embodiments, the suspicious code determination module 1603 may include: a second installation package acquisition sub-module and a display sub-module.

[0239] Wherein, the second installation package acquisition sub-module may be configured to determine a second malicious code and the second malicious installation package where the second malicious code is located in the first candidate malicious codes according to the malicious code baseline feature. The display sub-module may be configured to, in response to an installation request of a target object for the second malicious installation package, display a target warning work order according to the second malicious code and the second malicious installation package, and the target warning work order includes the second malicious code and the behavior scenario code segment where the second malicious code is located.

[0240] Since each functional module of the supply chain attack detection device 1600 of the exemplary embodiments of the present disclosure corresponds to the steps of the exemplary embodiments of the above supply chain attack detection method, details thereof will not be described herein again.

[0241] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solution of the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computing device (which can be a personal computer, a server, a mobile terminal, or a smart device, etc.) to execute the method according to the embodiments of the present disclosure, such as Figure 3 one or more of the steps shown.

[0242] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0243] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not claimed by the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0244] It should be understood that the present disclosure is not limited to the detailed structures, drawing methods, or implementation methods shown herein. On the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A supply chain attack detection method, characterized in that, Including: Obtain a target installation package; Scan the target installation package according to target static features in a target behavior scenario, and obtain first candidate malicious codes from the target installation package; Determine suspicious codes among the first candidate malicious codes according to malicious code baseline features, where the malicious code baseline features are known malicious code features; Determine a suspicious installation package where the suspicious codes are located from the target installation package; Inject a malicious code detection hook function into the suspicious installation package and simulate running the suspicious installation package; During the simulated running of the suspicious installation package, obtain second candidate malicious codes from the suspicious installation package through the malicious code detection hook function; Determine first malicious codes among the second candidate malicious codes according to the malicious code baseline features; Determine a first malicious installation package where the first malicious codes are located.

2. The method according to claim 1, characterized in that, The target behavior scenario includes a behavior scenario of a target malicious operation, and the target static features include suspicious features of malicious operations; among them, scanning the target installation package according to the target static features in the target behavior scenario and obtaining first candidate malicious codes from the target installation package includes: Scan the target installation package; Determine codes including the suspicious features of malicious operations in the target installation package as suspicious codes of malicious operations; Determine the first candidate malicious codes according to the suspicious codes of malicious operations.

3. The method according to claim 1, characterized in that, The target behavior scenario includes an external connection behavior scenario, and the target static features include target external connection features; among them, scanning the target installation package according to the target static features in the target behavior scenario and obtaining first candidate malicious codes from the target installation package includes: Scan the target installation package; Determine codes including the target external connection features in the target installation package as external connection codes; Determine the first candidate malicious codes according to the external connection codes.

4. The method according to claim 3, characterized in that Determining the first candidate malicious codes according to the external connection codes includes: Obtain target scenario parameters of the target behavior scenario; Determine the first candidate malicious codes according to the target scenario parameters and the external connection codes.

5. The method according to claim 4, wherein The target installation package includes a target file; among them, determining the first candidate malicious codes according to the target scenario parameters and the external connection codes includes: If the target scenario parameter is a first value, determine the external connection code as the first candidate malicious code; If the target scenario parameter is a second value, take the external connection code, the N lines of code above the external connection code, and the N lines of code below the external connection code as the first candidate malicious codes; where the second value is equal to N, and N is a positive integer greater than or equal to 1; If the target scenario parameter is a third value, determine the first candidate malicious codes according to the target file where the external connection code is located.

6. The method according to claim 1, wherein The target behavior scenario includes an obfuscation behavior scenario, and the target static features include target obfuscation features; among them, scanning the target installation package according to the target static features in the target behavior scenario and obtaining first candidate malicious codes from the target installation package includes: Scan the target installation package; Identify the code including the target obfuscation feature in the target installation package as the obfuscated code; Determine the first candidate malicious code according to the obfuscated code.

7. The method according to claim 1, wherein The target behavior scenario includes a system command execution scenario, and the target static feature includes a target system command execution feature; wherein, scanning the target installation package according to the target static feature in the target behavior scenario and obtaining the first candidate malicious code from the target installation package includes: Scan the target installation package; Identify the code including the target system command execution feature in the target installation package as the system command execution code; Determine the first candidate malicious code according to the system command execution code.

8. The method according to claim 7, wherein Identifying the code including the target system command execution feature in the target installation package as the system command execution code includes: Determine the target language of the target installation package; Link the target official library of the target language; Supplement the target system command execution feature according to the target official library; Identify the code of the target system command execution feature in the target installation package according to the supplemented target system command execution feature.

9. The method according to claim 1, wherein The target static features include a target black website feature, a target black IP feature, and a target black domain name feature; wherein, scanning the target installation package according to the target static feature in the target behavior scenario and obtaining the first candidate malicious code from the target installation package includes: Scan the target installation package; Identify the code including the target black website feature, the target black IP feature, and the target black domain name feature in the target installation package as the black address code; Determine the first candidate malicious code according to the black address code.

10. The method according to any one of claims 1 to 9, characterized in that, The first candidate malicious code includes a first target candidate malicious code and a second target candidate malicious code, and the first target candidate malicious code and the second target candidate malicious code correspond to the same target static feature; wherein, scanning the target installation package according to the target static feature in the target behavior scenario and obtaining the first candidate malicious code from the target installation package further includes: Perform hash encoding processing on the first target candidate malicious code to generate a first hash value of the first target candidate malicious code; Perform hash encoding processing on the second target candidate malicious code to generate a second hash value of the second target candidate malicious code; Determine the target distance between the first hash value and the second hash value; Obtain the target scenario parameter of the target behavior scenario and determine the target distance threshold according to the target scenario parameter; If the target distance between the first hash value and the second hash value is less than the target distance threshold, perform deduplication processing on the first target candidate malicious code and the second target candidate malicious code.

11. The method according to claim 1, characterized in that, The malicious code detection hook function includes a low-level malicious code detection hook function; wherein, injecting the malicious code detection hook function into the suspicious installation package and simulating the running of the suspicious installation package includes: Simulate the execution of the suspicious installation package in the target script sandbox and inject the underlying malicious code detection hook function into the suspicious installation package; Among them, during the simulation of the suspicious installation package, obtaining the second candidate malicious code from the suspicious installation package through the malicious code detection hook function includes: During the simulation of the suspicious installation package, determining whether the suspicious code is the second candidate malicious code according to the target static features through the underlying malicious code detection hook function.

12. The method according to claim 1, characterized in that, Determining suspicious code from the first candidate malicious code according to the malicious code baseline features includes: Determining the second malicious code and the second malicious installation package where the second malicious code is located from the first candidate malicious code according to the malicious code baseline features; In response to the installation request of the target object for the second malicious installation package, display a target warning work order according to the second malicious code and the second malicious installation package, where the target warning work order includes the second malicious code and the behavior scenario code segment where the second malicious code is located.

13. A supply chain attack detection device, characterized in that, Include: An installation package acquisition module configured to acquire a target installation package; A first candidate malicious code determination module configured to scan the target installation package according to the target static features in the target behavior scenario and obtain the first candidate malicious code from the target installation package; A suspicious code determination module configured to determine suspicious code from the first candidate malicious code according to the malicious code baseline features, where the malicious code baseline features are known malicious code features; A suspicious installation package determination module configured to determine the suspicious installation package where the suspicious code is located from the target installation package; A simulation execution module configured to inject a malicious code detection hook function into the suspicious installation package and simulate the execution of the suspicious installation package; A second candidate malicious code determination module configured to obtain the second candidate malicious code from the suspicious installation package through the malicious code detection hook function during the simulation of the suspicious installation package; A first malicious code determination module configured to determine the first malicious code from the second candidate malicious code according to the malicious code baseline features; A first malicious installation package determination module configured to determine the first malicious installation package where the first malicious code is located.

14. An electronic device, characterized in that, Include: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-12.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-12 is implemented.

Citation Information

Patent Citations

  • Method and device for automatically processing malicious code sample

    CN103761481A

  • Rule-based detection method of ATP attack behavior

    CN105376245A