Malicious package detection method and system for npm software package

Through a multi-level method of static rule base, string point analysis, dynamic analysis and ChatGPT verification, the false alarm and missed response problems of npm software package malicious packet detection in the existing technology are solved, and efficient and accurate malicious packet detection is achieved.

CN120145381APending Publication Date: 2025-06-13WUHAN UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510302271.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When detecting malicious packets in npm software packages, the existing technology has high false alarm rates and high false alarm rates, and it is impossible to effectively detect malicious code with bypass and obfuscation.

Method used

A static rule base is used to initially filter suspicious malicious packages, combine string-based point analysis methods and dynamic analysis schemes to further verify maliciousness, and use the large language model ChatGPT for verification and rule updates.

Benefits of technology

It realizes efficient detection of malicious packages in npm packages, with low false positive rate and high accuracy, and can detect new attacks and complex obfuscation packets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145381A_ABST
    Figure CN120145381A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious package detection method and system for npm software packages. The method comprises the following steps that suspicious malicious packages and confused software packages are preliminarily filtered out in a large sample data set through efficient static rule matching; the time-consuming taint analysis based on the character string further shrinks the range in the rule-matched suspicious malicious packet result; suspicious samples of two-step filtering results are submitted to ChatGPT through a constructed prompt to verify the maliciousness of the suspicious samples, and meanwhile new malicious features are learned to update existing static rules; and for the confused software package detected by static rule matching, the maliciousness of the confused software package is confirmed by adopting a dynamic analysis mode. According to the method, the malicious package in the npm software package can be efficiently detected, the false alarm rate is low, the accuracy rate is high, and new attacks can be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security detection technology, and more specifically, to a malicious package detection method and system for npm software packages. Background Art

[0002] The software supply chain is the core of modern software development, covering the entire process from code writing, testing, publishing to final delivery. With the increasing complexity of technology, the software development process is increasingly dependent on a large number of external components, especially open source code. The report pointed out that in the past five years, the proportion of open source code in software development has soared from 40% to 78%-90%, with an average of 147 open source components per application. In areas such as the Internet of Things and network security, the application of open source code is almost fully covered, which means that the construction of modern software has become heavily dependent on these open source resources.

[0003] This phenomenon brings convenience but also provides attackers with a new attack surface. Due to the lack of security audits on third-party libraries or management warehouse codes, security incidents in the software supply chain have occurred frequently in recent years, among which third-party package management warehouses have become an important target for attackers to pay attention to and carry out poisoning attacks.

[0004] Generally speaking, poisoning attackers mainly implant malicious code in three stages of the use of the open source software supply chain: 1. Package installation stage; 2. Package introduction stage; 3. Package function running stage. In order to ensure the security of the software supply chain, some security detection tools for the software supply chain have emerged, but the existing tools have a large number of false positives and missed positives. Most tools rely heavily on predefined rules for static detection and cannot detect bypassed and obfuscated code. The current identification scheme based on dynamic analysis mainly detects network behavior or file reading and writing behavior outside the package, which introduces a lot of unnecessary noise, resulting in a large number of false positives. Summary of the invention

[0005] An object of the present invention is to address the deficiencies of the prior art and provide a malicious package detection method for npm software packages. The method can efficiently detect malicious packages in npm software packages. Compared with the prior art, the method has the advantages of being applicable to large samples, having a low false alarm rate, a high accuracy rate, and being able to detect new attacks.

[0006] In order to achieve the above object, the first aspect of the present invention provides a malicious package detection method for npm software package, comprising:

[0007] Preliminarily filter out suspicious malicious packages and obfuscated software packages from a large sample data set based on a static rule base, where the static rule base is pre-built based on the features contained in different types of malicious packages;

[0008] Use a string-based point analysis method to further analyze the maliciousness of the filtered suspicious malicious packages;

[0009] Submit the suspicious malicious packages obtained from the preliminary filtering and the suspicious malicious packages obtained from the further analysis to a pre-constructed large model to further verify their maliciousness, and at the same time update the newly learned malicious features to the existing static rules;

[0010] For obfuscated software packages, use a dynamic analysis scheme to verify their maliciousness;

[0011] Integrate the malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis scheme as the detection results.

[0012] In one implementation, based on a static rule library, preliminarily filter out suspicious malicious packages and obfuscated software packages from a large sample dataset, including:

[0013] Perform string-based rule matching on the package.json of npm software packages to detect whether there are suspicious behaviors. For npm packages with suspicious behaviors, record them as suspicious malicious packages and mark the suspicious behaviors they satisfy in the results;

[0014] Scan the code files of npm software packages to detect whether they conform to obfuscation features or encoding features. If they do, perform corresponding decoding attempts according to the encoding features or call the de-obfuscation interface according to the obfuscation features. For code files that cannot be successfully decoded or de-obfuscated, record the corresponding npm software packages as obfuscated software packages. Otherwise, replace the original code file with the decoded or de-obfuscated code file and perform the next static matching process. The static matching process includes: scanning the replaced code file to detect whether there are suspicious behaviors. If there are, record the corresponding npm software package as a suspicious malicious package and mark the suspicious behaviors it satisfies in the results.

[0015] In one implementation, scan the code files of npm software packages to detect whether they conform to obfuscation features or encoding features, including:

[0016] Detect whether there are encoding features in the code file through regular expression rules;

[0017] Judge whether the code file conforms to the obfuscation feature by the occurrence frequency of the obfuscation feature string in the code file.

[0018] In one implementation, use a string-based point analysis method to further analyze the maliciousness of the filtered suspicious malicious packages, including:

[0019] Construct taint analysis rules that conform to malicious behaviors for some static rules in the pre - constructed static rule library;

[0020] Further analyze the maliciousness of the filtered suspicious malicious packages based on the taint analysis rules that conform to malicious behaviors.

[0021] In one implementation, further analyzing the maliciousness of the filtered suspicious malicious packages based on the taint analysis rules that conform to malicious behaviors includes:

[0022] Match the filtered suspicious malicious packages with the taint analysis rules that conform to malicious behaviors. If the match is successful, further confirm them as malicious packages; otherwise, eliminate the suspicious malicious packages.

[0023] In one implementation, submit the suspicious malicious packages obtained from the further analysis to the pre - constructed large - model to further verify their maliciousness, and at the same time update the learned new malicious features to the existing static rules, including:

[0024] Classify the malicious package types into six categories: information leakage, backdoor trojan, encryption ransomware, reverse shell, mining virus, and repository pollution. And submit the corresponding code examples of each type of malicious package to ChatGPT respectively to construct a prompt, requiring to first confirm the maliciousness of the submitted npm software package, and give its malicious package classification and the judgment basis for considering it as this type of malicious package;

[0025] For the suspicious malicious packages obtained from the preliminary filtering and the further analysis, conduct maliciousness verification analysis based on ChatGPT, and according to the maliciousness judgment basis output by ChatGPT, check whether it is an existing rule in the existing static rule library. If not, add it to the existing static rule library.

[0026] In one implementation, for the obfuscated software packages, adopt a dynamic analysis scheme to verify their maliciousness, including:

[0027] The obfuscated software packages obtained from the preliminary filtering are dynamically run in a sandbox environment;

[0028] Use a preset tool to monitor the runtime behaviors of the processes during the running of the npm software package;

[0029] According to the monitored runtime behaviors, detect whether there are malicious features in the npm software package to verify its maliciousness.

[0030] Based on the same inventive concept, the second aspect of the present invention provides a malicious package detection system for npm software packages, including:

[0031] A static rule matching module for preliminarily filtering out suspicious malicious packages and obfuscated software packages from a large sample dataset based on a static rule library, where the static rule library is pre-constructed according to the characteristics included in different types of malicious packages;

[0032] A string-based taint analysis module for further analyzing the maliciousness of the filtered suspicious malicious packages using a string-based taint analysis method;

[0033] A ChatGPT verification and rule update module for submitting the suspicious malicious packages obtained from the preliminary filtering and the suspicious malicious packages obtained from the further analysis to a pre-constructed large model to further verify their maliciousness, and at the same time updating the learned new malicious features to the existing static rules;

[0034] A dynamic analysis module for verifying the maliciousness of the obfuscated software packages using a dynamic analysis scheme;

[0035] A detection result obtaining module for integrating the malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis scheme as the detection results.

[0036] Based on the same inventive concept, a third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the malicious package detection method for npm software packages described in the first aspect.

[0037] Based on the same inventive concept, a fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the malicious package detection method for npm software packages described in the first aspect.

[0038] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0039] The present invention provides a malicious package detection method for npm software packages. Starting from the characteristics of six major types of malicious packages, suspicious samples (suspicious malicious packages and obfuscated software packages) are efficiently screened out from the input large sample data set through the rule matching of the static rule base, with high recall rate. Under the consideration of low false positives, the string-based taint analysis scheme is used as a supplement to the rule matching method to further reduce false positives. At the same time, a verification method based on a large language model ChatGPT is introduced to reduce the cost of manual screening, further reduce the false positive rate, and enable the large language model to empower our self-learning rules, thereby detecting new attacks. In order to further make up for the deficiency of detecting obfuscated malicious packages, a dynamic scheme is proposed to deal with complex obfuscated packages in a targeted manner. The present invention can efficiently detect malicious packages in npm software packages, and has the characteristics of being suitable for large samples, low false positive rate, high accuracy, and being able to detect new attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 This is a flow chart of a malicious package detection method for npm software packages according to an embodiment of the present invention;

[0042] Figure 2 An overall framework flow chart of a malicious package detection method for npm software packages according to an embodiment of the present invention;

[0043] Figure 3 The following is a flow chart of detecting and deobfuscating obfuscated software packages in an embodiment of the present invention;

[0044] Figure 4 It is a dynamic analysis flow chart in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Embodiment 1

[0047] This embodiment discloses a method for detecting malicious packages for npm packages. Please refer to Figure 1 , including:

[0048] S1: Based on a static rule library, initially filter out suspicious malicious packages and obfuscated packages from a large sample dataset. Among them, the static rule library is pre-constructed according to the characteristics included in different types of malicious packages;

[0049] S2: Adopt a string-based taint analysis method to further analyze the maliciousness of the filtered suspicious malicious packages;

[0050] S3: Submit the suspicious malicious packages initially filtered and the suspicious malicious packages obtained from further analysis to a pre-constructed large model to further verify their maliciousness, and at the same time update the newly learned malicious features to the existing static rules;

[0051] S4: For obfuscated packages, adopt a dynamic analysis scheme to verify their maliciousness;

[0052] S5: Integrate the malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis scheme as the detection results.

[0053] Specifically, step S1 is to perform rule matching between the pre-constructed static rule library and the packages in the large sample dataset, and initially screen out suspicious malicious packages and obfuscated packages according to the matching situation. S2 adopts a string-based taint analysis method to further screen the suspicious malicious packages, and then through S3, further verify the suspicious malicious packages obtained in S1 and S2, and give a maliciousness classification for manual review. S4 adopts a dynamic analysis scheme for obfuscated packages to confirm whether they are malicious packages. Finally, integrate the malicious package results confirmed by S3 and S4 to obtain the final detection results. In addition, record the reasons for determining their maliciousness and output a report to the user.

[0054] In one implementation, S1 can be implemented in the following way:

[0055] S1.1: Perform string-based rule matching on the package.json of the npm package to detect whether there are suspicious behaviors. For npm packages with suspicious behaviors, record them as suspicious malicious packages and mark the suspicious behaviors they satisfy in the results;

[0056] S1.2: Scan the code files of the npm packages to detect whether they conform to the obfuscation characteristics or encoding characteristics. If they do, perform corresponding decoding attempts according to the encoding characteristics or call the de-obfuscation interface according to the obfuscation characteristics. For the code files that cannot be successfully decoded or de-obfuscated, record the corresponding npm packages as obfuscated packages. Otherwise, replace the original code files with the decoded or de-obfuscated code files and proceed to the next static matching process. The static matching process includes: scanning the replaced code files to detect whether there are any suspicious behaviors. If there are, record the corresponding npm packages as suspicious malicious packages and mark the suspicious behaviors they satisfy in the results.

[0057] In the specific implementation process, the npm packages include the package.json file and the code files. For the files of the package.json type, use string-based rule matching to detect whether there are any suspicious behaviors. The suspicious behaviors include the existence of mining-sensitive strings, suspicious commands (including sensitive behaviors such as collecting sensitive information and transmitting it externally, accessing suspicious domain names or IPs, running suspicious executable programs, etc.).

[0058] For the code files, first detect whether they conform to the obfuscation characteristics or encoding characteristics. Among them, the obfuscation characteristics focus on changing the code structure to make it difficult to read and analyze, such as variable / function renaming, control flow flattening, dynamic execution (eval), etc.; the encoding characteristics are to transform the data to hide the string or key code content, such as Base64, Hex, Unicode escape, etc. Obfuscation usually affects the code logic, while encoding only affects the data representation.

[0059] Scan the code files of the npm packages to detect whether there are any suspicious behaviors (including sensitive behaviors such as encrypted files, obtaining sensitive information and transmitting it externally, running malicious programs, etc.).

[0060] In one implementation, scanning the code files of the npm packages to detect whether they conform to the obfuscation characteristics or encoding characteristics includes:

[0061] Detect whether there are encoding characteristics in the code files through regular expression rules;

[0062] Judge whether the code files conform to the obfuscation characteristics by the occurrence frequency of the obfuscation characteristic strings in the code files.

[0063] In the specific implementation, for the code files in S1.2, detecting whether they conform to the obfuscation characteristics or encoding characteristics and de-obfuscating or decoding them can be achieved through the following methods:

[0064] Step 1.2.1, detect whether there are encoding characteristics in the code files through regular expression rules;

[0065] Step 1.2.2: For the code files with encoding (features), decode them according to their encoding types, and replace the original code content with the decoded code content.

[0066] Step 1.2.3: Determine whether the code file is an obfuscated code file by confusing the occurrence frequency of the feature strings in the code file.

[0067] Step 1.2.4: For the obfuscated code files, de-obfuscate them according to their obfuscation features, and replace the original code content with the de-obfuscated code content.

[0068] It should be noted that in the npm software package, some key strings are protected by a certain encoding algorithm (such as base64), but these encoded strings are difficult to read and identify, so the corresponding decoding algorithm is needed for decoding.

[0069] Please refer to Figure 2 and Figure 3 , where Figure 2 is the overall framework flowchart of the malicious package detection method for the npm software package in the embodiment of the present invention. Figure 3 is the detection and de-obfuscation flowchart for the obfuscated software package in the embodiment of the present invention.

[0070] In one implementation, S2 can be implemented in the following manner:

[0071] S2.1: For some static rules in the pre-constructed static rule library, construct taint analysis rules that conform to malicious behaviors.

[0072] S2.2: Further analyze the maliciousness of the filtered suspicious malicious packages based on the taint analysis rules that conform to malicious behaviors.

[0073] In one implementation, S2.2 includes:

[0074] Match the filtered suspicious malicious packages with the taint analysis rules that conform to malicious behaviors. If the match is successful, further confirm them as malicious packages; otherwise, eliminate the suspicious malicious packages.

[0075] Specifically, some static rules with strong generality (such as the behavior of accessing the sensitive file / etc / passwd to obtain sensitive information) may cause a certain degree of false positives (i.e., judging a benign package as a malicious package), because benign packages may also contain similar operations. To reduce the false positives caused by these rules, the present invention proposes a taint analysis method based on string literals and combines it with Semgrep (an open-source static code analysis tool) for more accurate detection. For example, malicious behavior will only be reported when key sensitive strings (such as bash - i, cat, curl, etc.) go through taint propagation and finally enter the system execution function. The complete taint analysis rules cover the following scenarios:

[0076] · Encoding and obfuscation techniques: Detect Base64 encoding, obfuscation, encryption

[0077] · Sensitive information access: Environment variable acquisition, sensitive data leakage, suspicious links

[0078] · Command execution and hijacking: Command overwriting, code execution

[0079] · Malicious download and execution: Download executable files, suspicious EXE execution, suspicious installation scripts

[0080] For the above scenarios, corresponding semgrep scripts are customized and written in this embodiment. For static rules that are prone to false positives, such as the behavior of accessing the sensitive file / etc / passwd in obtaining sensitive information, since benign packages may also have such behaviors, they will be filtered by more stringent taint analysis rules. Therefore, this implementation maintains a list of static rules with a high false positive rate and customizes and writes semgrep taint analysis rules that conform to malicious behaviors for these rules respectively;

[0081] For the software packages that are identified as suspicious malicious packages in the scan result of step S1 only due to these static rules with strong generality and high false positive rate, semgrep taint analysis rule scans will be performed. For software packages that do not conform to the taint analysis rules, they will be filtered out from the suspicious malicious packages.

[0082] In one implementation, S3 can be achieved through the following steps:

[0083] S3.1: Classify malicious package types into six categories: information leakage, backdoor trojan, encryption ransomware, reverse shell, mining virus, and repository pollution, and submit corresponding code examples of each type of malicious package to ChatGPT respectively to construct a prompt, requiring that for the submitted npm software package, first confirm its maliciousness, and give its malicious package classification and the judgment basis for considering it as this type of malicious package;

[0084] S3.2: For the suspicious malicious packages obtained through preliminary filtering and the suspicious malicious packages obtained through further analysis, conduct maliciousness verification analysis based on ChatGPT, and according to the basis for judging maliciousness output by ChatGPT, check whether it is a rule existing in the existing static rule library. If not, add it to the existing static rule library.

[0085] Specifically, the existing static rule library includes a static rule list and a taint analysis rule.

[0086] In one implementation, S4 can be achieved through the following steps:

[0087] S4.1: The obfuscated software packages obtained through preliminary filtering are dynamically run in a sandbox environment;

[0088] S4.2: Use a preset tool to monitor the runtime behavior of the process during the running of the npm software package;

[0089] S4.3: According to the monitored runtime behavior, detect whether there are malicious features in the npm software package to verify its maliciousness.

[0090] Specifically, Figure 4 This is the flowchart of the dynamic analysis in the embodiment of the present invention.

[0091] In the specific implementation process, the strace tool can be used for monitoring. The runtime behavior of the process during the running of the software package includes system calls of the process during the running of the software package, received signals, network transmission, file operation conditions, etc.

[0092] According to the monitored runtime behavior, detecting whether there are malicious features in the npm software package to verify its maliciousness can be: for example, detecting the existence of malicious domains and wild IPs in the network transmission situation, and detecting the access situation to sensitive files and whether there is an encryption operation on sensitive files in the file operation situation.

[0093] Finally, malicious packages screened by static analysis verified by ChatGPT and obfuscated malicious packages matched with malicious features through dynamic running can be obtained. At the same time, the corresponding judgment basis of the malicious packages will be included in the output result for manual review by the user.

[0094] As can be seen from the above description, from the perspective of efficiently detecting malicious packages existing in a large sample set of npm software packages in the embodiments of the present invention, taking into account the consideration of a low false positive rate, a suspicious malicious package and a confusion package are first efficiently detected through a combination of dynamic and static solutions, and then a more fine-grained static taint analysis rule review and dynamic runtime feature maliciousness matching are respectively performed. At the same time, a large language model is introduced to assist in the confirmation of malicious packages and a rule self-learning system is implemented. Finally, a complete maliciousness description report is output for user review. Compared with traditional static analysis and dynamic behavior analysis methods, the method of the present invention has the characteristics of high efficiency, applicability to large sample sets, low false positive rate, high accuracy, and the ability to detect new attacks.

[0095] Embodiment 2

[0096] Based on the same inventive concept, this embodiment discloses a malicious package detection system for npm software packages, including:

[0097] A static rule matching module, configured to preliminarily filter out suspicious malicious packages and confusion software packages from a large sample dataset based on a static rule library, where the static rule library is pre-constructed according to the characteristics included in different types of malicious packages;

[0098] A string-based taint analysis module, configured to further analyze the maliciousness of the filtered suspicious malicious packages by using a string-based taint analysis method;

[0099] A ChatGPT verification and rule update module, configured to submit the suspicious malicious packages obtained by preliminary filtering and the suspicious malicious packages obtained by further analysis to a pre-constructed large model to further verify their maliciousness, and at the same time update the learned new malicious features to the existing static rules;

[0100] A dynamic analysis module, configured to use a dynamic analysis solution to verify the maliciousness of the confused software packages;

[0101] A detection result obtaining module, configured to integrate the malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis solution as the detection result.

[0102] Among them, the dynamic analysis module includes a dynamic malicious feature extraction module, configured to perform dynamic runtime feature extraction and maliciousness matching on the confused packages whose maliciousness cannot be determined by static feature recognition.

[0103] Since the system introduced in Embodiment 2 of the present invention is the system adopted for implementing the malicious package detection method for npm software packages in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the system, so it will not be elaborated here. Any system adopted by the method in Embodiment 1 of the present invention belongs to the scope protected by the present invention.

[0104] Embodiment III

[0105] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in Embodiment I is implemented.

[0106] Since the computer-readable storage medium introduced in Embodiment III of the present invention is the computer-readable storage medium used for the malicious package detection method for npm software packages in Embodiment I of the present invention, based on the method described in Embodiment I of the present invention, those skilled in the art can understand the specific structure and variations of this computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium used in the method of Embodiment I of the present invention falls within the scope of protection of the present invention.

[0107] Embodiment IV

[0108] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in Embodiment I is implemented.

[0109] Since the computer device introduced in Embodiment IV of the present invention is the computer device used for the malicious package detection method for npm software packages in Embodiment I of the present invention, based on the method described in Embodiment I of the present invention, those skilled in the art can understand the specific structure and variations of this computer device, so it will not be elaborated here. Any computer device used in the method of Embodiment I of the present invention falls within the scope of protection of the present invention.

[0110] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0111] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.

[0112] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A malicious package detection method for npm software package, characterized in that: include: Preliminarily filter out suspicious malicious packages and obfuscated software packages from a large sample data set based on a static rule base, where the static rule base is pre-built based on the features contained in different types of malicious packages; The maliciousness of the filtered suspicious malicious packets is further analyzed using a string-based point analysis method; Submit the suspicious malicious packets obtained through preliminary filtering and further analysis to the pre-built large model to further verify their maliciousness, and update the learned new malicious features to the existing static rules; For obfuscated software packages, dynamic analysis solutions are used to verify their maliciousness; The malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis solution are integrated as the detection results.

2. The malicious package detection method for npm software package according to claim 1, characterized in that: Based on the static rule base, suspicious malicious packages and obfuscated software packages are initially filtered out from a large sample data set, including: Perform string-based rule matching on the package.json of the npm software package to detect whether there is any suspicious behavior. For npm packages with suspicious behavior, record them as suspicious malicious packages and mark the suspicious behaviors they meet in the results; Scan the code files of the npm software package to detect whether they meet the obfuscation characteristics or encoding characteristics. If they meet the requirements, make corresponding decoding attempts based on the encoding characteristics or call the deobfuscation interface based on the obfuscation characteristics. For code files that cannot be successfully decoded or deobfuscated, record their corresponding npm software packages as obfuscated software packages. Otherwise, replace the original code files with the decoded or deobfuscated code files and proceed to the next static matching process. The static matching process includes: scanning the replaced code files to detect whether there is any suspicious behavior in them. If so, record the corresponding npm software package as a suspicious malicious package and mark the suspicious behavior it meets in the results.

3. The malicious package detection method for npm software package as claimed in claim 2, characterized in that: Scan the code files of the npm package to check whether they meet the obfuscation features or encoding features, including: Use regular expression rules to detect whether there are coding features in the code file; The frequency of occurrence of obfuscation feature strings in the code file is used to determine whether the code file meets the obfuscation feature.

4. The malicious package detection method for npm software package according to claim 1, characterized in that: The maliciousness of the filtered suspicious malicious packets is further analyzed using a string-based point analysis method, including: Construct taint analysis rules that match malicious behaviors for some static rules in the pre-built static rule library; The maliciousness of the filtered suspicious malicious packets is further analyzed based on the taint analysis rules that conform to malicious behavior.

5. The malicious package detection method for npm software package as claimed in claim 4, characterized in that: The maliciousness of the filtered suspicious malicious packets is further analyzed based on the taint analysis rules that match the malicious behavior, including: The filtered suspicious malicious packets are matched with the taint analysis rules that match the malicious behavior. If the match is successful, it is further confirmed as a malicious packet. Otherwise, the suspicious malicious packet is removed.

6. The malicious package detection method for npm software package according to claim 1, characterized in that: Submit the suspicious malicious packets obtained through further analysis to the pre-built large model to further verify their maliciousness. At the same time, the learned new malicious features are updated to the existing static rules, including: Malicious packages are classified into six categories: information leakage, backdoor Trojans, encryption ransomware, rebound shells, mining viruses, and warehouse pollution. The corresponding code examples of each type of malicious package are submitted to ChatGPT, and a prompt is constructed to confirm the maliciousness of the submitted npm software package, and give its malicious package classification and the basis for judging whether it belongs to this type of malicious package. For the suspicious malicious packets obtained through preliminary filtering and further analysis, malicious verification analysis based on ChatGPT is performed, and according to the malicious judgment basis output by ChatGPT, it is checked whether the rules already exist in the existing static rule base. If not, they are added to the existing static rule base.

7. The malicious package detection method for npm software package according to claim 1, characterized in that: For obfuscated software packages, dynamic analysis solutions are used to verify their maliciousness, including: The obfuscated software package obtained by preliminary filtering is dynamically run in a sandbox environment; Use pre-defined tools to monitor the runtime behavior of processes during npm package execution; Based on the monitored runtime behavior, detect whether there are malicious features in the npm package to verify its maliciousness.

8. A malicious package detection system for npm software packages, characterized in that: include: A static rule matching module is used to preliminarily filter out suspicious malicious packages and obfuscated software packages from a large sample data set based on a static rule base, wherein the static rule base is pre-built according to the features contained in different types of malicious packages; The string-based taint analysis module is used to further analyze the maliciousness of the filtered suspicious malicious packets using a string-based point analysis method; ChatGPT verification and rule update module, which is used to submit the suspicious malicious packets obtained through preliminary filtering and further analysis to the pre-built large model to further verify their maliciousness and update the learned new malicious features to the existing static rules; Dynamic analysis module, used to verify the maliciousness of obfuscated software packages using dynamic analysis solutions; The detection result acquisition module is used to integrate the malicious package results confirmed by the large model and the malicious package results confirmed by the dynamic analysis solution as the detection results.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the malicious package detection method for the npm software package as described in any one of claims 1 to 7 is implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the malicious package detection method for the npm software package is implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Detection method for APTs (Advanced Persistent Threat) based on instruction monitoring

    CN106850582A

  • Malicious code package detection method and device, computer equipment and storage medium

    CN114637993A

  • Malicious npm packet detection method, device, equipment and medium

    CN115718918A

  • Enhanced taint analysis Android malicious code detection method and device

    CN116451219A

  • Malicious software package detection method based on malicious behavior sequence feature modeling

    CN117056924A