Malicious code detection method and code detection device

Through the combination of large language model and neural network model, malicious code detection is carried out on source code products written in multiple languages, solving the problem of limited detection scope in the existing technology, and achieving wider and more accurate malicious code identification and false positive correction.

CN120337213APending Publication Date: 2025-07-18SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193296.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing malicious code detection methods can only detect code slices in one programming language, and the detection range is limited, so it is impossible to effectively identify malicious code in source code products written in multiple languages.

Method used

The large language model is used to process the programming language identification and similar samples of the source code products. The detection results are output through the first large language model, combined with the prompt information of the reference code type and similar samples, the detection range is expanded, and the neural network model is used to quickly filter legal code slices to reduce the number of detections and improve detection efficiency.

Benefits of technology

It realizes malicious code detection of source code products written in multiple languages, improves the accuracy and efficiency of detection, can identify new malicious code, correct false positives, and output detailed explanation of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337213A_ABST
    Figure CN120337213A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a malicious code detection method and a code detection device, which can be used for detecting source code products written in multiple languages and judging whether the source code products comprise malicious codes or not, so that the types of the detectable source code products are increased, and the detection range is expanded. The method comprises the steps that after a code detection device obtains code slices from a source code product, similar samples of the code slices are obtained, then the code slices, programming language identifiers of the source code product and the similar samples are input into a first large language model, and whether the code slices comprise malicious codes or not is judged through the first large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing, and particularly to a malicious code detection method and a code detection device. Background Art

[0002] Malicious code refers to a code segment or software specifically made to perform unauthorized actions.

[0003] Currently, a malicious code detection method is generally as follows: Locate the source points and sink points in the program code, construct a code slice based on the source points and sink points, determine the programming language corresponding to the code slice, input the code slice into a neural network classifier of the bidirectional encoder representation from transformers (BERT) architecture corresponding to the programming language identifier, and output whether the code slice contains malicious code through the neural network classifier.

[0004] Each neural network classifier can only detect code slices of one programming language and cannot detect malicious code in other programming languages. Therefore, its scope of application is limited. Summary of the Invention

[0005] This application provides a malicious code detection method, which can use a large language model to process the code slices of source code artifacts, the programming language identifiers of the source code artifacts, and similar samples, so as to detect source code artifacts written in multiple languages and determine whether they contain malicious code.

[0006] In a first aspect, a malicious code detection method is provided. The method includes: After a code detection device obtains code slices from a source code artifact, it obtains similar samples of the code slices, and then inputs the code slices, the programming language identifier of the source code artifact, and the similar samples into a first large language model. The detection result of the code slices is output through the first large language model, and the detection result is used to indicate whether the code slices contain malicious code. The code slices, the programming language identifier of the source code artifact, and the similar samples of the code slices include the programming language information of the code slices. The above information, as the input data of the first large language model, enables the first large language model to detect source code artifacts written in multiple languages and determine whether they contain malicious code, increasing the types of source code artifacts that can be detected and expanding the detection scope.

[0007] In some possible implementation manners, the above malicious code detection method further includes: the code detection device determines the reference code type corresponding to the code slice. When the reference code type belongs to the malicious code type, the code detection device obtains the first prompt information corresponding to the reference code type from the template library; the code detection device inputs the code slice, the programming language identifier of the source code artifact, and the similar sample into the first large language model, and the detection result of the code slice output by the first large language model includes: the code detection device inputs the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar sample into the first large language model, and the detection result of the code slice is output by the first large language model. The first prompt information may include one or more of the definition of the malicious code type, the analysis steps of the malicious code type, or the precautions of the malicious code type. After inputting the first prompt information related to the malicious code type into the first large language model, the first large language model can identify the malicious code type associated with the first prompt information, and can improve the classification accuracy of the first large language model.

[0008] In some possible implementation manners, the code detection device obtaining the reference code type corresponding to the code slice includes: the code detection device inputs the application programming interface (API) function of the code slice into the neural network model, and the reference code type corresponding to the code slice is output by the neural network model. The neural network model is a small model trained based on the few shot learning algorithm and malicious samples. Using it, the reference code type of the code slice can be quickly determined, and legal code slices can be quickly screened out. Moreover, the model parameters of the neural network model are less, which is convenient for deployment.

[0009] In some possible implementation manners, the above malicious code detection method further includes: when the target malicious code from the user does not belong to the template library, the code detection device obtains the prompt information of the target malicious code and the malicious code type of the target malicious code, establishes the corresponding relationship between the malicious code type of the target malicious code and the prompt information of the target malicious code in the template library, and adds the target malicious code to the sample library where the similar samples belong. The user can add new malicious codes (i.e., target malicious codes) to the template library and the sample library. For malicious codes similar to the target malicious code, the code detection device can obtain the prompt information of the target malicious code and the target malicious code, and use them as the input data of the large language model, so as to quickly identify malicious codes similar to the new malicious code (such as the same type of malicious codes as the new malicious code).

[0010] In some possible implementation manners, the above malicious code detection method further includes: The code detection device determines a document set corresponding to the programming language identifier from the API document library. After obtaining the target API document corresponding to the API function of the code slice from the document set, the code slice, the programming language identifier of the source code artifact, the first prompt message, the similar sample, and the target API document are input into the first large language model. Using the target API document associated with the code slice as the input data of the large language model can improve the recognition accuracy of malicious code.

[0011] In some possible implementation manners, the input data of the first large language model further includes the identifier of the source code artifact. Inputting the identifier of the source code artifact into the first large language model can improve the classification ability of the first large language model for the code slice.

[0012] In some possible implementation manners, the code detection device obtains the code slice from the source code artifact by: The code detection device determines the source point and the convergence point of the source code artifact according to the sensitive interface function in the source code artifact, and selects the code slice from the source code artifact according to the source point and the convergence point of the source code artifact. In this way, the code slice associated with the sensitive interface function can be selected for malicious code detection, and the code slice not associated with the sensitive interface function can be not detected, so the number of code slices to be detected can be reduced, and the code detection efficiency can be improved.

[0013] In some possible implementation manners, the above malicious code detection method further includes: When the code slice includes malicious code, the code detection device extracts the first vector feature from the vector of the code slice, searches for the target false positive code corresponding to the first vector feature in the false positive sample library, inputs the code slice and the target false positive code into the second large language model, and determines whether the malicious code of the code slice is a false positive through the second large language model. Among them, the similarity value between the first vector feature and the vector of the target false positive code is greater than the reference similarity threshold. In this way, the large language model can be used for false positive detection to correct the false positive malicious code slice, and the correct rate of malicious code detection can be improved.

[0014] In some possible implementation manners, the above malicious code detection method further includes: when the malicious code in the code slice is not a false positive and the detection result includes the target malicious code type, the code detection device extracts a second vector feature from the vector of the code slice, and then obtains second prompt information including key code features, source code artifact identifiers, function names in the source code artifact, and output languages from the template library according to the target malicious code type and the second vector feature. Then, the code slice and the second prompt information are input into a third large language model, and a detection result explanation is output through the third large language model. The detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code. The judgment result of the key code features includes that the key code feature is a malicious code feature or the key code feature is not a malicious code feature. In this way, a detection result explanation can be output by the large language model, and it can be checked whether there is an error in the process of the large language model detecting malicious code according to the detection result explanation.

[0015] In some possible implementation manners, the second prompt information further includes analysis steps corresponding to the target malicious code type, and the detection result explanation further includes a summary of the analysis steps corresponding to the target malicious code type.

[0016] A second aspect provides a code detection device, which includes a slicing module, a prompting module, and a code analysis module. The slicing module is used to obtain a code slice from a source code artifact; the prompting module is used to obtain a similar sample of the code slice; the code analysis module is used to input the code slice, the programming language identifier of the source code artifact, and the similar sample into a first large language model, and output a detection result of the code slice through the first large language model. The detection result is used to indicate whether the code slice includes malicious code.

[0017] In some possible implementation manners, the prompting module is further used to determine a reference code type corresponding to the code slice; when the reference code type is a malicious code type, obtain first prompt information corresponding to the reference code type from the template library, and the code analysis module is specifically used to input the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar sample into the first large language model, and output a detection result of the code slice through the first large language model.

[0018] In some possible implementation manners, the prompting module is specifically used to input API functions of the code slice into a neural network model, and output a reference code type corresponding to the code slice through the neural network model.

[0019] In some possible implementation manners, when the target malicious code from the user does not belong to the template library, the prompting module is further used to obtain prompt information of the target malicious code and the malicious code type of the target malicious code, establish a corresponding relationship between the prompt information of the target malicious code and the malicious code type of the target malicious code in the template library, and add the target malicious code to the sample library to which the similar sample belongs.

[0020] In some possible implementation manners, the prompting module is further configured to determine a document set corresponding to the programming language identifier from the API document library, and obtain a target API document corresponding to the API function of the code slice from the document set; the code analysis module is specifically configured to input the code slice, the programming language identifier of the source code artifact, the first prompting information, the similar samples, and the target API document into the first large language model.

[0021] In some possible implementation manners, the first prompting information includes at least one of a definition of a reference code type, an analysis step of a reference code type, or a note of a reference code type.

[0022] In some possible implementation manners, the input data of the first large language model further includes an identifier of the source code artifact.

[0023] In some possible implementation manners, the slicing module is specifically configured to determine a source point of the source code artifact and a convergence point of the source code artifact according to sensitive interface functions in the source code artifact; and select a code slice from the source code artifact according to the source point and the convergence point of the source code artifact.

[0024] In some possible implementation manners, the code detection device further includes a false alarm detection module. When the code slice includes malicious code, the false alarm detection module is configured to extract a first vector feature from the vector of the code slice; search for a target false alarm code corresponding to the first vector feature in the false alarm sample library, input the code slice and the target false alarm code into the second large language model, and determine whether the malicious code in the code slice is a false alarm through the second large language model.

[0025] In some possible implementation manners, the code detection device further includes an explanation module. When the malicious code in the code slice is not a false alarm and the detection result includes a target malicious code type, the explanation module is configured to extract a second vector feature from the vector of the code slice; obtain second prompting information from the template library according to the target malicious code type and the second vector feature, input the code slice and the second prompting information into the third large language model, and output a detection result explanation through the third large language model.

[0026] In some possible implementation manners, the second prompting information further includes an analysis step corresponding to the target malicious code type, and the detection result explanation further includes a summary of the analysis step corresponding to the target malicious code type.

[0027] For the noun explanations, the specific steps executed by each module, and the beneficial effects in the second aspect, reference may be made to the corresponding descriptions in the first aspect.

[0028] A third aspect provides a computing device cluster, which includes at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the malicious code detection method as described in the first aspect or any possible implementation manner of the first aspect.

[0029] A fourth aspect provides a computer-readable storage medium, which includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the malicious code detection method in the first aspect or any possible implementation manner of the first aspect.

[0030] A fifth aspect provides a computer program product, which includes computer program instructions; when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the malicious code detection method in the first aspect or any possible implementation manner of the first aspect.

[0031] Based on the implementation manners provided in the above aspects of this application, further combinations can be made to provide more implementation manners. Description of the Drawings

[0032] Figure 1 It is a schematic diagram of a malicious code detection scenario in an embodiment of this application;

[0033] Figure 2 It is another schematic diagram of a malicious code detection scenario in an embodiment of this application;

[0034] Figure 3 It is a flowchart of a malicious code detection method in an embodiment of this application;

[0035] Figure 4 It is a schematic diagram of a malicious code detection method in an embodiment of this application;

[0036] Figure 5 It is a schematic diagram of a false positive detection method in an embodiment of this application;

[0037] Figure 6 It is a schematic diagram of explaining a detection result in an embodiment of this application;

[0038] Figure 7 It is a structural diagram of a code detection device in an embodiment of this application;

[0039] Figure 8 It is a schematic diagram of a code detection device detecting malicious code in an embodiment of this application;

[0040] Figure 9 It is another schematic diagram of a code detection device detecting malicious code in an embodiment of this application;

[0041] Figure 10 It is a structural diagram of a computing device in an embodiment of the present application;

[0042] Figure 11 It is a structural diagram of a computing device cluster in an embodiment of the present application;

[0043] Figure 12 It is another structural diagram of a computing device cluster in an embodiment of the present application. Detailed implementation manners

[0044] The malicious code detection method of the present application can be applied to the code detection scenario of the software supply chain. The software supply chain refers to the integration of all relevant resources and activities in the software development process, which includes but is not limited to code writing, component selection, tool use, interaction between developers and other stakeholders, software maintenance and update. The security of the software supply chain includes the security of the coding process, tools, devices, or the code, modules, and services from the upstream of the supply chain in all stages of software design and development on the supply chain, as well as the security during the software delivery channel and the usage process.

[0045] Code poisoning in the software supply chain means that an attacker implants malicious code into the software through hijacking, tampering, etc. during the development, dissemination, and upgrade processes of the software, so as to achieve the purpose of stealing information or destroying the system. Due to the lack of supervision in the third-party central repository, the package manager and the central repository have become the main targets for current code poisoning attacks. That is, the attacker uploads third-party software containing malicious code to the central repository through hijacking, tampering, etc. When the victim uses the package manager to install third-party software, the software containing malicious code will be downloaded and installed, and the malicious code will be triggered during the use or operation of the software.

[0046] The following introduces the code detection scenario of the software supply chain. Refer to Figure 1 , in a malicious code detection scenario, the user sends the source code artifact to the code detection device 110, and the code detection device 110 detects the source code artifact. When the source code artifact does not include malicious code, the source code artifact is stored in the central repository 121 of the server 120. When the source code artifact includes malicious code, the source code artifact is isolated from the central repository 121. In this way, the code detection device 110 can perform a full-scale analysis on the source code artifacts uploaded to the central repository 121, timely discover malicious code, prevent malicious code from entering the central repository 121, and solve malicious code at the entrance of the central repository 121.

[0047] The central repository 121 is an open software repository where third-party developers or third-party organizations can upload and release third-party packages. Third-party packages are also known as third-party software packages.

[0048] The package manager 122 is a software tool for managing third-party software packages, which can automate the installation, upgrade, configuration, and deletion of third-party packages. For example, the most popular package manager for the Python language is PIP, and the most popular package manager for the JavaScript language is NPM. The package manager 130 can search for and install corresponding third-party packages in the central repository based on third-party package identifiers (such as name, version, etc.). It can perform functions such as automatically downloading and installing third-party software packages.

[0049] In Figure 1 the scenario shown, the code detection device 110 and the server 120 are different computing devices. In actual applications, the code detection device 110 can be integrated into the server 120. The package manager 122 and the central repository 121 can also be deployed on different computing devices. It should be understood that the central repository 121 can also be deployed in a server cluster.

[0050] Refer to Figure 2 , in another malicious code detection scenario, for third-party libraries introduced by developers, the code detection device 210 detects all source code artifacts of the third-party library. When all source code artifacts of the third-party library do not include malicious code, the third-party library is integrated into the development environment 221 of the computing device 220. When the source code artifacts include malicious code, the source code artifacts are rejected from entering the development environment 221 of the computing device 220. It should be understood that for third-party libraries that do not include malicious code, the third-party library can also be integrated into its related test environment or production environment. In Figure 2 the scenario shown, the code detection device 210 and the computing device 220 are different computing devices. In actual applications, the code detection device 210 can also be integrated into the computing device 220. The computing device 220 can be, but is not limited to, a server or a terminal. The terminal can be, but is not limited to, a desktop computer, a mobile phone, a tablet computer, a wearable device, an in-vehicle computer, or an Internet of Things device.

[0051] Regarding the problem of the small detection scope of the current malicious code detection model, the present application provides a malicious code detection method, which can detect source code artifacts written in various languages and determine whether they include malicious code, having a larger detection scope.

[0052] The malicious code detection method of the present application will be introduced below with the code detection device as the execution subject. Refer to Figure 3, in one embodiment, the malicious code detection method of the present application includes the following steps:

[0053] S301. Obtain a code slice from the source code artifact.

[0054] In this embodiment, the source code artifact refers to a binary file generated by compiling and packaging the source code. Different development languages correspond to binary files in different formats, and these binary files can usually be directly run on the server or terminal. The source code artifact may be, but is not limited to, a software package. Optionally, S301 includes: determining the source point and the sink point of the source code artifact according to the sensitive interface functions in the source code artifact, and selecting a code slice from the source code artifact according to the source point and the sink point of the source code artifact. In this way, suspicious code slices can be selected from the source code artifact according to the sensitive interface functions, and the code slices not related to the sensitive interface functions can be excluded, reducing the number of code slices to be detected and improving the detection speed of the source code artifact.

[0055] Among them, the sensitive interface function is a function used to perform sensitive operations. Sensitive operations include, but are not limited to, system calls or sending information to the outside through the network. A system call is a type of operation provided by the operating system to the application program, such as reading and writing files, executing system commands, etc. System calls can be encapsulated into application programming interfaces (APIs) for developers to call. Malicious code can also achieve attack purposes by calling some dangerous API functions without directly calling the underlying system calls. Such API functions can be called sensitive interface functions. In the source code artifact, the sink is the call point of the sensitive interface function, and this call point will receive one or more variables as parameters. The source is the generation point of the above variables. For a set of sinks and sources, one or more program execution paths can be obtained. Extracting the code of this program execution path as a program code segment, this process is called program slicing, and the obtained program code segment can be simply called a slice or a code slice.

[0056] S302. Obtain a similar sample of the code slice.

[0057] Specifically, calculate the similarity value between the vector of the code slice and the vector of the code sample in the sample library. For the code sample with a similarity value greater than the similarity threshold, randomly select one of them as the similar sample of the code slice, or determine the similar sample of the code slice as the similar sample corresponding to the maximum similarity value. The similarity threshold can be set according to the actual situation, and the present application does not make a limit.

[0058] S303. Input the code slice, the programming language identifier of the source code artifact, and the similar samples into the first large language model, and output the detection result of the code slice through the first large language model. The detection result is used to indicate whether the code slice contains malicious code.

[0059] The programming language identifier of the source code artifact can be, but is not limited to, C, C++, C#, Java, or Python, and can be specifically set according to the actual situation. When the code slice contains malicious code, the detection result can also include the malicious code subclass of the malicious code. Optionally, the input data of the first large language model can also include the identifier of the source code artifact. The identifier of the source code artifact can include, but is not limited to, the name of the source code artifact. For example, the identifier of the source code artifact can also include the version of the source code artifact.

[0060] In this embodiment, the code slice, the programming language identifier of the source code artifact, and the similar samples of the code slice include the programming language information of the code slice. The above information, as the input data of the first large language model, enables the first large language model to detect source code artifacts written in multiple languages. Compared with the malicious code detection model that detects a single language, the detection range of malicious code is expanded, providing good scalability.

[0061] Secondly, both the programming language identifier of the source code artifact and the similar samples are data highly related to the code slice. Using the programming language identifier of the source code artifact and the similar samples as the input data of the first large language model can improve the classification ability of the first large language model for the code slice, thereby improving the accuracy of malicious code detection. This application can also input other data related to the code slice into the first large language model to further improve the accuracy of malicious code detection. In an optional embodiment, the malicious code detection method of this application further includes: determining the reference code type corresponding to the code slice; when the reference code type belongs to the malicious code type, obtaining the first prompt information corresponding to the reference code type from the template library; S303 includes: inputting the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar samples into the first large language model, and outputting the detection result of the code slice through the first large language model.

[0062] Optionally, determining the reference code type corresponding to the code slice includes: inputting all the API functions of the code slice into the neural network model, and outputting the reference code type corresponding to the code slice through the neural network model. All the API functions of the code slice can be, but are not limited to, the API function sequence, and the order of the API functions in the API function sequence is the same as the order of the API functions in the code slice. Optionally, the neural network model is trained based on the few-shot learning algorithm and malicious samples. The model parameters of this neural network model are less and the computational amount is smaller, so the detection efficiency is high.

[0063] The reference code type may include a malicious code type or a legitimate code type. When the reference code type is a legitimate code type, the code slice can be considered a legitimate slice. The malicious code type can be divided into multiple major malicious code categories, and each major malicious code category includes one or more malicious code subcategories. Optionally, the reference code type can be a malicious code subcategory. For example, Trojan.1, where Trojan represents the major malicious code category of Trojan horses, and Trojan.1 is the malicious code subcategory. The major malicious code categories include, but are not limited to, computer viruses, Trojan horses, worms, logic bombs, bacteria, malicious scripts, malicious controls, or spyware, which can be specifically set according to the actual situation and are not limited in this application. The number of major malicious code categories and the number of malicious code subcategories in each major malicious code category are set according to the actual situation and are not limited in this application.

[0064] Optionally, when the reference code type is a malicious code subcategory, each malicious code subcategory in the template library corresponds to a first prompt word template, and the first prompt word template includes the first prompt information of the malicious code subcategory. The first prompt information includes one or more of the definition of the malicious code subcategory, the analysis steps of the malicious code subcategory, or the precautions of the malicious code subcategory. When the reference code type is a major malicious code category, this application can also set a first prompt word template for each major malicious code category. The first prompt word template includes the first prompt information of the major malicious code category, and the first prompt information of the major malicious code category includes one or more of the definition of the major malicious code category, the analysis steps of the major malicious code category, or the precautions of the major malicious code category. When the reference code type is a subcategory of a malicious code subcategory, a first prompt word template can also be set for each subcategory of the malicious code subcategory. By analogy, this application can also set a first prompt word template for other subordinates of the malicious code subcategory.

[0065] In this embodiment, the first prompt information is data highly related to the malicious code type. Using the first prompt information as the input data of the first large language model can improve the classification ability of the first large language model for code slices, thereby improving the accuracy of malicious code detection. Even when the malicious code type is not used as training data when training the first large language model, the first large language model still has the ability to recognize the malicious code type.

[0066] Moreover, the neural network model trained by few-shot learning can be used to perform the first detection on the code slice, and then the first large language model can be used to perform the second detection on the code slice, which can improve the accuracy of identifying malicious code.

[0067] In another alternative embodiment, the malicious code detection method of the present application further includes: determining a document set corresponding to a programming language identifier from an API document library, obtaining a target API document corresponding to the API function of the code slice from the document set, and then inputting the code slice, the programming language identifier of the source code artifact, the first prompt information, the similar sample, and the target API document into a first large language model, and outputting a detection result of the code slice through the first large language model, where the detection result is used to indicate whether the code slice includes malicious code.

[0068] The API document includes one or more of the basic information, usage, usage guide, endpoints and methods, authentication methods, request parameters, response objects, error codes and messages, example code, and common problems when using the API of the API. The basic information of the API includes but is not limited to the name, version number, author, or release log of the API.

[0069] In this embodiment, adding the API document of the code slice as a prompt word for the first large language model can further improve the classification ability of the first large language model for the code slice, thereby improving the accuracy of malicious code detection.

[0070] For ease of understanding, the malicious code detection method of the present application is introduced below with another embodiment. Refer to Figure 4 , in one embodiment, the malicious code detection method of the present application includes the following steps:

[0071] S401. Program slicing.

[0072] Specifically, determine the source point of the source code artifact and the convergence point of the source code artifact according to the sensitive interface functions in the source code artifact, and select a code slice from the source code artifact according to the source point and the convergence point of the source code artifact.

[0073] S402. Retrieve a prompt word template in the template library.

[0074] Input all API functions of the code slice into a neural network model, and output the code type corresponding to the code slice through the neural network model. After obtaining the code type of the code slice, retrieve a prompt word template in the template library according to the code type of the code slice. The prompt word template includes the first prompt information corresponding to the code type.

[0075] S403. Retrieve a similar sample in the sample library.

[0076] Calculate the similarity value between the vector of the code slice and the vectors of the code samples in the sample library. For the code samples with a similarity value greater than the similarity threshold, randomly select one of them as the similar sample of the code slice, or determine the similar sample of the code slice as the similar sample corresponding to the maximum similarity value.

[0077] S404. Retrieve the target API documentation in the document library.

[0078] Determine the document set corresponding to the programming language identifier from the API document library, and obtain the target API documentation corresponding to the API function of the code slice from the document set.

[0079] S405. Data assembly.

[0080] Assemble the code slice, the programming language identifier of the source code artifact, the first prompt message, the similar sample, and the target API documentation into a prompt word in a preset order.

[0081] S406. Detect malicious code using the first large language model.

[0082] After inputting the assembled prompt word into the first large language model, use the first large language model to detect whether the code slice contains malicious code.

[0083] S407. Does the detection result include malicious code? If yes, output the malicious code slice and the malicious code type; if no, output no malicious code.

[0084] In this embodiment, the first large language model can accurately identify the malicious code of the source code artifact based on the code slice, the programming language identifier of the source code artifact, the first prompt message, the similar sample, and the target API documentation.

[0085] Before executing step S303, the first large language model needs to be trained first. The following introduces the training process of the first large language model. In one embodiment, training the first large language model includes establishing a first training sample set and training the model.

[0086] In the process of establishing the first training sample set, obtain source code artifacts including malicious code and source code artifacts without malicious code from the open source library, and label the code types for the code slices of the source code artifacts. Extract the code slice and the programming language identifier of the source code artifact from the source code artifact, obtain the similar sample of the code slice, and then form a query text with the code slice, the programming language identifier of the source code artifact, and the similar sample of the code slice. When the code type of the code slice is a legal code type, use the legal code type as the response text. When the code type of the code slice belongs to the malicious code type, use the malicious code type of the code slice as the response text, and then generate a training sample including the query text and the response text. Optionally, the query text can also include the first prompt message for each malicious code type, and the first prompt message includes one or more of the definition of the malicious code type, the analysis steps of the malicious code type, or the precautions of the malicious code type.

[0087] In the process of training the model, the first training sample set is used to perform supervised fine tuning (SFT) training on the base large model to obtain a first large language model. The base large model can be, but is not limited to, a generative pre-trained transformer (GPT) model, a large language model Meta artificial intelligence (LLAMA) model, or a deepseek model.

[0088] There may be false detections when detecting malicious code through a large language model. The present application can further check the code slices to reduce the false detection rate. The following is a detailed introduction. In another optional embodiment, the malicious code detection method of the present application also includes: when the code slice includes malicious code, extracting a first vector feature from the vector of the code slice, searching for a target false alarm code corresponding to the first vector feature in the false alarm sample library, inputting the code slice and the target false alarm code into the second largest language model, and judging whether the malicious code in the code slice is a false alarm by the second largest language model.

[0089] In this embodiment, an embedding model can be used to extract a first vector feature from the vector of the code slice, and then the similarity value between the first vector feature and the false alarm code vector feature in the false alarm sample library is calculated. For false alarm samples whose similarity value is greater than the reference similarity threshold, the application can select any one of them as the target false alarm code, or use the false alarm code corresponding to the maximum similarity value as the target false alarm code. The reference similarity threshold can be set according to the actual situation, and this application does not limit it.

[0090] The following introduces the prompt word task including the target false positive code and the code slice. In one embodiment, the prompt word task is described as follows:

[0091]

[0092]

[0093] After constructing the above prompt words, the above prompt words are input into the second largest language model, and the second largest language model is used to determine whether the malicious code in the code slice is a false alarm. For ease of understanding, the false alarm detection method of the present application is introduced in another embodiment. Figure 5 In one embodiment, the malicious code detection method of the present application includes the following steps:

[0094] S501, searching for a near false alarm code.

[0095] Specifically, after obtaining the malicious code slices, calculate the similarity value between the vector features of the malicious code slices and the vector features of the false positive samples in the false positive sample library. The false positive samples with similarity values greater than the reference similarity threshold are approximate false positive samples. In this application, any one of the approximate false positive samples can be selected as the target false positive code, or the approximate false positive code corresponding to the maximum similarity value can be used as the target false positive code.

[0096] S502. Prompt word assembly.

[0097] Assemble the malicious code slice and its corresponding target false positive code into a prompt word.

[0098] S503. Use the second large language model to perform false positive detection on the prompt word.

[0099] After inputting the assembled prompt word into the second large language model, use the second large language model to output the false positive detection result. If the malicious code slice is a false positive, correct it to a legal code slice. If the malicious code slice is a false positive, its corresponding explanation text can be generated.

[0100] In this embodiment, the second large language model can identify whether a code slice is a false positive based on the code slice and false positive samples, and can correct the wrong detection result of the malicious code.

[0101] Before performing step S503, the second large language model needs to be trained first. The training process of the second large language model is introduced below. In one embodiment, training the second large language model includes establishing a second training sample set and training the model.

[0102] In the process of establishing the second training sample set, obtain source code artifacts including malicious code and source code artifacts without malicious code from the open source library, and label the code types for the code slices of the source code artifacts. Extract code slices and false positive code slices from the source code artifacts, and then form query texts with the code slices and false positive code slices. When the code type of the code slice is a legal code type, use the legal code type as the response text. When the code type of the code slice belongs to the malicious code type, use the malicious code type of the code slice as the response text, and then generate training samples including query texts and response texts. In the process of training the model, use the second training sample set to perform SFT training on the base large model to obtain the second large language model.

[0103] The method for explaining the detection result is introduced below. In another optional embodiment, the malicious code detection method of the present application further includes: when the malicious code in the code slice is not a false positive and the detection result includes the target malicious code type, extracting a second vector feature from the vector of the code slice; obtaining second prompt information from the template library according to the target malicious code type and the second vector feature, and then inputting the code slice and the second prompt information into a third large language model, and outputting a detection result explanation through the third large language model, where the detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code.

[0104] In this embodiment, a second prompt word template can be configured in the template library. The second prompt word template includes second prompt information of malicious code subclasses, and the second prompt information includes but is not limited to key code features, source code artifact identifiers, function names in source code artifacts, and output languages. Among them, the key code feature is the call point of a sensitive interface function, and the sensitive interface function can be pre-configured by an expert.

[0105] Optionally, the second prompt information further includes analysis steps corresponding to the malicious code type, and the detection result explanation further includes a summary of the analysis steps corresponding to the malicious code type.

[0106] Before inputting the code slice and the second prompt information into the third large language model, the third large language model needs to be trained first. The training process of the third large language model is introduced below. In one embodiment, training the third large language model includes establishing a third training sample set and training the model.

[0107] In the process of establishing the third training sample set, obtain source code artifacts including malicious code and source code artifacts without malicious code from the open source library, and label the code type for the code slices of the source code artifacts. Extract the code slice and the second prompt information of the code slice from the source code artifact, and then form a query text with the code slice and the second prompt information of the code slice. When the code type of the code slice is a legal code type, use the legal code type or the legal code explanation text as the response text of the code slice. When the code type of the code slice belongs to the malicious code type, use the malicious code explanation text corresponding to the code slice as the response text, and then generate a training sample including the query text and the response text. In the process of training the model, use the third training sample set to perform SFT training on the base large model to obtain the third large language model.

[0108] It should be noted that the training methods of the first large language model, the second large language model, and the third large language model in the present application can be but are not limited to SFT training.

[0109] After the explanation of the detection results of the third large language model, the user can check whether the judgment results of the key code features and the harm of malicious code are correct, and can further judge whether there are problems in the malicious code detection process of the large language model. If there are problems, the problems can be corrected manually. The following uses another embodiment to introduce the method for explaining the detection results. Refer to Figure 6 In another embodiment, the malicious code detection method in this application may include the following steps:

[0110] S601. Retrieve the prompt word template.

[0111] When the detection results include malicious code slices and malicious code types, obtain the vector features of the malicious code slices, use the malicious code type and the vector features of the malicious code slices as keywords, and retrieve in the template library with the keywords. The fields of the template library include malicious code types, vector features, and prompt word templates. The prompt word templates include but are not limited to key code features, source code artifact identifiers, function names in source code artifacts, and output languages.

[0112] S602. Assemble the prompt words.

[0113] Assemble the key code features, source code artifact identifiers, function names in source code artifacts, and output languages into prompt words in a preset order.

[0114] S603. Process the prompt words with the third large language model.

[0115] After inputting the prompt words into the third large language model, obtain the detection result explanation through the third large language model. The format of the detection result explanation can be but is not limited to JSON.

[0116] S604. Convert the format of the detection result explanation.

[0117] Convert the format of the detection result explanation from JSON to a user-defined format, and generate a malicious code recognition report (including the detection result explanation in the user-defined format). The user-defined format can include but is not limited to the TXT format and the DOC format.

[0118] The following introduces the hardware for implementing the malicious code detection method in this application. Refer to Figure 7, in one embodiment, the code detection device 700 of the present application includes a slicing module 701, a prompting module 702, a code analysis module 703, a false positive detection module 704, and an explanation module 705. The slicing module 701 is used to obtain code slices from source code artifacts; the prompting module 702 is used to obtain similar samples of the code slices; the code analysis module 703 is used to input the code slices, the programming language identifier of the source code artifacts, and the similar samples into a first large language model, and output a detection result of the code slices through the first large language model, and the detection result is used to indicate whether the code slices include malicious code.

[0119] Taking the module as an example of a software functional unit, the code analysis module 703 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the above computing instance may be one or more. For example, the code analysis module 703 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally one region may include multiple AZs.

[0120] Similarly, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Among them, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0121] As an example of a hardware functional unit, the code analysis module 703 may include at least one computing device, such as a server, etc. Alternatively, the code analysis module 703 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0122] The multiple computing devices included in the code analysis module 703 may be distributed in the same region or in different regions. The multiple computing devices included in the code analysis module 703 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the code analysis module 703 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0123] In other embodiments, the code analysis module 703 may be used to execute Figures 3 to 6 any step performed by the code detection device in the malicious code detection method shown, the slicing module 701 may be used to execute Figures 3 to 6 any step performed by the code detection device in the malicious code detection method shown, the prompting module 702 may be used to execute Figures 3 to 6 any step performed by the code detection device in the malicious code detection method shown, the false alarm detection module 704 may be used to execute Figures 3 to 6 any step performed by the code detection device in the malicious code detection method shown, the explanation module 705 may be used to execute Figures 3 to 6 any step performed by the code detection device in the malicious code detection method shown. The steps to be implemented by the slicing module 701, the prompting module 702, the code analysis module 703, the false alarm detection module 704, and the explanation module 705 can be specified as needed and are implemented by the slicing module 701, the prompting module 702, the code analysis module 703, the false alarm detection module 704, and the explanation module 705 respectively Figures 3 to 6 to implement all the functions of the code detection device through different steps in the malicious code detection method shown.

[0124] The code detection device 700 can implement the functions of the code detection device 110. Refer to Figure 8 , in one embodiment, the slicing module 701 is used to obtain code slices from the source code artifact; the hint module 702 is used to obtain similar samples of the code slices; the code analysis module 703 is used to input the code slices, the programming language identifier of the source code artifact, and the similar samples into the first large language model, and output the detection result of the code slices through the first large language model. When the detection result shows that all the code slices of the source code artifact are legal, the legal source code artifact is saved in the central warehouse 121.

[0125] When the detection result is used to indicate that the code slice includes malicious code, the malicious code slice is sent to the false positive detection module 704. The false positive detection module 704 extracts the first vector feature from the vector of the code slice; searches for the target false positive code corresponding to the first vector feature in the false positive sample library, and inputs the code slice and the target false positive code into the second large language model, and judges whether the malicious code of the code slice is a false positive through the second large language model. When all the malicious codes of the source code artifact in the false positive detection result are false positives, it is determined that the source code artifact is a legal source code artifact, and the legal source code artifact is saved in the central warehouse 121.

[0126] When the false positive detection result shows that the source code artifact includes malicious code slices, the malicious code slices are sent to the explanation module 705. The explanation module 705 can also obtain the malicious code type of the malicious code slice from the code analysis module 703, and extract the second vector feature from the vector of the malicious code slice; obtain the second prompt information from the template library according to the malicious code type and the second vector feature of the malicious code slice, and input the malicious code slice and the second prompt information into the third large language model, and output the malicious code slice and the detection result explanation through the third large language model. The detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code. The reviewer can view the malicious code slice and the detection result explanation to judge whether the above detection result is correct. When the reviewer judges that all the malicious codes of the source code artifact are false positives, it can be determined that the source code artifact is a legal source code artifact, and the legal source code artifact is saved in the central warehouse 121.

[0127] The code detection device 700 can also implement the functions of the code detection device 210. Refer to Figure 9 , in another embodiment, the slicing module 701 is used to obtain code slices from the source code artifact; the hint module 702 is used to obtain similar samples of the code slices; the code analysis module 703 is used to input the code slices, the programming language identifier of the source code artifact, and the similar samples into the first large language model, and output the detection result of the code slices through the first large language model. When the detection result shows that all the code slices of the source code artifact are legal, it can be determined that the source code artifact is a legal source code artifact, and the legal source code artifact is integrated into the development environment 221.

[0128] When the detection result is used to indicate that the code slice includes malicious code, the malicious code slice is sent to the false positive detection module 704. The false positive detection module 704 extracts the first vector feature from the vector of the code slice; searches for the target false positive code corresponding to the first vector feature in the false positive sample library, inputs the code slice and the target false positive code into the second large language model, and determines whether the malicious code of the code slice is a false positive through the second large language model. When the second large language model determines that all the malicious code of the source code artifact is a false positive, it determines that the source code artifact is a legal source code artifact and integrates the legal source code artifact into the development environment 221.

[0129] When the false positive detection result shows that the source code artifact includes a malicious code slice, the malicious code slice is sent to the explanation module 705, and the second vector feature is extracted from the vector of the malicious code slice; the explanation module 705 obtains the malicious code type of the malicious code slice from the code analysis module 703, obtains the second prompt information from the template library according to the malicious code type and the second vector feature of the malicious code slice, inputs the malicious code slice and the second prompt information into the third large language model, and outputs the malicious code slice and the detection result explanation through the third large language model. The detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code. The reviewer can check the malicious code slice and the detection result explanation to judge whether the above detection result is correct. When the reviewer determines that the malicious code of the source code artifact is a false positive, the reviewer can determine that the source code artifact is a legal source code artifact and integrate the legal source code artifact into the development environment 221.

[0130] This application also provides a computing device 1000 that can implement Figures 1 to 9 the functions of the code detection device in Figure 10 As shown, in one embodiment, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. It should be understood that this application does not limit the number of processors and memories in the computing device 1000.

[0131] The bus 1002 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10It is represented by only one line in the figure, but it does not mean that there is only one bus or one type of bus. The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (for example, the memory 1006, the processor 1004, and the communication interface 1008).

[0132] The processor 1004 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc. The processor includes a plurality of processing cores.

[0133] The memory 1006 may include a volatile memory, such as a random access memory (RAM). The memory 1006 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). In some embodiments, executable program codes are stored in the memory 1006, and the processor 1004 executes the executable program codes to implement the functions of the slicing module 701, the prompting module 702, the code analysis module 703, the false alarm detection module 704, and the interpretation module 705, so as to implement the above-mentioned malicious code detection method.

[0134] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 1000 and other devices or a communication network.

[0135] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0136] As Figure 11 shown, the computing device cluster includes at least one computing device 1000. The memories 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the malicious code detection method.

[0137] Please refer toFigure 12 , Figure 12 This is a schematic diagram showing the connection of computing devices in the computing device cluster of the embodiment of the present application through a network. As Figure 12 shown, two computing devices 1000A and 1000B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0138] In a possible implementation, the memory in computing device 1000A stores instructions for executing the functions of the slicing module 701 and the prompting module 702. At the same time, the memory in computing device 700B stores instructions for executing the functions of the code analysis module 703, the false alarm detection module 704, and the interpretation module 705. It should be understood that Figure 12 the functions of computing device 1000A shown in

[0139] can also be completed by multiple computing devices. Similarly, the functions of computing device 1000B can also be completed by multiple computing devices.

[0140] The embodiment of the present application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the malicious code detection method of the present application.

[0141] The terms "first", "second", etc. in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.

[0142] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A malicious code detection method, characterized in that, The method is applied to a code detection device for providing cloud services, and the method includes: The code detection device obtains a code slice from a source code artifact; The code detection device obtains a similar sample of the code slice; The code detection device inputs the code slice, the programming language identifier of the source code artifact, and the similar sample into a first large language model, and outputs a detection result of the code slice through the first large language model, where the detection result is used to indicate whether the code slice includes malicious code.

2. The method according to claim 1, wherein: The method further includes: the code detection device determines a reference code type corresponding to the code slice; when the reference code type belongs to a malicious code type, the code detection device obtains first prompt information corresponding to the reference code type from a template library; The code detection device inputs the code slice, the programming language identifier of the source code artifact, and the similar sample into a first large language model, and the output of the detection result of the code slice through the first large language model includes: the code detection device inputs the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar sample into the first large language model, and outputs the detection result of the code slice through the first large language model.

3. The method according to claim 2, wherein The code detection device determines the reference code type corresponding to the code slice, including: The code detection device inputs API functions of the code slice into a neural network model, and outputs the reference code type corresponding to the code slice through the neural network model, where the neural network model is trained based on a few-shot learning algorithm and malicious samples.

4. The method according to claim 3, characterized in that The method further includes: When a target malicious code from a user does not belong to the template library, the code detection device obtains prompt information of the target malicious code and the malicious code type of the target malicious code; The code detection device establishes a correspondence between the prompt information of the target malicious code and the malicious code type of the target malicious code in the template library; The code detection device adds the target malicious code to the sample library to which the similar sample belongs.

5. The method according to any one of claims 2 to 4, characterized in that, The first prompt information includes at least one of a definition of the reference code type, an analysis step of the reference code type, or a note of the reference code type.

6. The method according to any one of claims 2 to 5, wherein: The method further includes: the code detection device determines a document set corresponding to the programming language identifier from an API document library; the code detection device obtains a target API document corresponding to the API function of the code slice from the document set; The code detection device inputs the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar sample into a first large language model, including: the code detection device inputs the code slice, the programming language identifier of the source code artifact, the first prompt information, the similar sample, and the target API document into the first large language model.

7. The method according to any one of claims 1 to 6, characterized in that The input data of the first large language model further includes the identifier of the source code artifact.

8. The method according to any one of claims 1 to 7, characterized in that, The code detection device obtaining a code slice from a source code artifact includes: The code detection device determining a source point and a convergence point of the source code artifact according to sensitive interface functions in the source code artifact; The code detection device selecting the code slice from the source code artifact according to the source point and the convergence point of the source code artifact.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: When the code slice includes malicious code, the code detection device extracts a first vector feature from the vector of the code slice; The code detection device searches in a false positive sample library for a target false positive code corresponding to the first vector feature, where the similarity value between the first vector feature and the vector feature of the target false positive code is greater than a reference similarity threshold; The code detection device inputs the code slice and the target false positive code into a second large language model, and determines whether the malicious code in the code slice is a false positive through the second large language model.

10. The method according to claim 9, wherein The method further includes: When the malicious code in the code slice is not a false positive and the detection result includes the target malicious code type of the code slice, the code detection device extracts a second vector feature from the vector of the code slice; The code detection device obtains second prompt information from a template library according to the target malicious code type and the second vector feature, where the second prompt information includes key code features, source code artifact identifiers, function names in the source code artifact, and output languages; The code detection device inputs the code slice and the second prompt information into a third large language model, and outputs a detection result explanation through the third large language model, where the detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code.

11. The method according to claim 10, wherein The second prompt information further includes analysis steps corresponding to the target malicious code type, and the detection result explanation further includes a summary of the analysis steps corresponding to the target malicious code type.

12. A code detection device, characterized in that, Including: A slicing module, configured to obtain a code slice from a source code artifact; A prompt module, configured to obtain a similar sample of the code slice; A code analysis module, configured to input the code slice, the programming language identifier of the source code artifact, and the similar sample into a first large language model, and output a detection result of the code slice through the first large language model, where the detection result is used to indicate whether the code slice includes malicious code.

13. The device according to claim 12, wherein The prompt module is further configured to determine a reference code type corresponding to the code slice; when the reference code type is a malicious code type, obtain first prompt information corresponding to the reference code type from a template library; The code analysis module is specifically configured to input the code slice, the programming language identifier of the source code artifact, the first prompt information, and the similar sample into a first large language model, and output a detection result of the code slice through the first large language model.

14. The device according to claim 13, characterized in that, The prompt module is specifically configured to input the application programming interface (API) functions of the code slice into a neural network model, and output the reference code type corresponding to the code slice through the neural network model. The neural network model is trained based on a few-shot learning algorithm and malicious samples.

15. The device according to claim 14, characterized in that, When the target malicious code from the user does not belong to the template library, the prompt module is further configured to obtain the prompt information of the target malicious code and the malicious code type of the target malicious code, establish a correspondence between the prompt information of the target malicious code and the malicious code type of the target malicious code in the template library, and add the target malicious code to the sample library to which the similar samples belong.

16. The device according to any one of claims 13 to 15, characterized in that The first prompt information includes at least one of the definition of the reference code type, the analysis steps of the reference code type, or the precautions of the reference code type.

17. The device according to any one of claims 13 to 16, wherein The prompt module is further configured to determine the document set corresponding to the programming language identifier from the API document library, and obtain the target API document corresponding to the API function of the code slice from the document set. The code analysis module is specifically configured to input the code slice, the programming language identifier of the source code artifact, the first prompt information, the similar samples, and the target API document into a first large language model.

18. The device according to any one of claims 12 to 17, characterized in that, The input data of the first large language model further includes the identifier of the source code artifact.

19. The device according to any one of claims 12 to 18, characterized in that, The slicing module is specifically configured to determine the source point and the convergence point of the source code artifact according to the sensitive interface functions in the source code artifact; select the code slice from the source code artifact according to the source point and the convergence point of the source code artifact.

20. The device according to any one of claims 12 to 19, characterized in that The device further includes: A false positive detection module, configured to, when the code slice includes malicious code, extract a first vector feature from the vector of the code slice; search for a target false positive code corresponding to the first vector feature in the false positive sample library, where the similarity value between the first vector feature and the vector feature of the target false positive code is greater than a reference similarity threshold; input the code slice and the target false positive code into a second large language model, and determine whether the malicious code in the code slice is a false positive through the second large language model.

21. The device according to claim 20, characterized in that, The device further includes: An explanation module, configured to, when the malicious code in the code slice is not a false positive and the detection result includes the target malicious code type of the code slice, extract a second vector feature from the vector of the code slice; obtain second prompt information from the template library according to the target malicious code type and the second vector feature, where the second prompt information includes key code features, the identifier of the source code artifact, the function name in the source code artifact, and the output language; input the code slice and the second prompt information into a third large language model, and output a detection result explanation through the third large language model, where the detection result explanation is used to indicate the judgment result of the key code features and the harm of the malicious code.

22. The device according to claim 21, characterized in that, The second prompt information further includes the analysis steps corresponding to the target malicious code type, and the detection result explanation further includes an outline of the analysis steps corresponding to the target malicious code type.

23. A cluster of computing devices, characterized in that, Comprising at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 11.

24. A computer-readable storage medium, characterized in that, Including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 11.

25. A computer program product comprising instructions, characterized in that, When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 11.