Malicious code detection method and device, electronic equipment and storage medium

CN122818355APending Publication Date: 2026-09-25CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610933966.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供一种恶意代码检测方法、装置、电子设备和存储介质,用于提供一种更智能、更高效、且具备自适应学习能力的检测方法,以专门解决大语言模型生成恶意代码的安全问题

Benefits of technology

[0016]本申请提供的恶意代码检测方法、装置、电子设备和存储介质,第一检测模块能从代码语法、逻辑和深层语义层面深度挖掘经过混淆、变形的新型恶意代码,而第二检测模块则能利用其在通用文本审核方面的优势,快速识别已知的恶意关键词或不合规内容;通过结合专用的代码深度分析能力和通用的内容审核能力,构建了并行的、多层次的检测架构;通过设置风险等级并引入代码沙箱对中风险样本进行精准验证,有效平衡了检测效率与准确性,避免了对所有代码都进行耗时的沙箱分析;通过利用沙箱验证结果对代码语义分析模型和规则库进行闭环反馈和动态调整,使得整个检测系统具备了持续学习和自我进化的能力,能够有效对抗不断变化的恶意代码变种,能够有效地检测和拦截恶意代码,显著提升了大语言模型输出内容的安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818355A_ABST
    Figure CN122818355A_ABST
Patent Text Reader

Abstract

The application provides a malicious code detection method and device, electronic equipment and storage medium, relates to the technical field of artificial intelligence, and comprises the following steps: inputting a code generated by a large language model into a first detection module and a second detection module in parallel to obtain a first risk detection result and a second risk detection result; the first detection module comprises a code semantic analysis model and a malicious code rule library; the second detection module is a general content review engine; determining the risk level of the code based on the first risk detection result and the second risk detection result; inputting the code into a code sandbox to obtain a code verification result output by the code sandbox; and adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule library based on the code verification result. The method and device provided by the application can effectively counter the constantly changing malicious code variants, effectively detect and intercept malicious codes, and significantly improve the security of the content output by the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for detecting malicious code. Background Technology

[0002] In recent years, artificial intelligence technologies, represented by Large Language Models (LLMs), have made groundbreaking progress and have been widely applied in various scenarios such as content creation, intelligent question answering, and code generation. However, LLMs may generate malicious code, such as Structured Query Language (SQL) injection statements, cross-site scripting (XSS), and attack payloads, under the inducement or unintentional prompting of users. Once executed, this malicious code poses a serious threat to downstream applications and user data security.

[0003] Related technologies are typically based on rule bases, keyword matching, and general natural language semantic understanding to identify and intercept known explicit malicious code. However, they cannot identify unknown implicit malicious code, resulting in low identification efficiency, high resource consumption, and an inability to effectively cope with rapidly iterating attack methods.

[0004] Therefore, how to provide a more intelligent, efficient detection method with adaptive learning capabilities to specifically address the security problem of malicious code generated by large language models has become an urgent technical issue for the industry. Summary of the Invention

[0005] This application provides a malicious code detection method, apparatus, electronic device, and storage medium, which are used to provide a more intelligent, efficient detection method with adaptive learning capabilities, specifically to solve the security problem of malicious code generated by large language models.

[0006] This application provides a method for detecting malicious code, including: The code generated by the large language model is input in parallel into the first detection module and the second detection module to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base; the second detection module is a general content moderation engine. Based on the first risk detection result and the second risk detection result, the risk level of the code is determined; If the risk level of the code is a preset risk level, the code is input into the code sandbox, and the code verification result output by the code sandbox is obtained; Based on the code verification results, adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base.

[0007] In some embodiments, adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results includes: A reward signal is determined based on the code verification results; Reinforcement learning is performed based on the reward signal, state space, and action space to adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base; The actions in the action space are used to trigger rule adjustments to the malicious code rule base and / or parameter adjustments to the code semantic analysis model. The state vector in the state space includes at least one of the following: The code verification result includes at least one of the following: malicious label, malicious behavior type, and malicious behavior confidence level. The code semantic analysis model has at least one of the following: detection accuracy, false negative rate, false positive rate, and obfuscated code recognition rate. The malicious code rule base includes at least one of the following: template hit rate, proportion of high-risk templates, coverage rate of benign templates, and template mismatch rate. Accuracy changes after the last round of optimization; The previous round of optimization took a long time.

[0008] In some embodiments, determining the reward signal based on the code verification result includes: Based on the code verification results, the correctly identified code samples, the samples missed by the code semantic analysis model, and the samples misjudged by the code semantic analysis model are determined. Based on the correctly identified code samples, the detection accuracy rate is determined; Based on the samples missed by the code semantic analysis model, the false detection rate is determined. Based on the samples misjudged by the code semantic analysis model, the detection misjudgment rate is determined; Based on the total time spent adjusting the malicious code rule base and adjusting the code semantic analysis model, the processing latency cost is determined. The reward signal is determined based on the detection accuracy, the false negative rate, the false positive rate, and the processing delay cost.

[0009] In some embodiments, adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results includes: Based on the code verification results, identify the malicious code samples that the code semantic analysis model missed and / or the benign code samples that it misjudged, and construct an incremental training dataset. The code semantic analysis model is fine-tuned based on the incremental training dataset.

[0010] In some embodiments, the code semantic analysis model includes a feature extraction network and a feature classification network; The fine-tuning of the code semantic analysis model based on the incremental training dataset includes: Freeze the model parameters of the feature extraction network; The model parameters of the feature classification network are adjusted based on the incremental training dataset.

[0011] In some embodiments, the fusion loss function for fine-tuning the code semantic analysis model is determined based on the following steps: The log-likelihood terms corresponding to malicious code samples and benign code samples are calculated by category weighting to determine the weighted cross-entropy loss term; The log-likelihood terms of the obfuscated code samples are summed to determine the obfuscated sample loss term; The weighted cross-entropy loss term and the confused sample loss term are weighted and fused to determine the fusion loss function.

[0012] In some embodiments, the category weights of the log-likelihood terms of the malicious code sample and the benign code sample are dynamically adjusted based on the current false negative rate of the code semantic analysis model and the probability of an increase in the false positive rate of the benign code sample after fine-tuning. The weight of the obfuscated sample loss term in the fusion loss function is dynamically adjusted based on the current obfuscated code recognition rate of the code semantic analysis model and the probability of an increase in the false positive rate of benign code samples after fine-tuning.

[0013] This application provides a malicious code detection device, including: The detection module is used to input the code generated by the large language model into the first detection module and the second detection module in parallel to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base; the second detection module is a general content moderation engine. A determination module is used to determine the risk level of the code based on the first risk detection result and the second risk detection result; The verification module is used to input the code into the code sandbox when the risk level of the code is a preset risk level, and obtain the code verification result output by the code sandbox; An adjustment module is used to adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results.

[0014] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the malicious code detection method described above.

[0015] This application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the malicious code detection method described above.

[0016] The malicious code detection method, device, electronic device, and storage medium provided in this application have the following features: The first detection module can deeply mine obfuscated and transformed novel malicious code from the levels of code syntax, logic, and deep semantics; the second detection module can leverage its advantages in general text review to quickly identify known malicious keywords or non-compliant content. By combining dedicated deep code analysis capabilities with general content review capabilities, a parallel, multi-layered detection architecture is constructed. By setting risk levels and introducing a code sandbox for accurate verification of medium-risk samples, detection efficiency and accuracy are effectively balanced, avoiding time-consuming sandbox analysis of all code. By utilizing sandbox verification results for closed-loop feedback and dynamic adjustment of the code semantic analysis model and rule base, the entire detection system possesses the ability to continuously learn and self-evolve, effectively combating constantly evolving malicious code variants, effectively detecting and intercepting malicious code, and significantly improving the security of the output content of the large language model. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the malicious code detection method provided in this application.

[0020] Figure 2 This is a schematic diagram of the malicious code detection process provided in this application.

[0021] Figure 3 This is a schematic diagram of the malicious code detection device provided in this application.

[0022] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps, units, or modules is not necessarily limited to those explicitly listed, but may include other steps, units, or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0025] In related technologies, large language models may generate malicious code or attack instructions based on user commands, as well as cross-site attack scripts based on user input. Due to their own characteristics, large language models can generate malicious variants, such as splitting SQL injection statements into multi-turn dialogue generation, using nested functions, mixed case, comment insertion, and replacing spaces with tabs / newlines, etc.

[0026] General-purpose content moderation engines in related technologies focus on semantic intent judgment through multi-round context analysis, as well as content moderation using rule bases, keyword matching, and complete semantic chains, with the aim of detecting malicious code. However, these engines suffer from a lack of code syntax parsing and semantic reasoning capabilities, and slow iteration and updates of adversarial examples can lead to detection failures.

[0027] In order to address the shortcomings of related technologies, Figure 1 This is a flowchart illustrating the malicious code detection method provided in this application, as follows: Figure 1 As shown, the method includes steps 110, 120, 130 and 140.

[0028] Step 110: Input the code generated by the large language model into the first detection module and the second detection module in parallel to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base; the second detection module is a general content moderation engine.

[0029] Specifically, the malicious code detection method provided in this application is executed by a malicious code detection device or system. This device can be implemented in software, such as a malicious code detection program; or it can be a device that executes the malicious code detection method, such as a terminal, computer, or server.

[0030] The first detection module is dedicated to in-depth analysis of code content, and it includes a code semantic analysis model and a malware rule base. The code semantic analysis model and the malware rule base can perform in-depth analysis of code content in parallel or sequentially.

[0031] Code semantic analysis models aim to understand the true intent of code at the semantic level, and are particularly adept at identifying malicious code that has been obfuscated, modified, or written in unconventional ways. In one specific embodiment, this model can be a Bidirectional Encoder Representations from Transformers (BERT) model based on the Transformer architecture. This model takes code generated by a large language model as input and outputs a probability value or risk score indicating that the code is malicious; this score is part of the first risk detection result.

[0032] The malware rule base (also known as the rule template library) stores a large number of known and clearly defined malware features, patterns, or signature templates. This rule base identifies whether code contains these known malicious features through rapid template matching. The detection result can be a boolean value (yes / no) or a risk score based on the severity level of the rule; this result is also part of the initial risk detection result. Finally, the initial detection module integrates the analysis results from the code semantic analysis model and the malware rule base to output a comprehensive initial risk detection result.

[0033] The second detection module is a general-purpose content moderation engine. Unlike the focused first detection module, the general-purpose content moderation engine has a broader detection scope and is typically used for risk assessment of multimodal content such as text, images, audio, and video. In the scenario of this application embodiment, it mainly utilizes its mature text moderation capabilities to analyze the code generated by the large language model from dimensions such as general semantics and keyword matching, especially the text content such as comments and strings in the code. The engine will output its own second risk detection result, such as a risk score or risk label.

[0034] By feeding the code into these two modules in parallel, their strengths can be complemented: the first detection module performs in-depth and professional analysis of the code's malicious behavior, while the second detection module conducts a broad-spectrum compliance review of the code's content.

[0035] Step 120: Determine the risk level of the code based on the results of the first and second risk detections.

[0036] Specifically, firstly, the risk detection results output by the two modules can be normalized. For example, they can be uniformly mapped to a score range of [0, 10], where a higher score indicates a greater risk.

[0037] Then, the final risk level is determined according to the preset merging strategy. The risk level can be divided into multiple levels, such as high risk, medium risk, and low risk.

[0038] (1) High Risk: When either the first risk detection result or the second risk detection result exceeds the preset high risk threshold (e.g., 8 points), the code is judged as high risk. This strategy is called "high risk fallback" to ensure that no clear risk discovered by any module is ignored. High-risk code is usually directly blocked and a compliance prompt is returned to the user.

[0039] (2) Low risk: When both the first risk detection result and the second risk detection result are lower than the preset low risk threshold (e.g., 3 points), the code is judged as low risk. Low-risk code is considered safe and can be directly output to the user.

[0040] (3) Medium risk: When the risk score of the code does not fall into the categories of high risk or low risk (for example, the overall score is between 4 and 7), it is judged as medium risk. Medium risk means that the code has suspicious characteristics, but the current module cannot determine its malice with 100% certainty and further verification is required.

[0041] Step 130: If the risk level of the code is the preset risk level, input the code into the code sandbox and obtain the code verification result output by the code sandbox.

[0042] Specifically, a code sandbox is a secure, isolated virtual environment that simulates a real code execution environment (such as a server environment or browser environment). A code sandbox can perform static analysis (such as syntax checking and logic vulnerability scanning) and dynamic execution on the input code. Through dynamic execution, the actual behavior of the code at runtime can be directly observed, such as whether it attempts to read or write files, initiates network connections, or performs database injection.

[0043] Code sandboxes analyze these actual behaviors to output a high-precision code verification result. This result is usually a clear label, such as "confirmed as malicious code" or "confirmed as benign code." For malicious code, the verification result can further include the specific type of malicious behavior, such as "SQL injection," "cross-site scripting," or "command execution." This verification result can be considered the ground truth label for that piece of code.

[0044] Step 140: Based on the code verification results, adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base.

[0045] Specifically, scenarios for adjusting the parameters of a code semantic analysis model can include: (1) Missed Detection: If a piece of code is judged as low-risk or medium-risk by the code semantic analysis model, but is verified as malicious code by the code sandbox, then this piece of code and its malicious label constitute a negative sample. This sample can be added to the model's training dataset, and the model's parameters can be updated through further training or fine-tuning, so that the model can identify similar malicious code in the future.

[0046] (2) False positives: If a piece of code is judged as high-risk by the model, but is verified as benign code by the sandbox, then this piece of code and its benign label constitute a positive sample. Similarly, this sample can be used for model retraining to reduce the false positive rate of the model for normal code.

[0047] Scenarios for adjusting rules in a malicious code rule base can include: (1) Rule optimization: If a rule frequently leads to misjudgment of benign code (verified by the sandbox), the system can reduce the priority of the rule, adjust its matching threshold, or have a security expert intervene to correct it.

[0048] (2) Rule supplementation: If the sandbox discovers a new type of malicious code pattern that is not covered in the rule base, the system can automatically or semi-automatically generate new rules based on the characteristics of the code and add them to the rule base, thereby enhancing the ability to detect unknown threats.

[0049] The malicious code detection method provided in this application has a first detection module that can deeply mine obfuscated and transformed new malicious code from the levels of code syntax, logic, and deep semantics, while the second detection module can quickly identify known malicious keywords or non-compliant content by leveraging its advantages in general text review. By combining dedicated deep code analysis capabilities with general content review capabilities, a parallel, multi-layered detection architecture is constructed. By setting risk levels and introducing a code sandbox to accurately verify medium-risk samples, the detection efficiency and accuracy are effectively balanced, avoiding time-consuming sandbox analysis of all code. By using the sandbox verification results to provide closed-loop feedback and dynamic adjustment to the code semantic analysis model and rule base, the entire detection system has the ability to continuously learn and self-evolve, effectively combating constantly changing malicious code variants, effectively detecting and intercepting malicious code, and significantly improving the security of the output content of the large language model.

[0050] It should be noted that each implementation method of this application can be freely combined, rearranged, or executed individually, and does not need to rely on or depend on a fixed execution order.

[0051] In some embodiments, based on code verification results, adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base includes: Reward signals are determined based on code verification results; Reinforcement learning is performed based on reward signals, state space, and action space to adjust the parameters of the code semantic analysis model and / or the rules of the malware rule base; In the action space, actions are used to trigger rule adjustments to the malicious code rule base and / or parameter adjustments to the code semantic analysis model; The state vector in the state space includes at least one of the following: At least one of the following in the code verification results: malicious label, malicious behavior type, and malicious behavior confidence level; At least one of the following: detection accuracy, false negative rate, false positive rate, and obfuscated code recognition rate of the code semantic analysis model; At least one of the following: template hit rate, proportion of high-risk templates, coverage rate of benign templates, and template mismatch rate in the malicious code rule base; Accuracy changes after the last round of optimization; The previous round of optimization took a long time.

[0052] Specifically, reinforcement learning can be performed by an agent configured within the system. The reward signal is the core driving force in the reinforcement learning process; it is used to quantitatively evaluate the quality of the decisions made by the system in the previous stage (i.e., the adjustments taken).

[0053] The reward signal is determined based on the highly reliable code verification results output by the code sandbox. Code verification results can include information such as malicious code and benign code. Based on the code verification results, further identification can be made of correctly identified code samples, samples missed by the code semantic analysis model, and samples misidentified by the code semantic analysis model.

[0054] In one specific embodiment, the reward signal can be designed as a comprehensive reward function, expressed by the formula: .

[0055] in, The value of the reward function; , , and These are weighting coefficients, all ranging from [0,1]. The accuracy rate is calculated as the number of correctly identified code samples divided by the total number of detected code samples. To detect the false negative rate, the calculation method is the number of samples that were identified as malicious in the sandbox but were missed by the code semantic analysis model / the total number of malicious code samples; To detect the false positive rate, the calculation method is the number of samples that were deemed benign in the sandbox validation but were misjudged by the code semantic analysis model / the total number of benign code samples; To account for latency costs, the calculation method is the total time / maximum time spent adjusting the malicious code rule base and the code semantic analysis model using reinforcement learning.

[0056] The goal of this reward function is to improve the accuracy of malware detection and reduce the false negative / false positive rate.

[0057] In a reinforcement learning framework, an agent observes the state of the environment at each time step, selects an action to perform based on its policy, and the environment transitions to the next state and provides a reward. This process iterates continuously, and the agent's policy is continuously optimized.

[0058] The state space is a collection of environmental states. In this embodiment, it is defined as a multi-dimensional vector that comprehensively describes the current performance and health of the detection system. This state vector provides the reinforcement learning agent with all the necessary information for decision-making.

[0059] In this embodiment of the application, a 13-dimensional normalized state vector can be defined. The definitions of each dimension are shown in Table 1.

[0060] Table 1 State Space Definition Table

[0061] Action space is the set of operations that a reinforcement learning agent can perform. In the embodiments of this application, these actions are used to specifically trigger rule adjustments to the malicious code rule base and / or parameter adjustments to the code semantic analysis model.

[0062] In this embodiment of the application, a 15-dimensional discrete action set can be defined. The definitions of each dimension are shown in Table 2.

[0063] Table 2 Action Space Definition Table

[0064] During operation, the reinforcement learning agent continuously observes the state vector composed of the above elements and selects the action most likely to yield a high reward based on its learned policy. After executing the action, the system state changes, the agent receives a reward, and uses this experience (state, action, reward, new state) to update its internal policy. This process is repeated until the agent's policy gradually converges, enabling it to make optimal adjustment decisions under different system states.

[0065] The malicious code detection method provided in this application transforms the optimization process, which originally required complex analysis and manual adjustment, into an automated closed-loop system driven by reinforcement learning. It can monitor its own performance in real time and make intelligent and dynamic self-adjustments based on actual data feedback, thereby responding more agilely to new attack methods, maintaining a high level of detection performance, and significantly improving the robustness and adaptability of malicious code detection.

[0066] In some embodiments, reinforcement learning is performed based on reward signals, state space, and action space to adjust the parameters of the code semantic analysis model and / or the rules of the malware rule base, including: The policy network and value network are determined; the policy network is used to determine the selection probability distribution of each action in the action space based on the state vectors in the state space; the value network is used to determine the state value estimate based on the state vectors in the state space. Advantage estimation is determined based on reward signals and state value estimation; Based on advantage estimation and action probability distribution, the shearing loss function is determined through shearing function constraints; The value loss function is determined based on advantage estimation and state value estimation; The parameters of the policy network and the value network are updated based on the shearing loss function and the value loss function.

[0067] Specifically, the environment in which the intelligent agent performs reinforcement learning The definition is as follows: (1) Define the state space Action space and reward function .

[0068] (2) Define the state update strategy Describes the state change caused by threshold adjustment after an action is performed, expressed mathematically as follows: .in, For state update functions; To perform the action After adjusting the trigger rule template library and fine-tuning the model, the state is recalculated.

[0069] The process of an agent performing reinforcement learning (taking a single step as an example) is as follows: (1) Enter the current time step status and actions ; (2) Execute actions to adjust the triggered state changes and calculate the new state. Recalculate status values; calculate rewards based on current status and actions. ; (3) Output and .

[0070] This application employs a reinforcement learning method based on the Proximal Policy Optimization (PPO) algorithm. The reinforcement learning agent consists of two deep neural networks: a policy network and a value network.

[0071] The policy network (Actor) outputs better model parameters or a rule template library to optimize actions. This network receives the state vector of the current environment. As input, output a probability distribution of each action that should be taken in this state. . For action, These are the model parameters (weights and biases) for the policy network.

[0072] The policy network can be a fully connected feedforward neural network. For example, it can contain an input layer (13 dimensions), two hidden layers (e.g., 256 and 128 dimensions, using ReLU as the activation function to introduce non-linearity), and an output layer.

[0073] The output layer dimension of the policy network corresponds to the dimension of the action space (e.g., 15 dimensions for 15 selectable actions), and a softmax activation function is used. The softmax function transforms the raw scores of the output layer into a probability distribution. Each element represents the probability of choosing the corresponding action, and the sum of all elements is 1. This can be represented as: .

[0074] in, , and These are the neural network weight matrices for the input layer and the two hidden layers, respectively. , and These are the bias vectors for the input layer and the two hidden layers, respectively.

[0075] The core function of a value network (Critic) is to evaluate the long-term value of being in a specific state under a given policy, that is, the expected cumulative reward that can be obtained starting from that state.

[0076] The input to the value network is the same as that to the policy network, which is also a state vector. .

[0077] The structure of a value network can be similar to that of a policy network; for example, it may also contain an input layer and two hidden layers. Unlike a policy network, the output layer of a value network has only one neuron, uses no activation function or uses a linear activation function, and its output is a scalar value. This scalar value is the state value estimate for the current state S. Here... These represent the model parameters of the value network. They can be expressed as: .

[0078] in, , and These are the neural network weight matrices for the input layer and the two hidden layers, respectively. , and These are the bias vectors for the input layer and the two hidden layers, respectively.

[0079] The agent interacts with the environment (i.e., the system processes code, the sandbox provides verification results, and the system state is updated) and collects a series of experiences (state). ,action ,award Next state After that, it is necessary to calculate the dominance estimate. Advantage estimation measures the performance of a given state. Next, take action How much better is it compared to the average performance of following the current strategy?

[0080] In a preferred embodiment, the Generalized Advantage Estimation (GAE) method can be used to calculate the bias and variance.

[0081] First, calculate the timing difference error. : .

[0082] in, It is to perform an action Immediate reward signals from post-environmental feedback; This is an optimization factor used to balance the importance of current and future rewards; for example, it can be set to 0.99. and These are the value estimates of the next state and the current state by the value network.

[0083] Then, the dominance estimate is calculated based on the time-series difference error. : .

[0084] in, Indicates the time step offset; This represents the total time steps of a single-round experience trajectory; This is a coefficient used to adjust the balance between bias and variance; for example, a value of 0.95.

[0085] To improve the stability of training, the calculated advantage estimate can be normalized, as expressed by the formula: .

[0086] in, This indicates that the advantage estimate is taken as the mean. This represents the standard deviation of the advantage estimate.

[0087] The core innovation of the PPO algorithm lies in its shearing loss function. It ensures training stability by limiting the update magnitude of the old and new strategies, which can be expressed by the formula: .

[0088] .

[0089] in, This represents the expectation of samples across all time steps within a batch. The probability ratio between the old and new strategies; This represents the probability distribution of each action under the new strategy; This represents the probability distribution of each action under the old strategy; These are the parameters of the policy network before the update; Let be the shearing function, and let the probability ratio be... Limited to [ Within the range of ]. is a hyperparameter called the shear coefficient, for example, with a value of 0.2. The negative sign indicates that minimizing the loss is equivalent to maximizing the cumulative reward.

[0090] Value loss function The goal is to make the predicted value of the value network To approximate the actual observed return as closely as possible, expressed by the formula: .

[0091] .

[0092] in, The actual observed return.

[0093] Define a composite loss function . It is a weighting factor (e.g., with a value of 1) used to balance the importance of the two loss terms.

[0094] Using a gradient descent optimization algorithm (such as the Adam optimizer), the gradient is calculated based on the composite loss function, and the parameters of the policy network are updated. and parameters of the value network .

[0095] Policy network update is represented as: .in This represents the policy learning rate, which can be 10. -4 ; Value network updates are represented as follows: ,in This represents the value learning rate, which can be 10. -3 .

[0096] Repeated updates Wheels can be set .

[0097] The malicious code detection method provided in this application provides a stable and efficient training method for the policy network and value network. This enables the system to intelligently select the most appropriate adjustment action (such as adjusting rules, fine-tuning the model, etc.) based on real-time changes in the system state (such as false negative rate, false positive rate, etc.). This achieves automated, refined, and intelligent adjustment of the parameters of the entire malicious code detection device, significantly improving the system's adaptability and defense effectiveness.

[0098] In some embodiments, based on code verification results, adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base includes: Based on the code verification results, identify the malicious code samples that the code semantic analysis model missed and / or the benign code samples that it misjudged, and construct an incremental training dataset. The code semantic analysis model was fine-tuned based on the incremental training dataset.

[0099] Specifically, based on the code verification results, malicious code samples that were missed (also known as negative samples or difficult negative samples) and benign code samples that were misjudged (also known as positive samples or difficult positive samples) can be identified.

[0100] When code generated by a large language model is judged as low-risk or no-risk by the code semantic analysis model during the first detection module, but is confirmed as malicious code by the code sandbox, a false positive occurs. The system will record this falsely identified malicious code and attach its malicious label determined by the sandbox.

[0101] A misjudgment occurs when a piece of code is identified as high-risk by the code semantic analysis model, but is subsequently verified as benign code by the code sandbox. The system will also record this misjudged benign code and assign it a benign label.

[0102] The system aggregates all the malicious code samples that were missed and the benign code samples that were misjudged during the above process to form an incremental training dataset. This dataset is specifically used for the subsequent fine-tuning phase, rather than for full training from scratch.

[0103] Once the incremental training dataset is obtained, the fine-tuning process for the code semantic analysis model can begin. Fine-tuning refers to retraining a pre-trained model on a small scale and for a short period using a new, specific dataset to adapt the model to new tasks or improve its performance in specific scenarios.

[0104] The specific process of fine-tuning may include: (1) Data preprocessing: Clean and format the code samples in the incremental training dataset as necessary. For example, unify the newline characters and spaces in the code; for samples that have been encoded and obfuscated, they can be decoded first, and then the original encoded text and the decoded text can be concatenated to help the model learn the semantic relationship before and after encoding. Subsequently, the dataset can be divided into training set, validation set and test set according to a certain ratio (e.g., 5:4:1).

[0105] (2) Model Loading and Training: Load the parameters of the code semantic analysis model of the current online service. Then, train the model using the training set prepared in the previous step. During training, the model processes these samples in batches, calculates the loss (i.e., error) between its prediction results and the true labels (from the sandbox), and adjusts the model's internal parameters (such as weights and biases) through backpropagation algorithm and optimizer (such as Adam) with the goal of minimizing the value of the loss function.

[0106] (3) Model Evaluation and Deployment: After several training rounds, the performance of the fine-tuned model is evaluated using a validation set to monitor whether it has improved in key metrics such as false negative rate and false positive rate. The fine-tuning process can be terminated when the model's performance reaches the preset target (e.g., a false negative rate of less than 1%, a false positive rate of less than 0.5%, and an obfuscated code recognition rate of more than 90% on the test set). Finally, the new model parameters that have been fine-tuned and validated are deployed online to replace the original model, thus completing one iteration of optimization. If the target is not achieved, the previous step can be returned, and the training parameters (such as the learning rate) can be adjusted or the incremental training dataset can be expanded before fine-tuning is performed again.

[0107] The malicious code detection method provided in this application accurately identifies and utilizes the missed and misjudged samples of the model for targeted fine-tuning, avoiding the high computational and time costs required for full retraining of the entire model; enabling the code semantic analysis model to quickly learn from errors and continuously adapt to the ever-emerging new and variant malicious code, thereby significantly improving the accuracy and response speed of malicious code detection.

[0108] In some embodiments, the code semantic analysis model includes a feature extraction network and a feature classification network; Fine-tuning of the code semantic analysis model based on the incremental training dataset includes: Freeze the model parameters of the feature extraction network; The model parameters of the feature classification network were adjusted based on the incremental training dataset.

[0109] Specifically, the code semantic analysis model can be divided into two parts: a feature extraction network and a feature classification network.

[0110] The core responsibility of the feature extraction network is to read the input code text sequence and, through complex self-attention mechanisms and feedforward networks, generate a feature vector rich in deep contextual semantic information for each word in the code sequence. These bottom and middle layers learn general code / language grammatical structures and basic semantic patterns, exhibiting strong generalization capabilities.

[0111] A feature classification network refers to one or more fully connected layers added on top of a feature extraction network, and is often called a classification head. This network receives an aggregated feature vector representing the entire code sequence from the output of the feature extraction network and maps it to the final classification task space.

[0112] Taking the BERT model as an example, the bottom 6-8 layers are usually feature extraction networks, while the top 3 layers and the classification head are usually feature classification networks.

[0113] Freezing refers to setting the parameters of a specific part of the network to an untrainable state. That is, during the backpropagation phase of the training process, the gradients of these frozen parameters are not calculated, and the optimizer does not update them.

[0114] In this embodiment, when fine-tuning begins, all or most of the model parameters of the feature extraction network are first frozen. The underlying principle is that the feature extraction network, through pre-training on massive amounts of data, has learned very powerful and general code semantic representation capabilities. These capabilities are also crucial for identifying malicious code. In fine-tuning scenarios with only a few incremental samples, completely retraining this part of the network would not only incur enormous computational overhead but also pose a risk of "catastrophic forgetting."

[0115] In a specific embodiment, the fine-tuning steps are as follows: (1) Model initialization: Load the BERT pre-trained model; freeze the parameters of the bottom 8 layers, and only fine-tune the top 3 layers + classification head.

[0116] (2) Input encoding: For cases where malicious code is usually short, we can assume that the maximum sequence length is 128 bits and retain the original form of special characters.

[0117] (3) Feature extraction: The input encoding is processed by the BERT encoder to output a semantic vector (768 dimensions); the semantic vector is pooled (mean pooling) to reduce the dimension to 128.

[0118] (4) Forward propagation: BERT outputs the predicted score (the raw score of the unactivated component); (5) Gradient clipping: Limit the gradient norm to within 1.0 to avoid gradient explosion; (6) Learning rate adjustment: The learning rate is gradually reduced using a linear scheduler, and in the last two rounds it is reduced to 10% of the initial value; (7) Model evaluation: Calculate the false negative rate, false positive rate, and obfuscated code recognition rate using the test set; if the indicators are not met, return to step (2) and adjust the data augmentation strategy.

[0119] The termination condition is: the number of training rounds reaches the preset number or the model performance reaches the predetermined target.

[0120] The malicious code detection method provided in this application significantly reduces the number of parameters that need to be updated during the fine-tuning process, thereby greatly improving training efficiency. By limiting the range of parameter updates, the model is forced to retain its general feature extraction capabilities while avoiding catastrophic forgetting, and only its top-level decision logic is adjusted, thereby improving the model's generalization ability on unseen samples.

[0121] In some embodiments, the fusion loss function for fine-tuning the code semantic analysis model is determined based on the following steps: The log-likelihood terms corresponding to malicious code samples and benign code samples are calculated by category weighting to determine the weighted cross-entropy loss term; The log-likelihood terms of the obfuscated code samples are summed to determine the obfuscated sample loss term; The weighted cross-entropy loss term and the confused sample loss term are weighted and fused to determine the fusion loss function.

[0122] Specifically, based on the semantic characteristics of malicious code identification, a fusion loss function can be constructed to fine-tune the code semantic analysis model. This can be expressed as a formula: .

[0123] in, The labels are the real labels for the samples: 0 = benign code, 1 = malicious code; This represents the probability value (probability of being predicted as malicious) of the code semantic analysis model classification head output after processing by the activation function, and its range is [0, 1]. Represents the natural logarithm function; This represents the summation function; Indicates the first The probability that an obfuscated sample in a code sandbox is predicted as malicious by the code semantic analysis model ranges from [0, 1].

[0124] This is the weighted cross-entropy loss term. This is the log-likelihood term for the malicious code sample. Calculate the weights for the category of the log-likelihood term of the malicious code sample. This is the log-likelihood term for benign code samples. Calculate the weights for the category of the log-likelihood term for benign code samples. Increase or decrease to control the impact when the proportion of malicious code samples is high or low.

[0125] To obfuscate the sample loss term. This is the log-likelihood term for the obfuscated code sample. To avoid obfuscating the weights of the sample loss term in the fusion loss function. Increasing or decreasing the value controls the impact of confused samples causing an increase or decrease in misjudgments.

[0126] In some embodiments, the category weights of the log-likelihood terms for malicious code samples and benign code samples are dynamically adjusted based on the current false negative rate of the code semantic analysis model and the probability of an increase in the false positive rate of benign code samples after fine-tuning. The weight of the obfuscated sample loss term in the fusion loss function is dynamically adjusted based on the current obfuscated code recognition rate of the code semantic analysis model and the probability of an increase in the false positive rate of benign code samples after fine-tuning.

[0127] Specifically, the weights of the aforementioned fusion loss function and The adjustments and optimizations are made using actions a9 (strengthening the semantic weight of malicious keywords) and a10 (reducing the semantic weight of obfuscated characters) of the PPO algorithm employed in reinforcement learning. The step size for each adjustment is calculated as follows: Weight Dynamic update step size .

[0128] in, For weight The base step size can be set to 0.1; Th20 is the current false negative rate of the code semantic analysis model; Th20 is the threshold of the false negative rate of the code semantic analysis model, which can be set to 10%; L1 represents the probability of an increase in the false negative rate of benign samples after adjustment, calculated based on historical values, and its value range is [0, 1].

[0129] Weight Dynamic update step size .

[0130] in, For weight The base step size can be set to 0.1; Th21 is the current obfuscated code recognition rate of the code semantic analysis model; Th21 is the threshold of the obfuscated code recognition rate of the code semantic analysis model, which can be set to 10%; L2 represents the probability of an increase in the false positive rate of benign samples after adjustment, calculated based on historical values, with a value range of [0, 1].

[0131] The malicious code detection method provided in this application, by fusing loss functions, achieves refined and intelligent guidance for the model fine-tuning process. It not only solves the potential imbalance between positive and negative samples in incremental training data but also specifically sets optimization targets for code obfuscation in malicious code attacks. The dynamic weight adjustment mechanism allows the loss function to perceive the model's most significant performance bottleneck and automatically allocate optimization resources in that direction. This improves overall detection accuracy while focusing on tackling specific types of difficult samples, making the model fine-tuning process more efficient and targeted.

[0132] Figure 2 This is a schematic diagram of the malicious code detection process provided in this application, such as... Figure 2 As shown, the code content generated by the large language model is first preprocessed by a general content moderation engine, then subjected to BERT detection and rule base correction, and finally detected again by the general content moderation engine. The detection results are then aggregated and processed by the aggregation module, and finally output after further processing. The code sandbox is mainly used for verifying the code execution results, and after being processed by the reinforcement learning module, it is fed back to the BERT model for fine-tuning and the rule base for correction.

[0133] The preprocessing of the general content moderation engine includes: encoding verification and normalization; text format normalization, such as unifying Chinese and English punctuation; filtering invalid characters; and uniformly truncating text into fixed-length characters. A unique identifier is added to the output code of each large language model.

[0134] The apparatus provided in the embodiments of this application is described below. The apparatus described below can be referred to in correspondence with the method described above.

[0135] Figure 3 This is a schematic diagram of the malicious code detection device provided in this application, such as... Figure 3 As shown, it includes: The detection module 310 is used to input the code generated by the large language model into the first detection module and the second detection module in parallel to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base; the second detection module is a general content moderation engine. The determination module 320 is used to determine the risk level of the code based on the first risk detection result and the second risk detection result; The verification module 330 is used to input the code into the code sandbox when the code risk level is a preset risk level, and obtain the code verification result output by the code sandbox. Adjustment module 340 is used to adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results.

[0136] The malicious code detection device provided in this application has a first detection module that can deeply mine new types of obfuscated and transformed malicious code from the levels of code syntax, logic, and deep semantics, while the second detection module can quickly identify known malicious keywords or non-compliant content by leveraging its advantages in general text review. By combining dedicated deep code analysis capabilities with general content review capabilities, a parallel, multi-layered detection architecture is constructed. By setting risk levels and introducing a code sandbox to accurately verify medium-risk samples, the detection efficiency and accuracy are effectively balanced, avoiding time-consuming sandbox analysis of all code. By using the sandbox verification results to provide closed-loop feedback and dynamic adjustment to the code semantic analysis model and rule base, the entire detection system has the ability to continuously learn and self-evolve, effectively combating constantly changing malicious code variants, effectively detecting and intercepting malicious code, and significantly improving the security of the output content of the large language model.

[0137] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor, communications interface, and memory communicate with each other via the communications bus. The processor can invoke logical commands stored in the memory to execute the methods described in the above embodiments, for example: The code generated by the large language model is input in parallel into the first detection module and the second detection module to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module. The first detection module includes a code semantic analysis model and a malicious code rule base. The second detection module is a general content moderation engine. Based on the first risk detection result and the second risk detection result, the risk level of the code is determined. If the risk level of the code is the preset risk level, the code is input into the code sandbox to obtain the code verification result output by the code sandbox. Based on the code verification result, the parameters of the code semantic analysis model and / or the rules of the malicious code rule base are adjusted.

[0138] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] The processor in the electronic device provided in this application embodiment can call logical instructions in the memory to implement the above method. Its specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effect, which will not be repeated here.

[0140] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided in the above embodiments.

[0141] The specific implementation method is the same as the aforementioned method implementation method and can achieve the same beneficial effects, so it will not be repeated here.

[0142] This application provides a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for detecting malicious code, characterized in that, include: The code generated by the large language model is input in parallel into the first detection module and the second detection module to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base. The second detection module is a general content moderation engine; Based on the first risk detection result and the second risk detection result, the risk level of the code is determined; If the risk level of the code is a preset risk level, the code is input into the code sandbox, and the code verification result output by the code sandbox is obtained; Based on the code verification results, adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base.

2. The malicious code detection method according to claim 1, characterized in that, The step of adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results includes: A reward signal is determined based on the code verification results; Reinforcement learning is performed based on the reward signal, state space, and action space to adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base; The actions in the action space are used to trigger rule adjustments to the malicious code rule base and / or parameter adjustments to the code semantic analysis model. The state vector in the state space includes at least one of the following: The code verification result includes at least one of the following: malicious label, malicious behavior type, and malicious behavior confidence level. The code semantic analysis model has at least one of the following: detection accuracy, false negative rate, false positive rate, and obfuscated code recognition rate. The malicious code rule base includes at least one of the following: template hit rate, proportion of high-risk templates, coverage rate of benign templates, and template mismatch rate. Accuracy changes after the last round of optimization; The previous round of optimization took a long time.

3. The malicious code detection method according to claim 2, characterized in that, The step of determining the reward signal based on the code verification result includes: Based on the code verification results, the correctly identified code samples, the samples missed by the code semantic analysis model, and the samples misjudged by the code semantic analysis model are determined. Based on the correctly identified code samples, the detection accuracy rate is determined; Based on the samples missed by the code semantic analysis model, the false detection rate is determined. Based on the samples misjudged by the code semantic analysis model, the detection misjudgment rate is determined; Based on the total time spent adjusting the malicious code rule base and adjusting the code semantic analysis model, the processing latency cost is determined. The reward signal is determined based on the detection accuracy, the false negative rate, the false positive rate, and the processing delay cost.

4. The malicious code detection method according to claim 1, characterized in that, The step of adjusting the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results includes: Based on the code verification results, identify the malicious code samples that the code semantic analysis model missed and / or the benign code samples that it misjudged, and construct an incremental training dataset. The code semantic analysis model is fine-tuned based on the incremental training dataset.

5. The malicious code detection method according to claim 4, characterized in that, The code semantic analysis model includes a feature extraction network and a feature classification network; The fine-tuning of the code semantic analysis model based on the incremental training dataset includes: Freeze the model parameters of the feature extraction network; The model parameters of the feature classification network are adjusted based on the incremental training dataset.

6. The malicious code detection method according to claim 1, characterized in that, The fusion loss function for fine-tuning the code semantic analysis model is determined based on the following steps: The log-likelihood terms corresponding to malicious code samples and benign code samples are calculated by category weighting to determine the weighted cross-entropy loss term; The log-likelihood terms of the obfuscated code samples are summed to determine the obfuscated sample loss term; The weighted cross-entropy loss term and the confused sample loss term are weighted and fused to determine the fusion loss function.

7. The malicious code detection method according to claim 6, characterized in that, The category weights of the log-likelihood terms of the malicious code samples and the benign code samples are dynamically adjusted based on the current false negative rate of the code semantic analysis model and the probability of an increase in the false positive rate of the benign code samples after fine-tuning. The weight of the obfuscated sample loss term in the fusion loss function is dynamically adjusted based on the current obfuscated code recognition rate of the code semantic analysis model and the probability of an increase in the false positive rate of benign code samples after fine-tuning.

8. A malicious code detection device, characterized in that, include: The detection module is used to input the code generated by the large language model into the first detection module and the second detection module in parallel to obtain the first risk detection result output by the first detection module and the second risk detection result output by the second detection module; the first detection module includes a code semantic analysis model and a malicious code rule base. The second detection module is a general content moderation engine; A determination module is used to determine the risk level of the code based on the first risk detection result and the second risk detection result; The verification module is used to input the code into the code sandbox when the risk level of the code is a preset risk level, and obtain the code verification result output by the code sandbox; An adjustment module is used to adjust the parameters of the code semantic analysis model and / or the rules of the malicious code rule base based on the code verification results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the malicious code detection method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the malicious code detection method according to any one of claims 1 to 7.