Detection Method, Device, Equipment and Computer Medium Applied to Program Code
By performing data cleaning and malicious code detection models on application code, identifying and repairing unrecorded malicious code, the problem of incomplete identification in the existing technology is solved and the security of hosts and network services is improved.
Patent Information
- Application Number
- CN202411035654.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-31
AI Technical Summary
In the prior art, when identifying malicious code, some malicious code is not recorded in the feature library and the rule library, resulting in reduced security of host activities and network services.
By obtaining the initial application code information, performing data cleaning and processing, inputting a pre-trained malicious code detection model, generating a malicious detection result set, and determining the result that meets the preset malicious conditions as the target malicious detection result set, and sending it to the program code maintenance terminal for repair.
Improves the security of host activities and network services, and improves the protection capabilities of the system by identifying and repairing unrecorded malicious code.
Smart Images

Figure CN118965348B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of code detection, and more particularly to a detection method, apparatus, device, and computer medium for program code. Background Art
[0002] By identifying malicious code in the application code of host activities and issuing alerts, the security of host activities and network services can be improved. Currently, when identifying malicious code, the commonly used method is to use a pre-set feature library and rule library to identify malicious code.
[0003] However, when using the above method to identify malicious code, there are often the following technical problems: when some malicious code is not recorded in the feature library and rule library, it is difficult to identify some malicious code, resulting in a reduction in the security of host activities and network services.
[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0005] The content part of the present disclosure is used to introduce the concepts in a brief form, and these concepts will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a detection method, apparatus, electronic device, and computer-readable medium for program code to solve one or more of the technical problems mentioned in the above background art section.
[0007] In a first aspect, some embodiments of the present disclosure provide a detection method for program code, the method including: obtaining initial application program code information; performing data cleaning processing on the initial application program code information to generate application program code information; inputting the application program code information into a pre-trained malicious code detection model to generate a malicious detection result set; determining each malicious detection result that meets a preset malicious condition in the malicious detection result set as a target malicious detection result set; and sending the target malicious detection result set to a program code maintenance terminal for repair processing.
[0008] Second aspect, some embodiments of the present disclosure provide a detection device for program code, the device comprising: an acquisition unit configured to acquire initial application program code information; a cleaning unit configured to perform data cleaning processing on the initial application program code information to generate application program code information; an input unit configured to input the application program code information into a pre-trained malicious code detection model to generate a malicious detection result set; a determination unit configured to determine each malicious detection result in the malicious detection result set that meets a preset malicious condition as a target malicious detection result set; and a sending unit configured to send the target malicious detection result set to a program code maintenance terminal for repair processing.
[0009] Third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having stored thereon one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described in any implementation manner of the first aspect above.
[0010] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.
[0011] The above various embodiments of the present disclosure have the following beneficial effects: Through the detection method for program code of some embodiments of the present disclosure, the security of host activities and network services is improved. Specifically, the reason for the reduction in the security of host activities and network services is that it is difficult to identify some malicious codes when some malicious codes are not recorded in the feature library and rule library, resulting in a reduction in the security of host activities and network services. Based on this, the detection method for program code of some embodiments of the present disclosure, first, acquires initial application program code information. Second, performs data cleaning processing on the initial application program code information to generate application program code information. Thus, the program code information after data cleaning can be obtained. Then, inputs the application program code information into a pre-trained malicious code detection model to generate a malicious detection result set. Thus, the malicious detection result set of the program code information can be detected through the malicious code detection model. Then, determines each malicious detection result in the malicious detection result set that meets a preset malicious condition as a target malicious detection result set. Thus, the target malicious detection result set representing malicious codes can be determined. Finally, sends the target malicious detection result set to a program code maintenance terminal for repair processing. Thus, the malicious codes can be repaired. Furthermore, the security of the host and network services can be improved. Description of the Drawings
[0012] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flowchart of some embodiments of a detection method applied to program code according to the present disclosure;
[0014] Figure 2 is a flowchart of some embodiments of a detection device applied to program code according to the present disclosure;
[0015] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0017] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0018] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0019] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0020] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0021] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0022] Figure 1The flowchart of some embodiments of the detection method applied to program code according to the present disclosure is shown. The process 100 of some embodiments of the detection method applied to program code according to the present disclosure is shown. The detection method applied to program code includes the following steps:
[0023] Step 101, obtain the initial application program code information.
[0024] In some embodiments, the execution subject (e.g., a computing device) of the detection method applied to program code may obtain the initial application program code information. Among them, the above-mentioned initial application program code information may include: domain name information, text information. For example, the domain name information and text information included in the initial user program code information may represent: wanting to obtain text information from the domain name represented by the specified domain name information. The initial application program code information may be the information of the application program code package to be detected, and may include the application program code text, domain name information, and text information. The domain name information may refer to the domain name information carried in the application program code text. The text information may refer to the code description information and domain name description information at the domain name.
[0025] Step 102, perform data cleaning processing on the above-mentioned initial application program code information to generate application program code information.
[0026] In some embodiments, the above-mentioned execution subject may perform data cleaning processing on the above-mentioned initial application program code information to generate application program code information. For example, data cleaning may be to remove redundant symbols, blank characters or duplicate characters in the initial application program code information.
[0027] In practice, the above-mentioned execution subject may perform removal processing on the code annotation information included in the above-mentioned initial application program code information to obtain application program code information. The code annotation information may refer to the comment information of the code.
[0028] Step 103, input the above-mentioned application program code information into a pre-trained malicious code detection model to generate a malicious detection result set.
[0029] In some embodiments, the above-mentioned execution entity may input the above-mentioned application program code information into a pre-trained malicious code detection model to generate a malicious detection result set. Among them, the above-mentioned malicious code detection model may be a neural network model that takes application program code information as input and a malicious detection result set as output. The malicious detection results in the malicious detection result set may be: whether the application program code information represents malicious code. For example, when the malicious detection result is that the user program code information represents malicious code, the application program code information may represent obtaining user information (such as user name, address, etc.) without the user's consent. One malicious detection result may correspond to a certain line of malicious code in the application program code information.
[0030] Optionally, the above-mentioned malicious code detection model may be trained through the following steps:
[0031] First step, obtain a program code training sample set. Among them, the program code training samples in the above-mentioned program code training sample set include: sample application program code information and a sample malicious detection result set. Here, the sample malicious detection result set may be the sample label corresponding to the sample application program code information.
[0032] Second step, determine an initial malicious code detection model. Among them, the above-mentioned initial malicious code detection model includes: an initial classification network, a first initial feature extraction network, a second initial feature extraction network, an initial fusion network, and an initial recognition network. The above-mentioned first initial feature extraction network includes: an initial division network and an initial domain name feature extraction network. Here, the above-mentioned initial classification network may be a network that takes sample application program code information as input and an initial classification information set as output. Among them, the initial classification information in the initial classification information set may represent the domain name information included in the sample application program code information. The initial classification information in the initial classification information set may also represent the text information included in the sample application program code information. For example, the initial classification model is used to: divide the domain name information and text information included in the sample application program code information. The above-mentioned first initial feature extraction network may be a neural network model that takes the initial classification information as input and the first initial feature information as output. Among them, the above-mentioned first initial feature extraction network may include: an initial division network and an initial domain name feature extraction network. Here, the initial division network may be a classification model that takes the initial classification information as input and outputs the first initial domain name division information, the second initial domain name division information, and the third initial domain name division information. For example, the first initial domain name division information may be: "Http:", "Https:". The second initial domain name division information may be, but is not limited to: "www.serve.com", "www.login.com". The third initial domain name division information may be, but is not limited to: ".com", ".net".
[0033] In the third step, select a target program code training sample from the above program code training sample set. One program code training sample can be randomly selected from the above program code training sample set as the target program code training sample.
[0034] In the fourth step, input the sample application program code information included in the target program code training sample into the initial classification network to obtain an initial classification information set.
[0035] In the fifth step, input the initial classification information that meets the preset classification conditions in the initial classification information set into the first initial feature extraction network to obtain the first initial feature information.
[0036] Among them, the above fifth step may include the following sub-steps:
[0037] In the first sub-step, input the initial classification information that meets the preset classification conditions in the initial classification information set into the initial partitioning network to obtain the first initial domain name partitioning information, the second initial domain name partitioning information, and the third initial domain name partitioning information.
[0038] In the second sub-step, in response to determining that the initial classification information includes preset information, determine the initial classification information as the initial malicious domain name information. Among them, the preset information can be an Internet protocol address. For example, the preset information can be "http: / / www.microsoft.com.cn".
[0039] In the third sub-step, in response to determining that the second initial domain name partitioning information meets the first preset domain name condition, determine the initial classification information as the initial malicious domain name information. Among them, the first preset domain name condition can be: the second initial domain name partitioning information includes preset sensitive information. For example, the preset sensitive information can be, but is not limited to, "login", "account".
[0040] In the fourth sub-step, in response to determining that the third initial domain name partitioning information meets the second preset domain name condition, determine the initial classification information as the initial malicious domain name information. Among them, the second preset domain name condition can be: the third initial domain name partitioning information is not the preset domain name suffix information. For example, the preset domain name suffix information can be, but is not limited to, ".com", ".cn", ".net", ".org".
[0041] In the fifth sub-step, input the initial malicious domain name information into the initial domain name feature extraction network to obtain the first initial feature information. Among them, the initial domain name feature extraction network is a custom network that takes the initial malicious domain name information as the input and the first initial feature information as the output.
[0042] Among them, the above-mentioned custom network is divided into three layers: The first layer is the embedding layer, which is used for: vectorizing the initial malicious domain name information to generate initial malicious domain name vectorized information; The second layer is the convolutional layer, including: the first convolutional network, the second convolutional network, and the third convolutional network. Among them, the first convolutional network is a convolutional network that takes the initial malicious domain name vectorized information as input and outputs the first initial malicious domain name convolutional information. The second convolutional network is a convolutional network that takes the first initial malicious domain name convolutional information as input and outputs the second initial malicious domain name convolutional information. The third convolutional network is a convolutional network that takes the second initial malicious domain name convolutional information as input and outputs the third initial malicious domain name convolutional information; The third layer is the fully connected layer, which is used for: converting the third initial malicious domain name convolutional information to generate the first initial feature information.
[0043] The sixth step is to input the initial classification information in the initial classification information set that does not meet the preset classification conditions into the second initial feature extraction network to obtain the second initial feature information. The above-mentioned second initial feature extraction network can be a neural network model that takes the initial classification information as input and outputs the second initial feature information. Among them, the preset classification condition can be that the initial classification information represents domain name information.
[0044] The seventh step is to input the first initial feature information and the second initial feature information into the initial fusion network to obtain the initial feature fusion information. The above-mentioned initial fusion network can be a model that takes the first initial feature information and the second initial feature information as input and outputs the initial feature fusion information. For example, the initial fusion network is used for: determining the first initial feature information and the second initial feature information as the initial feature fusion information.
[0045] The eighth step is to input the initial feature fusion information into the initial recognition network to obtain the initial malicious detection result set. The above-mentioned initial recognition network can be a neural network model that takes the initial feature fusion information as input and outputs the initial malicious detection result. For example, the initial recognition network can be a support vector machine (SVM) model.
[0046] The ninth step is to determine the loss value between the initial malicious detection result set and the corresponding sample malicious detection result set.
[0047] Among them, the above-mentioned ninth step may include the following sub-steps:
[0048] The first sub-step is to determine the operator loss value between the initial malicious detection result set and the corresponding sample malicious detection result set through a preset operator loss function.
[0049] The second sub-step is to determine the error loss value between the initial malicious detection result set and the corresponding sample malicious detection result set through a preset error loss function.
[0050] The third sub-step is to determine the loss value between the initial malicious detection result set and the corresponding sample malicious detection result set according to the above operator loss value and the above error loss value.
[0051] The tenth step is to determine the initially trained malicious code detection model as the trained malicious code detection model in response to determining that the loss value is less than the preset difference value.
[0052] Thus, the feature information in the initial malicious domain name information can be relatively accurately extracted by training the initial domain name feature extraction model including the embedding layer, convolutional layer, and fully connected layer. Therefore, relatively accurate malicious code can be identified through the relatively accurate first initial feature information extracted.
[0053] Optionally, in response to determining that the loss value is greater than or equal to the preset difference value, adjust the network parameters of the initial malicious code detection model, determine the adjusted initial malicious code detection model as the initial malicious code detection model, and select a target program code training sample from each of the unselected program code training samples, and train the initial malicious code detection model again. For example, the difference between the above loss value and the preset difference value can be calculated. On this basis, methods such as backpropagation and gradient descent are used to adjust the network parameters of the above initial malicious code detection model. It should be noted that the backpropagation algorithm and the gradient descent method are well-known technologies that have been widely studied and applied at present, and will not be elaborated here. Among them, there is no limitation on the setting of the preset difference value. For example, the preset difference value can be 0.1.
[0054] Step 104: Determine each malicious detection result that meets the preset malicious condition in the above malicious detection result set as the target malicious detection result set.
[0055] In some embodiments, the above execution subject may determine each malicious detection result that meets the preset malicious condition in the above malicious detection result set as the target malicious detection result set. Among them, the preset malicious condition may be: the malicious detection result represents malicious code for the application program code information.
[0056] Step 105: Send the above target malicious detection result set to the program code maintenance terminal for repair processing.
[0057] In some embodiments, the above execution subject may send the above target malicious detection result set to the program code maintenance terminal for repair processing. The program code maintenance terminal may refer to a terminal for maintaining and repairing code.
[0058] For further reference Figure 2, as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a detection device applied to program code. These embodiments of the detection device applied to program code correspond to Figure 1 the method embodiments shown, and the detection device applied to program code can be specifically applied to various electronic devices.
[0059] As Figure 2 shown, the detection device 200 applied to program code in some embodiments includes: an acquisition unit 201, a cleaning unit 202, an input unit 203, a determination unit 204, and a sending unit 205. Among them, the acquisition unit 201 is configured to acquire initial application program code information; the cleaning unit 202 is configured to perform data cleaning processing on the above initial application program code information to generate application program code information; the input unit 203 is configured to input the above application program code information into a pre-trained malicious code detection model to generate a malicious detection result set; the determination unit 204 is configured to determine each malicious detection result that meets the preset malicious condition in the above malicious detection result set as a target malicious detection result set; the sending unit 205 is configured to send the above target malicious detection result set to a program code maintenance terminal for repair processing.
[0060] It can be understood that the units described in the detection device 200 applied to program code correspond to Figure 1 each step in the method described in the reference. Thus, the operations, features, and beneficial effects described above for the method also apply to the detection device 200 applied to program code and the units included therein, and will not be elaborated herein.
[0061] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device (such as a computing device) 300 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0062] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0063] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in may represent one device or, as needed, multiple devices.
[0064] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0065] It should be noted that the computer-readable media described in some embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0066] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0067] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain initial application program code information; perform data cleaning processing on the above initial application program code information to generate application program code information; input the above application program code information into a pre-trained malicious code detection model to generate a malicious detection result set; determine each malicious detection result in the above malicious detection result set that meets a preset malicious condition as a target malicious detection result set; and send the above target malicious detection result set to a program code maintenance terminal for repair processing.
[0068] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0069] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0070] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes: an acquisition unit, a cleaning unit, an input unit, a determination unit, and a sending unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring initial application code information".
[0071] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0072] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A detection method applied to program code, comprising: Obtaining initial application program code information; Obtaining a program code training sample set; Determining an initial malicious code detection model, wherein the initial malicious code detection model includes: an initial classification network, a first initial feature extraction network, a second initial feature extraction network, an initial fusion network, and an initial recognition network; Selecting a target program code training sample from the program code training sample set; Inputting the sample application program code information included in the target program code training sample into the initial classification network to obtain an initial classification information set, wherein the initial classification information in the initial classification information set represents the domain name information included in the sample application program code information; Inputting the initial classification information in the initial classification information set that meets the preset classification condition into the first initial feature extraction network to obtain first initial feature information, including: Inputting the initial classification information in the initial classification information set that meets the preset classification condition into the initial classification network to obtain first initial domain name division information, second initial domain name division information, and third initial domain name division information; In response to determining that the second initial domain name division information meets the first preset domain name condition, determining the initial classification information as initial malicious domain name information, wherein the first preset domain name condition is that the second initial domain name division information includes preset sensitive information; In response to determining that the third initial domain name division information meets the second preset domain name condition, determining the initial classification information as initial malicious domain name information, wherein the second preset domain name condition is that the third initial domain name division information is not the preset domain name suffix information; Inputting the initial malicious domain name information into the initial domain name feature extraction network to obtain first initial feature information, wherein the initial domain name feature extraction network is a custom network that takes the initial malicious domain name information as input and the first initial feature information as output, and the custom network is divided into three layers: the first layer, the embedding layer, is used for: vectorizing the initial malicious domain name information to generate initial malicious domain name vectorized information; the second layer, the convolutional layer, includes: a first convolutional network, a second convolutional network, and a third convolutional network, wherein the first convolutional network is a convolutional network that takes the initial malicious domain name vectorized information as input and the first initial malicious domain name convolutional information as output, the second convolutional network is a convolutional network that takes the first initial malicious domain name convolutional information as input and the second initial malicious domain name convolutional information as output, and the third convolutional network is a convolutional network that takes the second initial malicious domain name convolutional information as input and the third initial malicious domain name convolutional information as output; the third layer, the fully connected layer, is used for: performing conversion processing on the third initial malicious domain name convolutional information to generate first initial feature information; Inputting the initial classification information in the initial classification information set that does not meet the preset classification condition into the second initial feature extraction network to obtain second initial feature information; Inputting the first initial feature information and the second initial feature information into the initial fusion network to obtain initial feature fusion information; Inputting the initial feature fusion information into the initial recognition network to obtain an initial malicious detection result set.
2. The method according to claim 1, wherein, The method further includes: Perform data cleaning processing on the initial application program code information to generate application program code information; Determine the loss value between the initial malicious detection result set and the corresponding sample malicious detection result set; In response to determining that the loss value is less than the preset difference value, determine the initial malicious code detection model as the trained malicious code detection model; Input the application program code information into the pre-trained malicious code detection model to generate a malicious detection result set; Determine each malicious detection result that meets the preset malicious condition in the malicious detection result set as the target malicious detection result set; Send the target malicious detection result set to the program code maintenance terminal for repair processing.
3. The method according to claim 2, wherein The performing data cleaning processing on the initial application program code information to generate application program code information includes: Perform removal processing on the code annotation information included in the initial application program code information to obtain application program code information.
4. The method according to claim 2, wherein The method further includes: In response to determining that the loss value is greater than or equal to the preset difference value, adjust the network parameters of the initial malicious code detection model, determine the adjusted initial malicious code detection model as the initial malicious code detection model, and select target program code training samples from each unselected program code training sample, and train the initial malicious code detection model again.
5. The method according to claim 2, wherein The determining the loss value between the initial malicious detection result set and the corresponding sample malicious detection result set includes: Determine the operator loss value between the initial malicious detection result set and the corresponding sample malicious detection result set through a preset operator loss function; Determine the error loss value between the initial malicious detection result set and the corresponding sample malicious detection result set through a preset error loss function; Determine the loss value between the initial malicious detection result set and the corresponding sample malicious detection result set according to the operator loss value and the error loss value.
6. A detection device applied to program code, including: An acquisition unit configured to acquire initial application program code information; A training unit configured to acquire a program code training sample set; Determine an initial malicious code detection model, where the initial malicious code detection model includes: an initial classification network, a first initial feature extraction network, a second initial feature extraction network, an initial fusion network, and an initial recognition network; select a target program code training sample from the program code training sample set; input the sample application program code information included in the target program code training sample into the initial classification network to obtain an initial classification information set, where the initial classification information in the initial classification information set represents the domain name information included in the sample application program code information; input the initial classification information that meets the preset classification conditions in the initial classification information set into the first initial feature extraction network to obtain first initial feature information, including: input the initial classification information that meets the preset classification conditions in the initial classification information set into the initial classification network to obtain a first initial domain name division information, a second initial domain name division information, and a third initial domain name division information; in response to determining that the second initial domain name division information meets the first preset domain name condition, determine the initial classification information as initial malicious domain name information, where the first preset domain name condition is: the second initial domain name division information includes preset sensitive information; in response to determining that the third initial domain name division information meets the second preset domain name condition, determine the initial classification information as initial malicious domain name information, where the second preset domain name condition is: the third initial domain name division information is not the preset domain name suffix information; input the initial malicious domain name information into the initial domain name feature extraction network to obtain first initial feature information, where the initial domain name feature extraction network is a custom network that takes the initial malicious domain name information as input and the first initial feature information as output, and where the custom network is divided into three layers: the first layer, the embedding layer, is used to: vectorize the initial malicious domain name information to generate initial malicious domain name vectorized information; the second layer, the convolutional layer, includes: a first convolutional network, a second convolutional network, and a third convolutional network, where the first convolutional network is a convolutional network that takes the initial malicious domain name vectorized information as input and the first initial malicious domain name convolutional information as output, the second convolutional network is a convolutional network that takes the first initial malicious domain name convolutional information as input and the second initial malicious domain name convolutional information as output, and the third convolutional network is a convolutional network that takes the second initial malicious domain name convolutional information as input and the third initial malicious domain name convolutional information as output; the third layer, the fully connected layer, is used to: perform a conversion process on the third initial malicious domain name convolutional information to generate the first initial feature information; input the initial classification information that does not meet the preset classification conditions in the initial classification information set into the second initial feature extraction network to obtain second initial feature information; input the first initial feature information and the second initial feature information into the initial fusion network to obtain initial feature fusion information; input the initial feature fusion information into the initial recognition network to obtain an initial malicious detection result set.
7. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs; When the one or more programs are executed by the one or more processors such that the one or more processors implement the method according to any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by a processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Code detection method and device, computer equipment and storage medium
CN114817913A
Industrial internet data security detection response and traceability method and device
CN117454376A
Code auditing method and device, terminal equipment and storage medium
CN118153048A