Methods, systems, apparatus, and computer readable media for training neural networks to learn computer code change representations

By performing function slicing and comparative learning of computer code, combined with a general defect list, unsupervised twin neural network is trained, and the problem of limited data of silent repair detection in open source software is solved, and efficient vulnerability detection and rating is achieved.

CN120380429APending Publication Date: 2025-07-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102661.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In open source software, the time for silently repairing vulnerabilities and the time gap for publicly disclosed vulnerabilities provide an opportunity for attack, and the existing technology is difficult to effectively detect and interpret silent repairs, and the data for training neural networks is limited, resulting in low reliability of the results.

Method used

By dividing computer code into multiple parts, generating function slices and using contrast learning to enhance data, training unsupervised twin neural networks, combining general defect lists and security announcement services, learning function change representations, realizing a training method that combines unsupervised and supervised.

Benefits of technology

Improves detection accuracy and reliability of silent repairs, can provide vulnerability categories and exploitability ratings, reduces the dependence of training data, and enhances the reliability of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380429A_ABST
    Figure CN120380429A_ABST
Patent Text Reader

Abstract

A method for training a neural network and a computer readable medium are described. A piece of computer code is divided into a plurality of computer code portions. A first alteration sample is generated, the first alteration sample comprising a first original computer code segment and a first modified computer code segment, the first alteration sample comprising at least one computer code portion of the plurality of computer code portions. A second alteration sample is generated, the second alteration sample including a second original computer code segment and a second modified computer code segment. A loss function is calculated based on the first altered sample and the second altered sample. The neural network is trained by minimizing the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to methods, systems, devices, and computer-readable storage media for training neural networks, and more particularly to methods, systems, devices, and computer-readable storage media for training neural networks to detect and characterize vulnerability fixes in computer code changes. Background Art

[0002] In a software project, there is usually a delay between the time when a developer of a software product fixes a vulnerability in the software product and the time when the vulnerability is publicly disclosed or otherwise becomes known to the public. This time gap provides an opportunity window for the vulnerability to be exploited. Since open-source software commits are public, malicious parties may discover vulnerabilities based on public software commits of fixes before the vulnerabilities are publicly disclosed to the public. Therefore, there is a need to determine whether a computer code change fixes a vulnerability. Summary of the Invention

[0003] Generally, according to some embodiments of the present invention, methods for training a neural network to detect vulnerability fixes in computer code changes are described. The response disclosure model for an open-source software project can involve the following three steps: (1) the vulnerability is secretly fixed without mention of the vulnerability; (2) the vulnerability is publicly disclosed through an announcement; (3) users of the software update the software in response to the vulnerability announcement. For users of a software system, it is crucial to be aware of the vulnerability and update their systems in a timely manner. In the context of open-source software, a vulnerability can be fixed in step (1) through a source code commit to a public source code repository as a silent fix. A silent fix is a commit used to fix a vulnerability that does not include any information about the vulnerability. Nevertheless, it is still possible for malicious users to reverse engineer the vulnerability based on changes to the computer code of the vulnerability fixed in step 1. Therefore, malicious users can exploit the vulnerability to attack software users who have not yet updated their software. There may be a time gap between step (1) when the vulnerability is fixed and step (2) when the vulnerability is publicly disclosed through an announcement. Therefore, for users of open-source software, it is important to detect a silent fix before it is publicly announced. Therefore, there is a need for a neural network that can take patch data as input and determine whether the patch is used to fix a vulnerability.

[0004] One of the problems in creating such a neural network is the lack of sufficient data to train the neural network. In some embodiments, computer code data can be augmented to increase the size of the training data. For example, for each function code change, the original computer code can be divided into multiple OriFSlice, and the modified computer code can be divided into multiple ModFSlice. Control flow graphs and data flow graphs can be used to generate slices based on the changed variables as anchors. OriFSlice and ModFSlice can be combined together into multiple function change samples. Automatically generated descriptions can also be included in the samples. Since OriFSlice and ModFSlice come from the same changed function, they may have the same semantic meaning. That is, they fix the same vulnerability. Therefore, these samples can be used to train the neural network to detect vulnerabilities through contrastive learning in an unsupervised manner because the user does not need to label function changes. A common weakness enumeration (CWE) can be used to assist in training. CWE provides a dictionary of common vulnerabilities, and the dictionary of common vulnerabilities can be used to classify vulnerabilities. Since a single patch may result in multiple samples, the available data has been augmented. The neural network can be trained by minimizing the differences between samples from the same function or the same CWE category. The neural network can also be trained by maximizing the distances between samples from different CWE categories.

[0005] According to a first aspect of the present invention, a method for training a neural network is described, including: dividing a section of computer code into multiple computer code parts; generating a first change sample; generating a second change sample; calculating a loss function based on the first change sample and the second change sample; and training the neural network by minimizing the loss function.

[0006] In a possible implementation, the first change sample includes a first original computer code segment and a first modified computer code segment.

[0007] Optionally, the first original segment and the first modified segment correspond to the same function.

[0008] In another possible implementation, the multiple computer code parts include multiple original computer code parts and multiple modified computer code parts, wherein the first original computer code segment includes the first original computer code part among the multiple original computer code parts, and wherein the first modified computer code segment includes the first modified computer code part among the multiple modified computer code parts.

[0009] In another possible implementation, the second change sample includes a second original computer code segment and a second modified computer code segment.

[0010] Optionally, the second original computer code segment includes a second original computer code part among a plurality of original computer code parts, and the second modified computer code segment includes a second modified computer code part among a plurality of modified computer code parts.

[0011] In another possible implementation, the first change sample and the second change sample correspond to the same function.

[0012] Optionally, the first change sample and the second change sample belong to the same category.

[0013] In another possible implementation, both the first change sample and the second change sample fix vulnerabilities of the same category.

[0014] In another possible implementation, the first change sample further includes: an automatically generated description, or a manually marked description, or a combination of an automatically generated description and a manually marked description.

[0015] Optionally, the computer code segment is a function.

[0016] In another possible implementation, the function is divided into multiple computer code parts based on the changed variables using a control flow graph or a data flow graph.

[0017] In another possible implementation, it further includes:

[0018] Generating a third change sample; calculating a loss function based on the first change sample and the third change sample; training a neural network by maximizing the loss function.

[0019] In another possible implementation, the computer code segment is obtained from a security bulletin service or a common vulnerability disclosure database.

[0020] In another possible implementation, the neural network is trained in an unsupervised manner.

[0021] In another possible implementation, the neural network is trained using contrastive learning, or the neural network is a siamese neural network.

[0022] Optionally, it further includes fine-tuning the neural network for a task.

[0023] In another possible implementation, the computer code is source code, intermediate code, or machine code.

[0024] According to another aspect of the present invention, there is provided a non-transitory computer-readable medium including computer program code stored therein for training a neural network, wherein when the code is executed by one or more processors, the one or more processors are caused to perform a method including: dividing a piece of computer code into multiple computer code parts; generating a first change sample including a first original computer code segment and a first modified computer code segment, the first change sample including at least one computer code part of the multiple computer code parts; generating a second change sample including a second original computer code segment and a second modified computer code segment; calculating a loss function based on the first change sample and the second change sample; and training the neural network by minimizing the loss function.

[0025] The method may further include performing any of the operations described above in connection with the first aspect of the present invention.

[0026] The neural network can be used to calculate the probability of a computer code change fixing a vulnerability. The neural network can be used to calculate the probability that a computer code change belongs to a certain category. The neural network can be used to assign a rating to a vulnerability. The rating can be an exploitability rating or a severity rating.

[0027] The present invention content does not necessarily describe all aspects of the full scope. Other aspects, features, and advantages will be apparent to those of ordinary skill in the art after reviewing the following description of the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more fully understand the present invention, reference is made to the following description and drawings, in which:

[0029] Figure 1 Schematic diagram of a computer network system for training a neural network to learn computer code change representations according to some embodiments of the present invention;

[0030] Figure 2A For Figure 1 Schematic diagram of the simplified hardware structure of a computing device of the computer network system shown in;

[0031] Figure 2B For Figure 1 Schematic diagram of the simplified software architecture of a computing device of the computer network system shown in;

[0032] Figure 3 Shows an exemplary timeline in response to a disclosure;

[0033] Figure 4 Shows an example of a submission;

[0034] Figure 5AFlowchart of three stages of a method for training a neural network to learn representations of computer code changes according to some embodiments of the present invention;

[0035] Figure 5B is Figure 5A Flowchart of some details of three stages of the method shown in;

[0036] Figure 6 Flowchart of a method for training a neural network to learn representations of computer code changes according to some embodiments of the present invention;

[0037] Figure 7 Schematic diagram of a method for enhancing computer code change data according to some embodiments of the present invention;

[0038] Figure 8 Schematic diagram of a method for training a neural network to learn representations of computer code changes according to some embodiments of the present invention;

[0039] Figure 9 Schematic diagram of source code changes;

[0040] Figure 10 Example of a function change description FCDesc showing a patch for fixing a cross - site scripting vulnerability in Apache ActiveMQ;

[0041] Figure 11 is Figure 5A Flowchart of the workflow for fine - tuning downstream tasks in stage 3 of the method shown in. Detailed Description

[0042] The embodiments disclosed herein relate to a neural network module or circuit for performing a neural network training process, and more particularly, to a neural network training process for detecting and characterizing vulnerability fixes in computer code changes. Herein, a vulnerability fix is a commit for fixing a vulnerability in a software product (e.g., a vulnerability in an open - source software product).

[0043] As will be described in more detail later, a "module" is an interpretive term that refers to a hardware structure, such as a circuit implemented using techniques such as electrical technology and / or optical technology (and more specific examples of semiconductors) for performing the defined operations or processing. Alternatively, a "module" can refer to a combination of a hardware structure and a software structure, where the hardware structure can be generally implemented using techniques such as electrical technology and / or optical technology (and more specific examples of semiconductors) to perform the defined operations or processing according to the software structure in the form of a set of instructions stored in one or more non - transitory computer - readable storage devices or media.

[0044] As will be described in more detail below, the neural network module can be part of a device, apparatus, system, etc., where the neural network module can be coupled to or integrated with other parts of the device, apparatus, or system such that their combination forms the device, apparatus, or system. Alternatively, the neural network module can be implemented as an independent neural network device or apparatus.

[0045] The neural network module performs a neural network training process for training a neural network to learn representations of computer code changes. Here, a process has the general meaning equivalent to a method and does not necessarily correspond to the concept of a computational process (which is an instance of a computer program being executed). More specifically, the processes herein are defined methods implemented using hardware components for processing data (such as computer code changes, source code changes, intermediate code changes, or machine code changes, etc.). A process can include or use one or more designed functions to process data. Here, a function is a defined sub - process or sub - method for computing (computing / calculating) or otherwise processing input data in a defined manner and generating or otherwise producing output data.

[0046] As will be understood by those skilled in the art, the neural network training processes disclosed herein can be implemented as one or more software and / or firmware programs having the necessary computer - executable code or instructions and stored in one or more non - transitory computer - readable storage devices or media, which can be any volatile and / or non - volatile, non - removable or removable storage devices such as RAM, ROM, EEPROM, solid - state memory devices, hard disks, CDs, DVDs, flash devices, etc. The neural network module can read the computer - executable code from the storage device and execute the computer - executable code to perform the neural network training process.

[0047] Alternatively, the neural network training processes disclosed herein can be implemented as one or more hardware structures having the necessary electrical and / or optical components, circuits, logic gates, integrated circuit (IC) chips, etc.

[0048] A. System Structure

[0049] Please refer to Figure 1 , Figure 1 which shows a computer network system for training a neural network and is generally identified by the reference numeral 100. In these embodiments, the neural network system 100 is used to train a neural network.

[0050] As Figure 1As shown, the neural network system 100 includes one or more server computers 102, multiple client computing devices 104, and one or more client computer systems 106, which are functionally interconnected through a network 108 such as the Internet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), etc., via appropriate wired and wireless network connections.

[0051] The server computer 102 can be a computing device specifically designed to serve as a server and / or a general-purpose computing device that acts as a server computer and is also used by various users. Each server computer 102 can execute one or more server programs.

[0052] The client computing device 104 can be a portable and / or non-portable computing device, such as a laptop computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a desktop computer, etc. Each client computing device 104 can execute one or more client applications, and sometimes one or more client applications can be referred to as "applications".

[0053] Generally, the computing devices 102 and 104 include similar hardware structures, such as Figure 2A the hardware structure 120 shown in. As shown, the hardware structure 120 includes a processing structure 122, a control structure 124, one or more non-transitory computer-readable memories or storage devices 126, a network interface 128, an input interface 130, and an output interface 132, which are functionally interconnected through a system bus 138. The hardware structure 120 may also include other components 134 coupled to the system bus 138.

[0054] The processing structure 122 can be one or more single-core or multi-core computing processors, generally referred to as a central processing unit (CPU), such as a microprocessor (Intel is a registered trademark of Intel Corporation, Santa Clara, California, USA), a microprocessor (AMD is a registered trademark of Advanced Micro Devices, Inc., Sunnyvale, California, USA), a microprocessor (ARM is a registered trademark of ARM Limited, Cambridge, UK) (manufactured by various manufacturers such as Qualcomm Incorporated, San Diego, California, USA) in manufactured under the architecture), etc. When the processing structure 122 includes multiple processors, the processors can cooperate through dedicated circuits such as dedicated buses or through the system bus 138.

[0055] The processing structure 122 may also include one or more real-time processors, programmable logic controllers (PLCs), microcontroller units (MCUs), μ-controllers (UCs), dedicated / custom processors, hardware accelerators, and / or control circuits (also referred to as "controllers"), for example, using field-programmable gate array (FPGA) or application-specific integrated circuit (ASIC) technology, etc. In some embodiments, the processing structure includes a CPU (otherwise referred to as the host processor) and a dedicated hardware accelerator, and the dedicated hardware accelerator includes circuits for performing calculations of neural networks such as tensor multiplication, matrix multiplication, etc. The host processor can offload some calculations to the hardware accelerator to perform the calculation operations of the neural network. Examples of hardware accelerators include graphics processing units (GPUs), neural processing units (NPUs), and tensor processing units (TPUs). In some embodiments, the host processor and the hardware accelerator (such as GPUs, NPUs, and / or TPUs) can generally be considered processors.

[0056] Generally, the processing structure 122 includes the necessary circuits implemented using technologies such as electrical and / or optical hardware components for performing encryption processes and / or decryption processes, and as a design purpose and / or use case, can be used to encrypt and / or decrypt data received from the input end and output the resulting encrypted or decrypted data through the output end.

[0057] For example, the processing structure 122 may include logic gates implemented by semiconductors to perform various computations, calculations, and / or processing. Examples of logic gates include AND gates, OR gates, XOR (exclusive OR) gates, and NOT gates, each of which requires one or more inputs and generates or otherwise produces an output based on the logic implemented therein. For example, NOT receives an input (such as a high voltage, a state with current, a state with light emission, etc.), inverts the input (such as forms a low voltage, a state without current, a state without light, etc.), and outputs the inverted input as the output.

[0058] Although the inputs and outputs of logic gates are typically physical signals, and their logical or processing operations are tangible operations with physical results (e.g., the output of a physical signal), their inputs and outputs are typically described using numbers (e.g., the digits "0" and "1"), and their operations are typically described as "computing" (from which "computer" or "computing device" is named) or "calculation", or more generally, as "processing", for generating or producing an output from their inputs.

[0059] Complex combinations of logic gates in the form of logic gate circuits can be formed using multiple AND, OR, XOR, and / or NOT gates, such as processing structure 122. These combinations of logic gates can be implemented using individual semiconductors, or more commonly, as an integrated circuit (IC).

[0060] The circuit of a logic gate can be a "hardwired" circuit that, once designed, can only perform the designed function. In this example, its process and function are "hardcoded" in the circuit.

[0061] With technological advancements, the circuit of a logic gate (such as processing structure 122) can typically alternatively be designed in a general-purpose manner such that it can perform various processes and functions according to a set of "programming" instructions that are implemented as firmware and / or software and stored in one or more non-transitory computer-readable storage devices or media. In this example, the circuit of a logic gate (such as processing structure 122) is typically useless without meaningful firmware and / or software.

[0062] Of course, those skilled in the art will understand that a process or function (and thus the processor 102) can be implemented using other technologies such as analog technologies.

[0063] Return reference Figure 1 The control structure 124 includes one or more control circuits (such as a graphics controller, an input / output chipset, etc.) for coordinating the operations of the various hardware components and modules of the computing devices 102 / 104.

[0064] The memory 126 includes one or more storage devices or media accessible by the processing structure 122 and the control structure 124 for reading and / or storing instructions for execution by the processing structure 122, and for reading and / or storing data, including input data and data generated by the processing structure 122 and the control structure 124. The memory 126 can be volatile and / or non-volatile, non-removable or removable memory, such as RAM, ROM, EEPROM, solid-state memory, hard disk, CD, DVD, flash memory, etc.

[0065] The network interface 128 includes one or more network modules for connecting to other computing devices or networks via a network 108 by using suitable wired or wireless communication technologies, such as Ethernet, (WI-FI is a registered trademark of the Wi-Fi Alliance in Austin, Texas, USA), (Bluetooth is a registered trademark of Bluetooth Sig Inc. in Kirkland, Washington, USA), Bluetooth Low Energy (BLE), Z-Wave, Long Range (LoRa), (ZIGBEE is a registered trademark of ZigBee Alliance Corp. in San Ramon, California, USA), wireless broadband communication technologies (such as Global System for Mobile (GSM)), Code Division Multiple Access (CDMA), Universal Mobile Telecommunications System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), CDMA2000, Long Term Evolution (LTE), 3GPP, 5G New Radio (5G NR), and / or other 5G networks, etc. In some embodiments, parallel ports, serial ports, USB connections, optical connections, etc. can also be used to connect to other computing devices or networks, although they are generally considered input / output interfaces for connecting input / output devices.

[0066] The input interface 130 includes one or more input modules for one or more users to input data through, for example, a touch-sensitive screen, a touch-sensitive whiteboard, a touchpad, a keyboard, a computer mouse, a trackball, a microphone, a scanner, a camera, etc. The input interface 130 can be a physically integrated part of the computing device 102 / 104 (for example, the touchpad of a laptop or the touch-sensitive screen of a tablet), or a device that is physically separated from other components of the computing device 102 / 104 (for example, a computer mouse) but functionally coupled. In some implementations, the input interface 130 can be integrated with the display output to form a touch-sensitive screen or a touch-sensitive whiteboard.

[0067] The output interface 132 includes one or more output modules for outputting data to a user. Examples of output modules include a display (such as a monitor, LCD display, LED display, projector, etc.), a speaker, a printer, a virtual reality (VR) headset, an augmented reality (AR) goggles, etc. The output interface 132 can be a physically integrated part of the computing device 102 / 104 (such as the display of a laptop computer or a tablet), or a device that is physically separated from other components of the computing device 102 / 104 (such as the display of a desktop computer) but functionally coupled.

[0068] The computing device 102 / 104 may also include other components 134, such as one or more positioning modules, a temperature sensor, a barometer, an inertial measurement unit (IMU), etc.

[0069] The system bus 138 interconnects various components 122 to 134, enabling them to send and receive data and control signals to each other.

[0070] Figure 2B A simplified software architecture 160 of the computing device 102 or 104 is shown. The software architecture 160 includes one or more application programs 164, an operating system 166, a logical input / output (I / O) interface 168, and a logical memory 172. One or more application programs 164, the operating system 166, and the logical I / O interface 168 are typically implemented as computer-executable instructions or code in the form of software programs or firmware programs stored in the logical memory 172, and the computer-executable instructions or code can be executed by the processing structure 122.

[0071] One or more application programs 164 are executed or run by the processing structure 122 to perform various tasks.

[0072] The operating system 166 manages various hardware components of the computing device 102 or 104 through the logical I / O interface 168, manages the logical memory 172, and manages and supports the application programs 164. The operating system 166 also communicates with other computing devices (not shown) through the network 108 to support communication between the application programs 164 and the application programs running on other computing devices. As those skilled in the art will recognize, the operating system 166 can be any suitable operating system, for example, (MICROSOFT and WINDOWS are registered trademarks of Microsoft Corporation in Redmond, Washington, USA) OS X iOS (APPLE is a registered trademark of Apple Inc., located in Cupertino, California, USA), Linux, (ANDROID is a registered trademark of Google LLC, located in Mountain View, California, USA), etc. The computing devices 102 and 104 of the mirror disinfection system 100 may have the same operating system or different operating systems.

[0073] The logical I / O interface 168 includes one or more device drivers 170 for communicating with the corresponding input and output interfaces 130 and 132 to receive data from them and send data to them. The received data can be sent to one or more applications 164 for processing by the one or more applications 164. The data generated by the applications 164 can be sent to the logical I / O interface 168 to be output (via the output interface 132) to various output devices.

[0074] The logical memory 172 is a logical mapping of the physical memory 126 to facilitate access by the applications 164. In this embodiment, the logical memory 172 includes a storage memory area that can be mapped to non-volatile physical memory such as a hard disk, a solid-state drive, a flash drive, etc., which is typically used for long-term data storage therein. The logical memory 172 also includes a working memory area that is typically mapped to high-speed physical memory and, in some implementations, is mapped to volatile physical memory such as RAM, and is typically used for the applications 164 to temporarily store data during program execution. For example, the applications 164 can load data from the storage memory area into the working memory area and can store the data generated during their execution into the working memory area. The applications 164 can also store some data into the storage memory area as needed or in response to user commands.

[0075] In the server computer 102, one or more applications 164 generally provide server functions to manage network communication with the client computing device 104 and to facilitate cooperation between the server computer 102 and the client computing device 104. Herein, the term "server", from a hardware perspective, can refer to the server computer 102, or from a software perspective, can refer to a logical server, depending on the context.

[0076] As described above, processing structure 122 is generally useless without meaningful firmware and / or software. Similarly, although a computer system such as neural network system 100 may have the potential to perform various tasks, it cannot perform any tasks and is useless without meaningful firmware and / or software. As will be described in more detail later, the neural network system 100 and its modules, circuits, and components (as a combination of hardware and software) described herein generally produce tangible results associated with the physical world, where those tangible results described herein can achieve improvements to the computer devices and systems themselves, modules, circuits, and their components, etc.

[0077] B. Responsive Disclosure Model

[0078] Responsive disclosure (also known as "coordinated vulnerability disclosure") is a vulnerability disclosure model in which a vulnerability or problem is only disclosed after a period of time during which the supporting vulnerability or problem has been fixed or patched. As Figure 3 shown, the responsive disclosure model for open-source software projects can involve the following three steps: (1) the vulnerability is secretly fixed without mention of the vulnerability; (2) the vulnerability is publicly disclosed through an announcement; (3) users of the software update the software in response to the vulnerability announcement. For users of software systems, it is crucial to be aware of vulnerabilities and update their systems in a timely manner.

[0079] In the context of open-source software, a vulnerability can be fixed in step (1) through a source code submission to a public source code repository as a silent fix. In this article, a submission includes three important pieces of information: (i) the commit message; (ii) the modified file name; (iii) the code changes for each file. Figure 4 An example of a submission with modified code (e.g., code added at line 18 and code deleted at line 21) is shown. A silent fix is a submission used to fix a vulnerability where the fix does not include any information that would indicate the vulnerability. For example, the commit message for the submission does not mention the name or nature of the vulnerability. Nevertheless, it is still possible for malicious users to reverse engineer the vulnerability based on changes to the computer code to fix the vulnerability in step (1). Therefore, malicious users can exploit the vulnerability to attack users who have not yet updated their software.

[0080] There may be a time gap between the steps for "silent" fixing of a vulnerability (1) and the steps for public disclosure of the vulnerability through an announcement (2). For example, there is usually a time gap of about 7 to 10 days between step (1) and step (2). This time gap provides an opportunity for malicious users to exploit. Since in the context of open-source software, the source code submissions for fixing vulnerabilities are public, malicious parties may discover and exploit the vulnerability during the time gap before users are notified of the vulnerability. Therefore, for users of open-source software, it is important to detect the silent fix before it is publicly announced.

[0081] In addition, merely identifying a silent fix for a vulnerability is not enough. An explanation of the silent fix should also be provided. Users of software may not be experts in every software they use, and it may be difficult for users to understand the nature of the vulnerability. If users do not understand the nature of the vulnerability, they may be likely to ignore the update, thus rendering such a warning system ineffective. Therefore, it is important to provide some explanation of the vulnerability. For example, the category or exploitability rating of the vulnerability can be provided to help users understand and evaluate the vulnerability.

[0082] The Common Vulnerabilities and Exposure (CVE) database provides a reference method for disclosing, identifying, and managing publicly known vulnerabilities. The National Vulnerability Database (NVD) of the United States is a popular CVE database that provides enhanced vulnerability information, such as the Common Weakness Enumeration (CWE). The CWE provides a dictionary of common weaknesses that may lead to software or hardware vulnerabilities. They include various detailed information about several types of vulnerabilities. By assigning a CWE to a CVE, the CWE can be used to classify the CVE, which provides additional information about the vulnerability. Depending on the nature of the vulnerability, multiple CWEs can be assigned to a CVE, but not every CVE in the NVD is assigned a CWE. Providing the CWE of a silent fix to users can help users understand the nature of the silent fix.

[0083] The Common Vulnerability Scoring System (CVSS) helps to define and classify vulnerabilities based on their potential impact and risk. There are two typical CVSS versions, namely CVSS 2.0 and 3.0. Exploitability is one of the basic group metrics in CVSS and is used to measure the risk of a vulnerability being exploited. The easier a vulnerability is to exploit, the higher the exploitability score of that vulnerability. Therefore, the exploitability metric reflects the risk of a vulnerability and supports users in prioritizing vulnerabilities. For example, a CVSS score can identify a vulnerability as having low, medium, or high risk. Providing users with the CVSS score for a silent fix can help users understand the nature of the silent fix.

[0084] There are many problems with the traditional methods that users use to monitor security updates. For example, users can monitor security bulletins from services such as NVD. However, as mentioned above, due to the response disclosure model, there is usually a gap between the vulnerability fix time and the vulnerability disclosure time. In addition, many vulnerabilities are never disclosed on NVD. Alternatively, users can monitor commits to public source code repositories to determine which commits are vulnerability fixes. The problem with this method is that many fixes are silent fixes, so it is not mentioned that the commit is for fixing a vulnerability. Since users are rarely proficient in the open-source software they are using, it may be difficult to determine which source code commits are for fixing vulnerabilities. In addition, any given software project may have many source code commits every day, most of which are not related to fixing vulnerabilities. This further increases the difficulty of trying to identify the commits used to fix vulnerabilities.

[0085] Another solution is to use VulFixMiner. VulFixMiner is a technical solution that identifies silent fixes for vulnerabilities based on commit-level or file-level code changes. VulFixMiner combines a deep learning solution to analyze the committed source code and then trains a neural network to identify vulnerability fixes. VulFixMiner consists of three phases:

[0086] 1. Fine-tuning phase: Fine-tune a pre-trained language model to learn the representation of file-level code changes.

[0087] 2. Training phase: The fine-tuned model is regarded as a file change transformer and collaborates with a commit change aggregator to encode commit-level code changes into commit-level code change representations. Then train a neural network classifier to use these representations to identify commits.

[0088] 3. Application phase: The trained VulFixMiner fetches new commits from the open-source software repository and calculates scores that indicate the likelihood that a commit is used to fix a vulnerability.

[0089] There are many drawbacks to using VulFixMiner. Due to limited and diverse data, it is challenging to identify silent fixes and provide explanations. The vast majority of source code commits are not related to vulnerability fixes. Therefore, the data available for training the neural network is limited. Additionally, the fixed vulnerabilities are associated with a wide range of CWE categories, indicating the diversity of the causes, behaviors, and consequences of vulnerabilities, which leads to the diversity of corresponding fix patterns. The limited and diverse training data results in the neural network being unable to produce reliable results.

[0090] VulFixMiner uses the code segments added and deleted in the entire commit to identify silent fixes, rather than using function-level changes. A single commit may address different issues. For example, a single commit can fix a vulnerability and add features. Due to the mixed information from the entire commit and the lack of code context information, it is difficult for VulFixMiner to provide explanations for different fixes. VulFixMiner can be used to identify vulnerability fixes, but it cannot provide explanations or ratings for these vulnerability fixes.

[0091] VulFixMiner requires supervised learning. VulFixMiner requires pre-labeling of code changes to understand which code changes are vulnerability fixes. VulFixMiner cannot be trained using unsupervised learning. Therefore, training VulFixMiner is very time-consuming, and the training data available for training VulFixMiner is limited, resulting in lower reliability of the results. In other words, the two main drawbacks of VulFixMiner are that it has no way to augment the limited available code change data and no way to train the model in an unsupervised manner.

[0092] According to some embodiments, contrastive learning is used to train the neural network. Contrastive learning is widely applied in the fields of computer vision and natural language processing (NLP). The key to contrastive learning is data augmentation. By applying augmentations to a data point to generate two different but semantically similar samples, contrastive learning attempts to learn the similar knowledge in the samples from the same data point and learn the differences between the samples generated from different data points. For example, in the NLP field, data augmentation is done by operating on tokens, such as token reordering and similar token replacements. In the field of software engineering, previous research has mainly focused on source code. Based on the methods in NLP, previous research has further proposed sampling / augmentation strategies based on the compilation mechanism to generate source code samples. For example, they use code compression, identifier modification, and regularization. These methods are able to learn source code representations, but none of them are able to learn source code change representations.

[0093] Now refer to Figure 5A andFigure 5B , which shows three stages of a method 400 for training a neural network according to some embodiments of the present invention. Stage 1 includes function change data augmentation 410. In Stage 1, code change data is augmented at the function level. More specifically, Stage 1 combines program slicing techniques and CWE category information to augment function changes using unsupervised (i.e., based on itself) and supervised (i.e., based on groups) methods. A single function change from a patch or commit is augmented into a set of semantically-preserved function change samples (FCSamples). Every two semantically similar or functionally similar FCSamples can be regarded as a positive pair for contrastive learning in the next stage. Stage 2 includes function change representation learning 420. The contrastive learner effectively learns the representations of different repair data by minimizing the distance between positive samples (similar data representations) and maximizing the distance between negative samples (dissimilar data representations). The contrastive learner learns function-level code change representations from different repair data and trains the neural network. Stage 3 includes downstream task fine-tuning 430. In Stage 3, the neural network can be further fine-tuned. In some embodiments, the neural network is fine-tuned to produce a silent repair identification model, a CWE classification model, and an exploitability rating classification model. The method is applicable to developing other types of models, such as a severity classification model.

[0094] Now refer to Figure 6 , which shows a method 500 for training a neural network to learn computer code change representations. Referring also to Figure 7, which shows a schematic diagram of a method for enhancing computer code change data, corresponding to the data enhancement step 410 of phase 1 of method 400. Method 500 includes dividing a segment of computer code 601 into multiple computer code parts 510. The computer code can be source code, intermediate code, machine code, or any other type of code that can be read, interpreted, or compiled by a computer. In a preferred embodiment, the segment of computer code 601 is a function. However, any other segment of computer code can be used, such as a file, class, or data structure. Dividing the segment of computer code 601 into multiple computer code parts may include using a program slicing module 604 to generate function slices 605 (FSlice) for the original function 602 and the modified function 603. The slices 605 correspond to the computer code parts. For each function change, function slices 605 are generated for the original function 602 (OriFSlice) and the modified function 603 (ModFSlice). Since the changed code statements between the original function 602 and the modified function 603 fix the same vulnerability, the changed variables in the changed code statements can be used as anchors for slicing. Other anchors can also be used for slicing. Function changes can be represented in a single file using change tracking or diff notation, which indicates which lines have been deleted and which lines have been added. Alternatively, function changes can be represented by two files, where one file represents the original computer code and the other file represents the modified computer code.

[0095] The slices 605 can be comprehensive slices 605 that incorporate aspects of both forward and backward slices. Functions can be divided into multiple computer code parts or slices based on the changed variables as anchors using a control flow graph or a data flow graph. A control flow graph (CFG) and a data flow graph (DFG) can be used to generate the slices 605 because the combination of these graphs preserves the structural integrity of the original program and extracts the data relationships between variables in the program. A source code parsing tool (such as TreeSitter) can be used to generate the CFG and DFG. Other types of computer graphs and parsing tools can be used to generate the slices 605. For each anchor, the corresponding code statements are extracted from these paths in order to create FSlice 605 for the function based on the changed variables.

[0096] Now refer to Figure 9, showing a schematic diagram of function code changes. The function code change 801 shows the source code lines deleted and added from the function. The function code change 801 involves two different variables: "serverId" and "base". The first OriFSlice 803 shows the slice generated based on the original function using the serverId variable as an anchor. The first ModFSlice 804 shows the slice generated based on the modified function using the serverId variable as an anchor. The second OriFSlice 806 shows the slice generated based on the original function using the base variable as an anchor. The second ModFSlice 807 shows the slice generated based on the modified function using the base variable as an anchor. In other words, the function code change 801 has been used to generate four slices, two original slices and two modified slices. It should be noted that not every function change includes a changed variable. For example, some function changes are related to function call renaming or operator changes. In this case, the function does not have slices based on the changed variable. Therefore, the complete function can be used without slicing. In other words, the function change will generate a single OriFSlice and a single ModFSlice.

[0097] Multimodal pre-training can help text-based models learn the implicit alignment between different modal inputs, e.g., the alignment between natural language and programming language. The FCSample 606 can include an automatically generated description 611. The function change description 611 (FCDesc) can be included in the sample 606 as supplementary information to enhance the enhanced function change sample 606. The function change description 611 can be generated using the function change description generator 610, such as GumTree Spoon ASTDiff. GumTree generates a list of change operations for each pair of original and modified functions. The GumTree tool is able to identify insert and delete change operations, as well as rename or move operations, providing detailed information about the changes. Figure 10 Showing an example of FCDesc for a patch that fixes a cross-site scripting vulnerability in Apache ActiveMQ.

[0098] Method 500 also includes generating a first change sample 606, the first change sample 606 including a first original computer code segment (e.g., OriFSlice) and a first modified computer code segment (e.g., ModFSlice), the first change sample 606 including at least one of a plurality of computer code portions (i.e., the first change sample includes at least one of the generated function slices 605) 520. Method 500 also includes generating a second change sample, the second change sample including a second original computer code segment (e.g., OriFSlice) and a second modified computer code segment (e.g., ModFSlice) 530. The function change enhancer module 612 can construct an FCSample 606 for the function change as follows:

[0099]

[0100] where ⊕ is the concatenation operator, and "i" and "j" are the i-th and j-th OriFSlice and ModFSlice respectively.

[0101] Figure 9 Two exemplary FCSamples are shown. The first FCSample 802 includes a first OriFSlice 803 and a first ModFSlice 804. The second FCSample 805 includes a second OriFSlice 806 and a second ModFSlice 807. The first original segment and the first modified segment can correspond to the same function. That is, the FCSample 606 can include slices generated from the same function. The FCSample 606 does not need to include slices with the same variables as the anchor point. The FCSample 606 can include any slices from the same function. For example, there can be an FCSample including the first OriFSlice 803 and the second ModFSlice 807. Since the slices are from the same function, they may have the same semantic meaning (i.e., they relate to the same computer code fix). In fact, in some embodiments, slices from the same class, data structure, or file can be combined together in the same sample. The FCDesc 611 for the function change can also be added to the sample 606. This way of generating the FCSample 606 enhances the available data for training the neural network.

[0102] The multiple computer code portions include multiple original computer code portions (e.g., OriFSlice) and multiple modified computer code portions (e.g., ModFSlice), wherein the first original computer code segment includes the first original computer code portion among the multiple original computer code portions, wherein the first modified computer code segment includes the first modified computer code portion among the multiple modified computer code portions, wherein the second original computer code segment includes the second original computer code portion among the multiple original computer code portions, and wherein the second modified computer code segment includes the second modified computer code portion among the multiple modified computer code portions. That is, each FCSample includes an OriFSlice and a ModFSlice.

[0103] If a commit includes several different function changes, a single patch or commit can result in several FCSample606. Additionally, if a single function change includes changes related to different variables, a single function change may result in several FCSample 606. For example, in the function code change 801, since the change is related to two different variables, four different FCSample 606 can be generated: OriFSlice 1+ModFSlice1, OriFSlice 2+ModFSlice 2, OriFSlice 1+ModFSlice 2, and OriFSlice 2+ModFSlice 1. Compare this with VulFixMiner, which uses an example of a single patch that includes changes to three functions, where each function has two variable changes. For VulFixMiner, this patch can generate a single training sample. In contrast, according to the present invention, this single patch can generate twelve training samples for training a neural network. This data augmentation technique improves the reliability of the trained neural network. To avoid the possibility of overfitting, the number of FCSample606 from a single function change can be restricted. For example, the number of FCSample 606 from a single function change can be restricted to four. The four selected FCSample 606 can be randomly selected from the total number of FCSample 606.

[0104] To train a neural network, FCSamples 606 can be combined into positive sample pairs by a related sample pair constructor module 607. Then, the neural network will attempt to minimize the differences between the positive sample pairs. Using the FCSamples 606 and the information (FC_CWE) of the CWE category 609 for each function change, the related sample pair constructor can generate positive FCSample pairs 608. If two FCSamples 606 are related (e.g., their semantic meanings are similar, or their functional meanings are similar), then they are positive function change sample pairs.

[0105] There are two ways to construct positive sample pairs. The first way is an unsupervised function-based method, similar to general data augmentation techniques. By this method, the first change sample and the second change sample correspond to the same function. If FCSamples 606 are generated from the same data instance, they can be combined into a positive sample pair. For example, if two FCSamples 606 are generated from the same function, they can be combined into a positive sample pair. Computer code from other segments can be used. For example, positive sample pairs can be constructed based on FCSamples 606 generated from the same file, class, or data structure. Since the two FCSamples 606 are derived from the same function, they can be semantically similar to each other (i.e., they fix the same type of vulnerability). If a function change cannot generate multiple FCSamples 606 (e.g., due to no changed variables), the function change cannot be used in this method.

[0106] The second way is a group-based supervised method, which uses the FC_CWE 609 information of function changes to construct positive pairs. By this method, the first change sample and the second change sample can belong to the same category, vulnerability category, or more specifically, the same CWE category 609. For example, for a group of FCSamples 606 of different function changes that fix the same type of vulnerability (i.e., the same FC_CWE 609), the FCSamples 606 in the same group can be functionally similar. Therefore, such FCSamples 606 in the same CWE category 609 can be used to create positive pairs 608. In addition to FC_CWE 609, other labels or groups can be used to group FCSamples 606. In some embodiments, the priority of the first method can be set higher than the priority of the second method so that the group-based method is only used when a function cannot generate more than one FCSample 606.

[0107] Method 500 also includes: calculating a loss function based on the first change sample and the second change sample 540, and training the neural network by minimizing the loss function 550. Now refer toFigure 8 shows a schematic diagram of a method 700 for training a neural network to learn representations of computer code changes, corresponding to the function change representation learning step 420 of stage 2 of method 400. To learn the representation of function changes, a contrastive learner can be used, which can effectively learn data representations by minimizing the distance between similar data (positive) and maximizing the distance between dissimilar data (negative). Therefore, using the constructed positive sample pairs 608, the contrastive learning method can effectively learn function change representations from different vulnerability fixes. The mini-batch permutator 702 can permute the inputs in a mini-batch 703, where all positive pairs within the mini-batch are associated with different CWE categories 609. In this way, any sample in a pair is negatively correlated with any sample in other pairs within the mini-batch. Next, the encoder 704 (e.g., FCBERT) is further pre-trained to encode the function changes into their embedded representation vectors 707. Then, the projection head 708 maps the vector 707 to the space where the contrastive loss is applied.

[0108] The mini-batch permutator 702 permutes n relevant sample pairs from the candidate pairs 608 into a mini-batch 703. The mini-batch permutator uses the CWE category 609 to ensure that each pair in a single mini-batch 703 corresponds to a different CWE category 609. That is, the sample pair 705 has a different CWE category 609 from the sample pair 706. Other methods for distinguishing the semantic meaning or function of the sample pairs 608 can be used instead of using the CWE category 609.

[0109] The pre-trained encoder 704 is used to encode each FCSample 606 in the positive sample pairs 608 into its corresponding function change representation vector 707. The pre-trained encoder FCBERT 704 with the same architecture and weights as CodeBERT can be used.

[0110] The non-linear projection head 708 helps to improve the representation quality of the layers in front of it. A multi-layer perceptron (MLP) with two hidden layers can be used to project the function change representation vector 707 into the space where the contrastive loss is applied.

[0111] A contrastive loss function can be defined to maximize the consistency of samples within the same relevant sample pairs and minimize the consistency between samples from different sample pairs. According to one embodiment, a noise contrastive estimate (NCE) loss function can be used to calculate the loss. For example, the loss function can be minimized between samples in the same positive sample pair 705, and the loss function can be minimized between samples in the same positive sample pair 706. Method 500 may further include: generating a third modified sample, calculating a loss function from the first modified sample and the third modified sample, and training a neural network by maximizing the loss function. That is, the loss function can be maximized between samples from sample pair 705 and samples from sample pair 706. Since they belong to different CWE categories 609, they may have different semantic meanings.

[0112] Method 500 may further include obtaining the segment of computer code from a security bulletin service or a common vulnerabilities and exposure (CVE) database. A common CVE is the NVD. CVEs (such as the NVD) and more general security bulletin services publish known software vulnerabilities. A CVE may publish the source code that caused the vulnerability or the source code changes used to fix the vulnerability. Thus, the source code provided on a CVE can be used to train a neural network to detect silent fixes for vulnerabilities. The source code obtained from a CVE can be manually downloaded and input into computer 104. Alternatively, computer 104 can automatically download the source code from a CVE server 102 via network 108.

[0113] Method 500 may further include training the neural network in an unsupervised manner. In some embodiments, contrastive learning can be used to train the neural network. A contrastive learner can effectively learn data representations by minimizing the distance between similar data (positive) and maximizing the distance between dissimilar data (negative). Since the semantic similarity of samples 606 is inferred based on samples 606 derived from the same function, class, or data structure, the user does not need to label the samples 606. Thus, the neural network can train itself in an unsupervised manner based on samples 606 without any input or labeling by the user. This reduces the amount of work required to train the neural network and increases the amount of training data that can be reasonably used, thereby improving the reliability of the neural network. As another option, the neural network can be a siamese neural network. In fact, any type of neural network can be used that requires two inputs.

[0114] Method 500 may also include the step of fine-tuning the neural network for a task, corresponding to the downstream task fine-tuning 430 in stage 3 of method 400. The encoder FCBERT 704 can be used as a pre-trained model to initialize other fine-tuned encoders by transferring weights from the pre-trained encoder 704 to the other fine-tuned encoders. For example, as Figure 11 shown, the encoder 704 can be used to initialize the FixEncoder, CWEEncoder, and EXPEncoder.

[0115] The goal of the silent fix identification task is to predict the probability that a commit is submitted to fix a vulnerability. VulFixMiner uses CodeBERT as a pre-trained model to fine-tune the task. CodeBERT in VulFixMiner may be replaced by the FixEncoder. Except for the pre-trained model, the architecture of VulFixMiner can remain unchanged, and the input structure can remain unchanged. The input to the task can be general commit data and patch data (i.e., the commit used to fix the vulnerability). For each commit, the neural network outputs a score, which represents the probability that the commit is used to fix the vulnerability. This neural network can be called CoLeFunDa_fix.

[0116] The goal of the CWE classification task is to predict the probability that a given function change in a patch is used to fix a specific CWE category. The input to this fine-tuning task can be patch data. More specifically, the function change description, the complete original function, and the complete source code of the modified function. This input is first encoded by the CWEEncoder into a function change representation vector. Then this vector is fed into a two-layer neural network to calculate the probability score for each CWE category. It should be noted that since a patch can be used to fix vulnerabilities with multiple CWE categories, this task can be regarded as a multi-label classification task and binary cross-entropy is used as the loss function. This neural network can be called CoLeFunDa_cwe.

[0117] The goal of the exploitability rating classification task is to predict the probability of the exploitability rating of the fixed vulnerability. The input and fine-tuning process in this task is similar to the CWE classification task, except for the loss function. Since a vulnerability has only one exploitability rating, this task can be considered a multi-class classification task and cross-entropy is used as the loss function. This neural network can be called CoLeFunDa_exp.

[0118] The neural network CoLeFunDa_fix can be used to calculate the probability that a computer code change fixes a vulnerability. Given a set of commits, CoLeFunDa_fix first calculates the probability scores and then outputs a list of commits sorted by the predicted probability. The higher the score of a commit, the higher the chance that the commit fixes the vulnerability.

[0119] The neural network CoLeFunDa_cwe can be used to calculate the probability that a computer code change belongs to a certain category (e.g., CWE category). Given a commit confirmed to fix a vulnerability, for each function change in the commit, CoLeFunDa_cwe calculates the score for each CWE category as follows:

[0120] CWE j Score i = CoLeFunDa cwe (FC i ) (2)

[0121] where FC i is the i-th function change in the commit, and CWE j Score i is the score for the j-th CWE category. The CWE score for the commit is calculated as follows:

[0122]

[0123] where n is the number of function changes in the commit. The CWE categories are sorted by score, and the higher the score, the greater the probability that the commit is used to fix a vulnerability in that specific CWE category.

[0124] The neural network CoLeFunDa_exp can be used to assign a rating to a vulnerability, such as an exploitability rating or a severity rating. The exploitability rating indicates how easy it is to exploit the vulnerability. The severity rating indicates how severe the consequences might be if the vulnerability is exploited. Given a commit confirmed to fix a vulnerability, for each function change in the commit, CoLeFunDa_exp calculates the score for each possible exploitability rating as follows:

[0125] EXP j Score i = CoLeFunDa exp (FC i ) (4)

[0126] where EXP j Score i is the score for the j-th exploitability rating of the i-th function change. The commit-level score for the exploitability rating is calculated as follows:

[0127]

[0128] where n is the number of function changes in the commit. The exploitability ratings are sorted by score, and the higher the score, the greater the probability that the commit is used to fix a vulnerability with a specific exploitability rating. A similar method can be used to calculate the severity rating.

[0129] It should be noted that CoLeFunDa_fix, CoLeFunDa_cwe, and CoLeFunDa_exp can be used individually or in sequence. To achieve better early vulnerability awareness, open-source software users can integrate CoLeFunDa_fix, CoLeFunDa_cwe, and CoLeFunDa_exp into the automated open-source software code repository monitoring pipeline. When new code changes are pushed to the public repository, CoLeFunDa_fix may first identify whether the commit is for fixing a vulnerability. If so, CoLeFunDa_cwe and CoLeFunDa_exp can also provide explanations of the relevant CWE categories of the vulnerability and exploitability ratings.

[0130] The neural network learns the representation of general function changes in computer code. This neural network can be used in other applications. For example, the neural network can be used for immediate defect prediction in computer code or for generating commit messages for source code commits to the source code repository. Other applications include detecting undisclosed vulnerabilities, summarizing the health of software projects, summarizing release goals, identifying project mentors or experts, generating documentation for software projects, CVE patch matching (i.e., determining the patches used to fix specific CVEs), and automated code review. The training methods disclosed herein can be used for various purposes, such as training a function change representation model, training a machine learning model, or training a generative adversarial network (GAN) model.

[0131] Method 500 can be executed by the processor of the client computing device 104. The security bulletin service or CVE can be hosted by one or more server computers 102. The client computing device 104 can download vulnerability information from the CVE server 102 for training the neural network through the network 108. Another server 102 can host a source code repository, such as GitHub. The client computing device 104 can monitor source code commits to the source code repository server 102 and use the neural network to determine whether the purpose of the source code commit is to fix a vulnerability. Thus, the client computing device 102 can provide an early warning of the existence of vulnerabilities to users of the software hosted on the source code repository server. Method 500 can be implemented in various forms, such as a cloud service, a plugin, or a client desktop application. Method 500 can also be executed by the processor of the server computer 102.

[0132] Although the embodiments have been described above with reference to the accompanying drawings, those skilled in the art will understand that changes and modifications can be made without departing from the scope defined by the appended claims.

Claims

1. A method for training a neural network, characterized in that, Comprising: Dividing a piece of computer code into multiple computer code parts; Generating a first changed sample; Generating a second changed sample; Calculating a loss function based on the first changed sample and the second changed sample; Training the neural network by minimizing the loss function.

2. The method according to claim 1, characterized in that The first changed sample includes a first original computer code segment and a first modified computer code segment.

3. The method according to claim 2, wherein The first original segment and the first modified segment correspond to the same function.

4. The method according to claim 2, wherein The multiple computer code parts include multiple original computer code parts and multiple modified computer code parts, wherein the first original computer code segment includes the first original computer code part among the multiple original computer code parts, and the first modified computer code segment includes the first modified computer code part among the multiple modified computer code parts.

5. The method according to claim 1, wherein The second changed sample includes a second original computer code segment and a second modified computer code segment.

6. The method according to claim 5, wherein The second original computer code segment includes the second original computer code part among the multiple original computer code parts, and the second modified computer code segment includes the second modified computer code part among the multiple modified computer code parts.

7. The method according to claim 1, characterized in that, The first changed sample and the second changed sample correspond to the same function.

8. The method according to claim 1, wherein The first changed sample and the second changed sample belong to the same category.

9. The method according to claim 1, characterized in that The first changed sample and the second changed sample both repair vulnerabilities of the same category.

10. The method according to claim 1, characterized in that, The first changed sample further includes: an automatically generated description, or a manually marked description, or a combination of the automatically generated description and the manually marked description.

11. The method according to claim 1, characterized in that, The piece of computer code is a function.

12. The method according to claim 11, characterized in that, The function is divided into multiple computer code parts based on a control flow graph or a data flow graph according to changed variables.

13. The method according to claim 1, wherein Also comprising: Generating a third changed sample; Calculating the loss function according to the first changed sample and the third changed sample; Training the neural network by maximizing the loss function.

14. The method according to claim 1, characterized in that, The piece of computer code is obtained from a security bulletin service or a common vulnerability disclosure database.

15. The method according to claim 1, wherein The neural network is trained in an unsupervised manner.

16. The method according to claim 1, characterized in that, The neural network is trained using contrastive learning, or the neural network is a siamese neural network.

17. The method according to claim 1, characterized in that, Also comprising fine-tuning the neural network for a task.

18. The method according to claim 1, wherein The computer code is source code, intermediate code, or machine code.

19. A non-transitory computer-readable medium, characterized in that, Including computer program code stored therein for training a neural network, which when executed by one or more processors causes the one or more processors to execute a method, the method comprising: Dividing a piece of computer code into multiple computer code parts; Generating a first changed sample; Generating a second changed sample; Calculating a loss function based on the first changed sample and the second changed sample; Training the neural network by minimizing the loss function.