Identifying security-relevant commits through architectural context

By integrating architecture context with commit data and employing machine learning models, the method effectively identifies security-related commits in open-source software, addressing the challenge of timely vulnerability detection and mitigation.

JP2025096224APending Publication Date: 2025-06-26エスアーペーエスエー
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024217524
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-12
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The software industry faces challenges in detecting security-related changes in open-source software libraries in a timely manner, as delays can allow attackers to exploit vulnerabilities between commit and release.

Method used

A method and apparatus that analyze commits by combining information about the architecture context with commit data, using machine learning models to predict the likelihood of security-related changes, and generating vectors to represent commits, messages, and architecture recognition for accurate identification.

Benefits of technology

This approach enables efficient and accurate identification of security-related commits, allowing for timely vulnerability detection and mitigation, thereby reducing the window of opportunity for attackers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096224000001_ABST
    Figure 2025096224000001_ABST
Patent Text Reader

Abstract

To enable identification of security-relevant commits that may provide an opportunity for a malicious party to identify vulnerabilities in software that a malicious party can take advantage of.SOLUTION: Methods and apparatuses described herein combine a commit with information related to architectural context in which changes appear to predict the likelihood that a given commit contains security-relevant software code changes. The security-relevant code changes may be code changes for fixing an existing security risk or code changes that may introduce a potential security risk.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Unless otherwise specified in this specification, the techniques described in this section are not prior art to the claims of this application and are not admitted to be prior art by inclusion in this section.

[0002] As more software products are built as a composition of proprietary closed-source components and free open-source software (OSS) libraries, the adoption of third-party open-source components has permeated the software industry today. Also, while the use of these OSS libraries can increase the speed of the development process, it is also a fact that the practices of quality assurance and the maturity of the developers of these third-party libraries can vary considerably. Third parties may provide updates to their software periodically, and it is important to detect security-related development changes in a timely manner. Delays in detecting and reacting to security-related changes leave an attacker with a prime opportunity, and the attacker monitors the repositories of popular components, observes security fixes in the updates, infers that the security fixes address security vulnerabilities, and may exploit the security vulnerabilities during the time period between when the commit is pushed to the repository and when the release is made public to the client projects that adopt it. Therefore, it is necessary to accurately and efficiently identify security-related commits.

Brief Description of the Drawings

[0003]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0004] This specification describes a method and apparatus for identifying security-related commits. A commit is an operation in a version control system that sends software (or source) code changes to a repository. A commit can be retrieved from the repository by a target system to update the software running on the target system at a later time. A commit created to fix a security problem in software is known as a security-related commit. A commit that may introduce a security problem into the software is also known as a security-related commit. Identifying security-related commits is important because they may provide an opportunity to identify vulnerabilities in software that can be exploited by malicious parties. This is becoming more of an issue as software solutions include more software components built by multiple independent third parties. The method and apparatus described in this specification combine information related to the context of the architecture in which those changes appear with the commit to predict the likelihood that a given commit contains security-related software code changes. To expand the characterization of a commit, annotations may be added to elements of the architecture (structurally such as components, connectors, interfaces, etc. or behaviorally such as protocols). This expanded representation (commit message + commit source code + architecture context) can be used to train more accurate machine learning models for different code analysis tasks. However, it will be apparent to those skilled in the art that the embodiments defined in the claims may include some or all of the features of these examples alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described in this specification.

[0005] Figure 1 shows a workflow for identifying security-related commits according to some embodiments. Workflow 100 may be executed on one or more computing systems. Exemplary computer systems are described in FIGS. 5 and 6. As shown, source code 105 may belong to a software program, may include a number of files, and each file may contain source code representing different parts of the software program. The software program may be a program developed by the target system or may be a program developed by a third party. When changes are made to the software program, a commit 110 is generated and may be stored in a version control system such as a version-controlled repository. Workflow 100 begins by receiving commit 110 from the version repository. Workflow 100 can analyze commit 110 considering the context of the architecture of source code 105 to predict whether commit 110 is a security-related commit. The prediction may be output from a machine learning model 160.

[0006] Commit 110 can include two parts, namely message 115 and source code change 130. Source code change 130 may store the actual changes made to the source code. In one example, the changes may be represented as lines of code being deleted and lines of code being added to a file of the source code. In another example, the changes may be represented as words being deleted and added to the software code. Message 115 may include natural language words, sentences, or paragraphs that provide an explanation of the purpose behind source code change 130. For example, the explanation may be that the source code change enables encryption or that the source code change corrects a security vulnerability in the software. Message 115 or source code change 130 may also include an overview of the files that were changed and / or the total number of additions and deletions.

[0007] Workflow 100 continues by processing message 115 with natural language processing (NLP) model 120 to generate message vector 125. The conversion from the text commit message to a numerical vector is performed using one of many existing text embedding methods (e.g., word2vec, etc.). The result of vectorizing the commit message is an n-dimensional numerical vector (where n is typically in the hundreds or thousands). Message vector 125 is one input to machine learning model 160.

[0008] Workflow 100 may also process source code change 130 with commit2vec 135 to generate commit vector 140. Commit2vec 135 is a neural network model or a machine learning model. The commit2vec model is fine-tuned to generate a commit vector representing the source code so that a classifier model can use the information encoded in those vectors to identify security-related code changes. Thus, the generated commit vector 140 may represent the source code itself. In one embodiment, representing the source code change as a commit vector is to protect and potentially highlight what is important for determining the security relevance of the code, while simply compressing the information. Commit vector 140 is a second input to machine learning model 160.

[0009] Workflow 100 can also process source code change 130 (at 150) with the annotated architecture model 145 to generate architecture recognition commit vector 155. Details regarding the process for generating architecture recognition commit vector 155 are described in FIG. 2. Architecture recognition commit vector 155 is a third input to machine learning model 160. Similar to the commit vector, the architecture recognition commit vector may be a multi-dimensional numerical representation.

[0010] In one embodiment, the generation of the message vector 125, the commit vector 140, and the architecture recognition commit vector 155 may occur in parallel, although in other embodiments the vectors may be generated in any order. Once all three vectors are generated, the workflow 100 may continue by utilizing a machine learning model 160 to process the three input vectors to generate a prediction as to whether the commit 100 is a security-related commit. In one example, the prediction may be a value between 0 and 1 (where 0 is the lowest and 1 is the highest) representing the likelihood that the commit is a security-related commit. In one embodiment, the workflow 100 predicts whether the commit includes source code changes that fix potential security risks. If a commit is likely to address a security issue in a software project, it may be desirable for the target system to prioritize incorporating that commit into its software. In another embodiment, the workflow 100 predicts whether the commit includes source code changes that may present a security risk. If a commit is likely to introduce a security issue with a software project, the target system may implement safeguards to protect against attacks by malicious actors. For example, if a commit is likely to have a security risk, the target system may decide not to incorporate the commit into its source code. The target system may also share this knowledge with the developers of the software project to investigate potential security risks. The target system may also consider updating its software or system to minimize potential security risks. It will be understood by those skilled in the art that the neural network model of the workflow 100 may be trained to predict source code changes that may introduce security risks or vulnerabilities, source code changes that may fix existing security risks or vulnerabilities, or both.

[0011] In one embodiment, the machine learning model may receive a concatenation of a message vector 125, a commit vector 140, and an architecture recognition commit vector 155. The machine learning model 160 may be a neural network model trained to identify security-related commits. Security-related commits are defined as commits having vulnerabilities or other security-related risks.

[0012] Figure 2 shows a workflow for generating an architecture recognition commit vector according to some embodiments. The workflow 200 may be executed on one or more computing systems having sufficient memory and computing power. The computing system may be one or more of the computing systems of FIG. 1. The workflow 200 begins with the architecture model extractor 210 receiving the source code 205 to generate an architecture model 215. In one example, the source code 205 may be a snapshot of the source code 105 at a particular point in time retrieved from a source code repository. The architecture model extractor 210 may extract a source code tree from the source code repository and generate an abstract representation of the architecture of the software project, known as the architecture model 215. In one example, an existing architecture model of the software project may be used. In other examples, the architecture model 215 may be automatically extracted from the source code 205 by using existing architecture reconstruction techniques.

[0013] In one embodiment, the mapping between the software code and the architecture model 215 is done when the architecture model 215 is extracted from the source code by the architecture model extractor 210. In one example, each line of the software code may be mapped to one architecture element. In other examples, the mapping may be a many-to-one mapping (multiple lines of software code are mapped to one architecture element, or one line of software code is mapped to multiple architecture elements).

[0014] In another embodiment, the architecture model extractor 210 includes a graphical user interface to assist the software developer / architect in manually linking architecture elements to several lines of code in step 210. The user interface may present a view of the architecture model in which elements without corresponding source code elements associated with them are highlighted in a suitable diagrammatic way (such as icons, colors, line types or thicknesses). For each of these elements, the software developer / architect can obtain a set of code elements that can be assigned to it. The set may be presented in a hierarchical view so that the software developer / architect can navigate the code base by following the natural structure of folders and files, packages, classes, data type definitions, methods / functions, etc.

[0015] Workflow 200 can continue by processing architecture model 215 with architecture model security analyzer (analyzer) 220. Architecture model security analyzer 220 may be configured to annotate the architecture model to identify one or more architecture elements in architecture model 215 that are security risks or potential security risks. The output of architecture model security analyzer 220 is annotated architecture model 225. Annotated architecture model 225 is a model in which each of the related architecture elements (components and connectors) is an annotated model annotated with security-related information. For example, the annotation can indicate that the component is a source or sink of confidential data, a data storage element, a component or communication channel protected by authentication or authorization or encryption, a channel that crosses a trust boundary, etc. This additional information may be provided as an annotation added manually by security architect 290, by an automated mechanism, or by a combination of the two.

[0016] Workflow 200 can continue with a contextualiser 240. The goal of the contextualiser 240 is to identify architectural elements in the annotated architectural model 225 affected by the code change 230. These architectural elements may be flagged so that they can be analyzed to predict whether they pose security-related risks. The code change 230 may be a source code change stored in a commit. For example, the code change 230 may be the source code change 130 of FIG. 1. In one example, each line of the modified software code may be mapped to one architectural element. In other examples, the mapping may be a many-to-one mapping (multiple lines of software code are mapped to one architectural element, or one line of software code is mapped to multiple architectural elements). In one embodiment, the architectural model extractor contextualiser 240 may include a graphical user interface to assist software developers / architects in flagging architectural elements of the annotated architectural model 225 affected by the code change 230. After the mapping, the contextualiser 240 generates a contextualized architectural model 245 that is simply the annotated architectural model 225 with information about the code change.

[0017] Workflow 200 can continue with vectorizer 250. Vectorizer 250 is configured to generate a distributed representation from the contextualized architecture model 245. The distributed representation may be a fixed-size vector. In one embodiment, vectorizer 250 can execute at least two tasks. The first task is the extraction of graph paths. In the extraction of graph paths, all paths connecting two specified nodes of the contextualized architecture model 245 are extracted. This may be done for all pairs of nodes in the graph that meet pre-established conditions. For example, the pre-established conditions may be that all paths must be between a minimum and maximum threshold and must include at least a node or edge affected by code changes. As another example, the pre-established conditions may be that all paths must start from and end at a terminal node and must include at least a node or edge affected by code changes. FIG. 3 shows an example of the extraction of graph paths.

[0018] The second task performed by the vectorizer 250 is the generation of a distributed representation. The generation of the distributed representation can include an embedding step that encodes each of the elements (nodes or edges) of the extracted path. In one embodiment, the embedding matrix is initialized from the vocabulary of the graph's nodes and edges. Training may be applied to the embedding matrix to generate meaningful representations, i.e., similar elements are close to each other in the embedding space. In one example, the embedding matrix may be pre-trained in the manner of a masked language model (MLM) where each of the components of the path is treated similarly to the words of a natural language sentence. Randomly selected words may then be omitted, and the model is trained to predict these omissions using the context provided by the surrounding words (here nodes and edges). The pre-training is applied to a large dataset of the architecture model. These models do not need to be annotated. The result of the encoding step may be a vector representation corresponding to the extracted path. The global attention mechanism may then aggregate the set of vectors corresponding to the set of extracted paths and generate an architecture recognition commit vector 255 that represents potential security risks related to the source code change 230. In one example, the architecture recognition commit vector 255 may be a single fixed-size vector. In another example, the architecture recognition commit vector 255 may be the architecture recognition commit vector 155 of FIG. 1. The architecture recognition commit vector 255 may then be fed into the machine learning pipeline 260, and in some examples, the machine learning pipeline 260 may be the machine learning model 160 of FIG. 1.

[0019] In some implementations, the annotated architecture model 225 may be stored and utilized in preparation for future software code changes made to a software project. For example, during the evaluation of a subsequent commit, the contextualizer 240 may receive the stored annotated architecture model 225 along with the code changes in the subsequent commit in order to analyze the commit without generating the annotated architecture model twice. In order to reuse the annotated architecture model in this way, the changes made to the software code will be used to update the annotated architecture model to keep it up-to-date. For example, in workflow 200, since the source code has been changed due to a change in the software code, the annotated architecture model 225 may be updated to account for the code change 230. This update step may be performed at any time after the contextualizer 240. For example, the update step may be performed before the vectorizer 250, after the vectorizer 250, or after the machine learning pipeline 260.

[0020] Figure 3 shows an example of the extraction of a graph path according to some embodiments. As shown, the contextualized architecture model 300 is a graph structure that includes nodes and edges. Node A 320, Node B 330, and Node C 370 have tags in their upper right corners indicating that they perform functions that may potentially be affected by security or other vulnerabilities. In this example, Node A 320 provides a user authentication function, Node B 330 performs a major business logic operation, and Node C 370 stores and provides access to business confidential data. Some other functions that may be related to security include access control, cryptography, injection, design flaws, security misconfigurations, legacy components, identification, authentication, data integrity, monitoring, and server-side forgery.

[0021] During the extraction of graph paths, all paths connecting two specified nodes of the contextualized architecture model 245 are extracted. This may be done for all pairs of nodes in the graph that meet pre-established conditions. The pre-established conditions may be that the paths involve nodes (i.e., nodes A, B, and C) that are affected by source code changes and potentially have security implications.

[0022] Now, assume that a code change affects node A. If the pre-established condition is that the extracted paths (where the nodes are endpoints of the graph) including the terminal nodes contain the nodes affected by the code change and are potential security risks, two paths may be retrieved from the graph path extraction. Cloud 310 - Edge 1 315 - Node A 320 - Edge 2 325 - Node B 330 - Edge 4 365 - Node C 370 - Edge 5 375 - VM1 380 Cloud 310 - Edge 1 315 - Node A 320 - Edge 2 325 - Node B 330 - Edge 3 335 - Node D 340 - Edge 6 345 - Node E 350 - Edge 7 355 - VM2 360

[0023] As another example, assume that a code change affects node D. If the pre-established condition is that the extracted paths (where the nodes are endpoints of the graph) including the terminal nodes contain the nodes affected by the code change and are potential security risks, two paths may be retrieved from the graph path extraction. VM1 380 - Edge 5 375 - Node C 370 - Edge 4 365 - Node B 330 - Edge 3 335 - Node D 340 - Edge 6 345 - Node E 350 - Edge 7 355 - VM2 360 Cloud 310 - Edge 1 315 - Node A 320 - Edge 2 325 - Node B 330 - Edge 3 335 - Node D 340 - Edge 6 345 - Node E 350 - Edge 7 355 - VM2 360

[0024] These two above-mentioned paths are then converted into a distributed representation using a pre-trained embedding matrix. Next, an attention mechanism aggregates them into a single vector. This vector can then be fed into a machine learning model trained to identify changes related to vulnerabilities. The choice of classifier is open and ranges from a simple linear classifier to a complex highly non-linear classifier. Furthermore, it can be decided to fix this representation or to fine-tune it to enable development. In some examples, the dataset can be split into three sets, namely a training set used to fit the model parameters, a validation set used for hyperparameter tuning, and a test set used to objectively evaluate the performance of the classifier.

[0025] Figure 4 shows a workflow according to some embodiments. Workflow 400 begins by receiving a source code commit at step 410. The source code commit may include at least one source code change to a source code repository and a natural language message describing the at least one source code change. Workflow 400 can continue by generating a message vector at step 420. The message vector may be based on the message part of the source code commit. Workflow 400 can continue by generating a commit vector at 430. The commit vector may be based on at least one source code change. In one example, the commit vector may be generated by processing at least one source code change by a neural network model.

[0026] Workflow 400 can continue by generating an architecture recognition commit vector at 440. The architecture recognition commit vector may be based on at least one source code change and an annotated architecture model of the source code repository. In one embodiment, the annotated architecture model is generated by first analyzing the source code repository to extract the architecture model and then annotating the architecture model to indicate or highlight architecture elements that may have potential security risks or vulnerabilities. Annotating may be automated by a security analyzer tool, or done manually by a security architect, or may be done by a combination of the two. In another embodiment, the annotated architecture model is retrieved from memory for use in generating the architecture recognition commit vector, updated to account for source code changes, and then stored back in memory for future use.

[0027] Workflow 400 can continue by processing the commit vector, message vector, and architecture recognition commit vector with a neural network model at 450 to determine the likelihood that a source code change creates a potential security risk or vulnerability. The output of the neural network model may be a value between 0 and 1, where 0 indicates a very low likelihood and 1 indicates a high likelihood that a software code change creates a potential security risk or vulnerability. In one embodiment, commits that include source code changes with a high likelihood of creating a security risk may be flagged for further analysis. In another embodiment, commits that include source code changes with a high likelihood of creating a security risk may not be deployed on the target system.

[0028] FIG. 5 shows the hardware of a dedicated computer configured to implement a code commit registry according to some embodiments. Specifically, computer system 501 includes a processor 502 in electronic communication with a non-transitory computer-readable storage medium including a database 503. This computer-readable storage medium stores thereon code 505 corresponding to a registry engine. Code 504 corresponds to a hash. The code may be configured to reference data stored in a database of a non-transitory computer-readable storage medium, for example, locally or at a remote database server. Software servers together may form a cluster or logical network of computer systems programmed with software programs that communicate with each other and operate together to process requests.

[0029] The embodiments described herein may provide one or more advantages. One potential benefit is enhanced collaboration with expected customers and vendors. That is, customers can freely engage vendors to develop beneficial add-ons, and thus it can be trusted that the beneficial add-ons operate seamlessly with the basic framework.

[0030] In view of the implementations of the subject matter described above, this application discloses the following list of examples. However, one feature of an isolated example, or a combination, and in some cases two or more features of the above examples combined with one or more features of one or more additional examples are further examples that also fall within the scope of the disclosure of this application.

[0031] An exemplary computer system 600 is shown in FIG. 6. The computer system 610 includes a bus 605 or other communication mechanism for communicating information, and a processor 601 coupled to the bus 605 for processing information. The computer system 610 also includes a memory 602 coupled to the bus 605 for storing information and instructions executed by the processor 601, including, for example, information and instructions for performing the techniques described above. This memory can also be used to store variables or other intermediate information during the execution of instructions executed by the processor 601. Possible implementations of this memory can include, but are not limited to, random access memory (RAM), read-only memory (ROM), or both. A storage device 603 is also provided for storing information and instructions. Common forms of storage devices include, for example, hard drives, magnetic disks, optical disks, CD-ROMs, DVDs, flash memories, USB memory cards, or other media readable by a computer. The storage device 603 may include, for example, source code, binary code, or software files for performing the above techniques. Both the storage device and the memory are examples of computer-readable media.

[0032] The computer system 610 can be coupled via the bus 605 to a display 612, such as a light emitting diode (LED) or liquid crystal display (LCD), for displaying information to a computer user. Input devices 611, such as a keyboard and / or a mouse, are coupled to the bus 605 for communicating information and command selections from the user to the processor 601. Combinations of these components enable a user to communicate with the system. In some systems, the bus 605 may be split into multiple dedicated buses.

[0033] Computer system 610 also includes a network interface 604 coupled to bus 605. Network interface 604 can provide bi-directional data communication between computer system 610 and local network 620. Network interface 604 may be, for example, a digital subscriber line (DSL) or modem that provides a data communication connection over a telephone line. Another example of a network interface is a local area network (LAN) card that provides a data communication connection to a compatible LAN. A wireless link is another example. In any such implementation, network interface 604 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0034] Computer system 610 can transmit and receive information, including messages or other interface actions, over local network 620, an intranet, or Internet 630 via network interface 604. In the case of a local network, computer system 610 can communicate with a plurality of other computer machines, such as server 615. Thus, the server computer systems represented by computer system 610 and server 615 may form a cloud computing network, and the cloud computing network may be programmed with the processes described herein. In the example of the Internet, software components or services may be on multiple different computer systems 610 or servers 631-635 across the network. The processes described above may be implemented, for example, on one or more servers. Server 631 may transmit an action or message from one component to a component on computer system 610 via Internet 630, local network 620, and network interface 604. The software components and processes described above may be implemented on any computer system and may, for example, send and / or receive information across the network.

[0035] The above description shows various embodiments of the present invention, along with examples of how aspects of the invention may be implemented. The above examples and embodiments should not be considered the only embodiments, but are presented to illustrate the flexibility and advantages of the present invention as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations, and equivalents will be apparent to those skilled in the art and may be employed without departing from the spirit and scope of the invention as defined by the claims.

[0036] Further examples Each of the following non-limiting features of the following examples may be independent or may be combined in various permutations or combinations with one or more of the other features of the following examples. In various embodiments, the present disclosure may be implemented as a processor or a method.

[0037] In some embodiments, the present disclosure is to receive a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message describing the at least one source code change, to generate a message vector based on the message, to generate a commit vector based on the at least one source code change, to generate an architecture recognition commit vector based on the at least one source code change and an annotated architecture model of the source code repository, and to process the message vector, the commit vector, and the architecture recognition commit vector by a neural network model to determine the likelihood that the at least one source code change creates a potential security risk.

[0038] In one embodiment, generating the message vector includes processing the message by a natural language model.

[0039] In one embodiment, generating the commit vector includes processing the at least one source code change by a second neural network model.

[0040] In one embodiment, a proof result is associated with a portion of a software artifact.

[0041] In one embodiment, generating an architecture recognition commit vector includes identifying architecture elements in an annotated architecture model associated with at least one source code change, extracting at least one path in the annotated architecture model through the architecture elements, and aggregating at least one of the extracted paths to generate an architecture recognition commit vector.

[0042] In one embodiment, generating an architecture recognition commit vector further includes extracting a source code tree from a source code repository, generating an architecture model from the source code tree, and annotating the architecture model to identify at least one architecture element in the architecture model that is a potential security risk.

[0043] In one embodiment, the extracted path connects two terminal nodes in the annotated architecture model.

[0044] In one embodiment, the method further includes updating the annotated architecture model according to at least one source code change.

[0045] In one embodiment, aggregating at least one of the extracted paths includes encoding each path of the at least one extracted path to generate a corresponding vector representation.

[0046] In one embodiment, aggregating at least one of the extracted paths further includes aggregating the vector representations corresponding to the at least one extracted path.

[0047] In some embodiments, the present disclosure includes a system comprising one or more processors and a non-transitory computer-readable medium storing a program executable by the one or more processors, the program including instructions for receiving a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message explaining the at least one source code change; generating a message vector based on the message; generating a commit vector based on the at least one source code change; generating an architecture recognition commit vector based on the at least one source code change and an annotated architecture model of the source code repository; and processing the message vector, the commit vector, and the architecture recognition commit vector by a neural network model to determine whether the at least one source code change creates a potential security risk.

[0048] In some embodiments, the present disclosure includes a system comprising one or more processors and a non-transitory computer-readable medium storing a program executable by the one or more processors, the program including instructions for receiving a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message explaining the at least one source code change; generating a message vector based on the message; generating a commit vector based on the at least one source code change; generating an architecture recognition commit vector based on the at least one source code change and an annotated architecture model of the source code repository; and processing the message vector, the commit vector, and the architecture recognition commit vector by a neural network model to determine whether the at least one source code change creates a potential security risk.

Explanation of Signs

[0049] 100 Workflow 105 Source Code 110 Commit 115 Message 120 Natural Language Processing (NLP) Model 125 Message Vector 130 Source Code Change 135 commit2vec 140 Commit Vector 145 Annotated Architecture Model 155 Architecture Recognition Commit Vector 160 Machine Learning Model 200 Workflow 205 Source Code 210 Architecture Model Extractor 215 Architecture Model 220 Architecture Model Security Analyzer 225 Annotated Architecture Model 230 Code Change 240 Contextualizer 245 Contextualized Architecture Model 250 Vectorizer 255 Architecture Recognition Commit Vector 260 Machine Learning Pipeline 290 Security Architect 300 Contextualized Architecture Model 310 Cloud 315 Edge 1 320 Node A 325 Edge 2 330 Node B 335 Edge 3 340 Node D 345 Edge 6 350 Node E 355 Edge 7 360 VM2 365 Edge 4 370 Node C 375 Edge 5 380 VM1 501 Computer System 502 Processor 503 Database 504 Code 505 Code 600 Computer System 601 Processor 602 Memory 603 Storage Device 604 Network Interface 605 Bus 610 Computer System 611 Input Device 612 Display 615 Server 620 Local Network 630 Internet 631 Server 632 Server 633 Server 634 Server 635 Server

Claims

1. 1. A method comprising: receiving a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message describing the at least one source code change; generating a message vector based on the messages; generating a commit vector based on the at least one source code change; generating an architecture-aware commit vector based on the at least one source code change and an annotated architecture model of the source code repository; processing the message vector, the commit vector, and the architecture-aware commit vector through a neural network model to determine a likelihood that the at least one source code change creates a potential security risk; A method comprising:

2. The method of claim 1 , wherein generating the message vector comprises processing the message through a natural language model.

3. 2. The method of claim 1, wherein generating the commit vector comprises processing the at least one source code change through a second neural network model.

4. generating the architecture-aware commit vector comprises: identifying an architecture element in the annotated architecture model that is associated with the at least one source code change; extracting at least one path in the annotated architecture model that passes through the architecture element; aggregating the at least one extracted path to generate the architecture-aware commit vector; 2. The method of claim 1, comprising:

5. generating the architecture-aware commit vector comprises: extracting a source code tree from the source code repository; generating an architecture model from the source code tree; annotating the architecture model to identify at least one architectural element in the architecture model that is a potential security risk; 5. The method of claim 4, further comprising:

6. The method of claim 4 , wherein the extracted path connects two terminal nodes in the annotated architectural model.

7. The method of claim 5 , further comprising updating the annotated architecture model according to the at least one source code change.

8. 5. The method of claim 4, wherein aggregating the at least one extracted path comprises encoding each leg of the at least one extracted path to generate a corresponding vector representation.

9. The method of claim 8 , wherein aggregating the at least one extracted path further comprises aggregating the vector representations corresponding to the at least one extracted path.

10. 1. A system comprising: one or more processors; and a non-transitory computer-readable medium having stored thereon a program executable by said one or more processors, said program comprising: receiving a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message describing the at least one source code change; generating a message vector based on the messages; generating a commit vector based on the at least one source code change; generating an architecture-aware commit vector based on the at least one source code change and an annotated architecture model of the source code repository; processing the message vector, the commit vector, and the architecture-aware commit vector through a neural network model to determine a likelihood that the at least one source code change creates a potential security risk; A system including a set of instructions for performing the steps of the method.

11. generating the architecture-aware commit vector, identifying an architecture element in the annotated architecture model that is associated with the at least one source code change; extracting at least one path in the annotated architecture model that passes through the architecture element; aggregating the at least one extracted path to generate the architecture-aware commit vector; and The system of claim 10, comprising:

12. generating the architecture-aware commit vector, extracting a source code tree from the source code repository; generating an architecture model from the source code tree; annotating the architecture model to identify at least one architectural element in the architecture model that is a potential security risk; The system of claim 11 further comprising:

13. The system of claim 12 , wherein the program further comprises a set of instructions for updating the annotated architecture model according to the at least one source code change.

14. 12. The system of claim 11, wherein aggregating the at least one extracted path comprises encoding each path of the at least one extracted path to generate a corresponding vector representation.

15. The system of claim 14 , wherein aggregating the at least one extracted path further comprises aggregating the vector representations corresponding to the at least one extracted path.

16. A non-transitory computer readable medium having stored thereon a program executable by one or more processors, the program comprising: receiving a source code commit, the source code commit including at least one source code change to a source code repository and a natural language message describing the at least one source code change; generating a message vector based on the messages; generating a commit vector based on the at least one source code change; generating an architecture-aware commit vector based on the at least one source code change and an annotated architecture model of the source code repository; processing the message vector, the commit vector, and the architecture-aware commit vector through a neural network model to determine a likelihood that the at least one source code change creates a potential security risk; A non-transitory computer-readable medium comprising a set of instructions for performing the steps of:

17. generating the architecture-aware commit vector, identifying an architecture element in the annotated architecture model that is associated with the at least one source code change; extracting at least one path in the annotated architecture model that passes through the architecture element; aggregating the at least one extracted path to generate the architecture-aware commit vector; and 20. The non-transitory computer-readable medium of claim 16, comprising:

18. generating the architecture-aware commit vector, extracting a source code tree from the source code repository; generating an architecture model from the source code tree; annotating the architecture model to identify at least one architectural element in the architecture model that is a potential security risk; 20. The non-transitory computer-readable medium of claim 17, further comprising:

19. 20. The non-transitory computer-readable medium of claim 18, further comprising updating the architecture model according to the at least one source code change.

20. 20. The non-transitory computer-readable medium of claim 17, wherein aggregating the at least one extracted path comprises encoding each path of the at least one extracted path to generate a corresponding vector representation.