Source code vulnerability detection method and system based on multi-modal feature fusion

Through the multimodal feature fusion method, combined with semantics, syntax, loop invariants and state transition diagram features, the problems of low accuracy and high false alarm rate in source code vulnerability detection in the existing technology are solved, and more accurate vulnerability detection and code change tracking are achieved.

CN120654238APending Publication Date: 2025-09-16STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510495895.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing technology is too simplistic in feature representation in source code vulnerability detection, unable to effectively identify loop invariants and program state migration rules, and unable to track code changes, resulting in low vulnerability detection accuracy.

Method used

A multimodal feature fusion method is used to extract semantic, grammatical, loop invariant and state transition diagram features, which are then combined with difference features to perform local vulnerability detection and correct the global detection results.

Benefits of technology

It improves the accuracy and adaptability of vulnerability detection, reduces the false positive rate and missed negative rate, and can effectively track the impact of code changes on vulnerability detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654238A_ABST
    Figure CN120654238A_ABST
Patent Text Reader

Abstract

The invention relates to a source code vulnerability detection method and system based on multi-modal feature fusion. The method comprises the steps of obtaining semantic features, grammar features, loop invariant features and state transition diagram features based on a source code to be detected; obtaining logic features and structural features based on the grammatical features, the state transition diagram features, the semantic features and the loop invariant features, and generating a vulnerability probability matrix of the source code to be detected based on the structural features and the logic features; obtaining difference characteristics of the code segments before and after modification of the modified code segments, and generating a vulnerability detection result based on the difference characteristics; generating a source code vulnerability detection result based on the vulnerability detection result and the vulnerability probability matrix; the system provided by the invention is used for implementing the method. Compared with the prior art, more accurate source code vulnerability detection and identification are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a source code vulnerability detection method and system based on multimodal feature fusion. Background Art

[0002] With the rapid development of Internet technology, cyberspace security incidents occur frequently, posing a serious threat to users' information and property security. Among them, with the exponential growth of software system complexity, code vulnerabilities have become the core threat to network security. Traditional vulnerability detection technologies include the following methods: Static analysis: Analyzing code logic through abstract syntax trees (AST) and control flow graphs (CFG), but there is a high false alarm rate problem, making it difficult to identify deep semantic vulnerabilities; Dynamic analysis: Based on fuzzing or symbolic execution, although it can discover runtime defects, it is limited by path explosion and execution efficiency; Introducing deep learning or machine learning to analyze based on code features, such as the Chinese patent application "CN118395448A", which provides a training method for a code vulnerability detection model, using the GRU model to extract long-term semantic information of the code, which can mine the complex correlation between multiple features. On the basis of the GIN model extracting the grammatical information and code structure features of the code, the long-term semantic information and grammatical information are fused in a channel cascade manner, paying full attention to the correlation between the code syntax and semantic information, and making up for the deviation in the feature quantity. The spliced ​​training feature vector includes richer structural and textual features, and code vulnerability detection is performed based on the obtained training feature vector. The above methods exist:

[0003] 1) The feature representation is too simplistic, relying only on syntax or control flow features, ignoring the dynamic migration of program states, and lacking information on key vulnerabilities such as loop invariants and variable dependencies;

[0004] 2) In the actual code detection process, the detected code vulnerability may be the previous vulnerability code segment that has been repaired. Traditional code detection methods cannot effectively track code changes and identify the causal relationship between code changes and vulnerability introduction, thus affecting the accuracy of code vulnerability detection.

[0005] Therefore, a technical problem that needs to be solved is to provide a method that can correct vulnerability detection results based on multimodal feature data and track code changes, thereby achieving more accurate source code vulnerability detection. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a source code vulnerability detection method and system based on multimodal feature fusion. The method performs source code segment vulnerability detection by acquiring multimodal data including semantics, syntax, loop invariants and state transition diagrams, intercepts modified code segments for local vulnerability detection, and corrects the detection results of the entire source code segment based on the results of the local vulnerability detection, thereby achieving more accurate source code vulnerability detection.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] According to a first aspect of the present invention, a method for detecting source code vulnerabilities based on multimodal feature fusion is provided, comprising:

[0009] Extract input data based on the source code to be tested and preprocess it to obtain semantic features, grammatical features, loop invariant features and state transition diagram features;

[0010] The grammatical features are integrated with the state transition diagram features to obtain structural features, the semantic features are combined with the loop invariant features to obtain logical features, and the vulnerability probability matrix of the source code to be detected is generated based on the structural features and the logical features;

[0011] Intercepting a code segment of the source code to be detected in which code modification has occurred, obtaining difference features based on the code segment before and after the modification, and generating a vulnerability detection result based on the difference features;

[0012] Based on the vulnerability detection results and the vulnerability probability matrix, a source code vulnerability detection result is generated, wherein the detection result includes the vulnerability location and the danger level.

[0013] As a preferred technical solution, the loop invariant feature is a condition or expression that always holds true in each iteration, and its generation method includes:

[0014] Identify loop entry, exit, and iteration path in the source code to be tested by data flow analysis to obtain the loop body, and mark the variables inside and outside the loop body;

[0015] Using interval analysis to determine the value ranges of variables inside and outside the loop body, and generating candidate loop invariants in combination with loop conditions;

[0016] Traversing all loop execution paths, if the candidate loop invariant is established in all loop execution paths, then determining that the candidate loop invariant is a loop invariant; otherwise, eliminating the corresponding candidate loop invariant;

[0017] A loop invariant feature is generated based on the loop invariant.

[0018] As a preferred technical solution, the method for obtaining the state transition diagram characteristics includes:

[0019] Segmenting the source code to be detected according to the function structure, determining the logical relationships between and within each code segment, and constructing a control flow graph based on the logical relationships;

[0020] Based on the control flow graph, jump relationships between state nodes and special nodes, or inclusion relationships between state nodes and special nodes, are obtained, and edges are constructed according to the jump relationships and inclusion relationships; the state nodes are functions, and their attributes include ID, variables, and variable operations; the special nodes are loop structures, and their attributes include loop start conditions and loop end conditions; and the attributes of the edges include relationship types and variable operations;

[0021] A state transition graph is constructed according to the state nodes, special nodes and edges, and state transition graph features are extracted based on the state transition graph.

[0022] As an optimal technical solution, the method for obtaining the structural features is: the shape of the grammatical features is code length × 1, the shape of the state transition diagram features is the number of graph nodes × 1, and the transpose of the grammatical features and the state transition diagram features is fused into a structural feature with a shape size of: code length × number of graph nodes.

[0023] As an optimal technical solution, the method for obtaining the logical feature is: the shape of the loop invariant feature is the number of loop invariants × 1, the shape of the semantic feature is the code length × 1, and the transpose of the semantic feature and the loop invariant feature is fused into a logical feature of code length × the number of loop invariants.

[0024] As a preferred technical solution, the method for obtaining the vulnerability probability matrix and the corresponding vulnerability types includes:

[0025] After mapping the logical features and structural features to the same dimension, dynamic weighted fusion is performed, and the expression is:

[0026] g=σ[W g (F1+F2)],

[0027] F3=g⊙F1+(1-g)⊙F2,

[0028] Where g represents the gate vector; W g represents the learnable weight matrix; F1 represents the logical feature; F2 represents the structural feature; σ[·] represents the Sigmoid function calculation; F3 represents the fusion feature;

[0029] The fusion features are processed using a Softmax function to obtain a vulnerability probability matrix of code length × vulnerability type; the vulnerability types include logic errors, memory management errors, buffer overflows, attention vulnerabilities, and no vulnerabilities; the (i, j) element in the vulnerability probability matrix represents the probability that the i-th row of code is the j-th vulnerability type.

[0030] As a preferred technical solution, the method for obtaining the vulnerability detection results and the corresponding vulnerability types includes:

[0031] Comparing the modified code segment before and after the modification, extracting the differences and encoding them to obtain difference features, wherein the difference features include semantic difference features, control flow difference features, modified variable features, and loop invariant difference features;

[0032] Each type of difference feature is projected to the same dimension and fused to obtain a fused difference feature, and the Softmax function is used to obtain the vulnerability detection result of the fused difference feature.

[0033] As a preferred technical solution, the method for obtaining the difference characteristics includes:

[0034] Extract the abstract syntax trees of the code segments before and after the modification and align the syntax units of the two codes. Based on the alignment results, mark the syntax units that have been added, deleted, or re-edited. Count the frequency of addition, deletion, or re-editing. Encode based on the marking results and frequency to obtain semantic difference features.

[0035] Obtain the control flow graphs of the code segments before and after modification, identify new paths, deleted paths, and logic correction paths based on the control flow graphs, calculate the lengths of the corresponding modified paths, and normalize the calculated lengths to obtain control flow difference features;

[0036] Obtain variables with changed values, types, ranges, and dependencies in the code segments before and after modification, and perform one-hot encoding based on the modified types to obtain the modified variable features.

[0037] Obtain the loop invariants that are corrected, destroyed, or newly added in the code segments before and after modification, and perform one-hot encoding based on the modification type to obtain the loop invariant difference features.

[0038] As a preferred technical solution, the method for obtaining the hazard level includes:

[0039] Correcting the vulnerability probability matrix based on the vulnerability detection results;

[0040] When the vulnerability detection result is the same as the detection result of the corresponding position in the vulnerability probability matrix, the frequency of vulnerability occurrence in each code segment is obtained based on the vulnerability probability matrix, and a vulnerability distribution density map is drawn based on the frequency. The code segment with high density is marked as high risk, and the code segment with low density is marked as low risk.

[0041] When the vulnerability detection result is different from the detection result of the corresponding position in the vulnerability probability matrix, the result of the corresponding position in the vulnerability probability matrix is ​​corrected according to the vulnerability detection result so that it is the same as the vulnerability detection result, and the corrected code vulnerability matrix is ​​obtained. The frequency of vulnerability occurrence in each code segment is obtained based on the corrected code vulnerability matrix, and a vulnerability distribution density map is drawn based on the frequency. The code segment with high density is marked as high risk, and the code segment with low density is marked as low risk.

[0042] According to a second aspect of the present invention, a source code vulnerability detection system based on multimodal feature fusion is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the method when executing the program.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1) When performing source code vulnerability detection in the prior art, the entire source code to be detected is often scanned and detected repeatedly, and it is unable to identify vulnerability transmission effects caused by local code modifications or residual unfixed vulnerabilities in local code segments. Therefore, in order to solve the problems existing in the prior art, the present invention performs global vulnerability detection and local modified code vulnerability detection respectively, and corrects the global code detection results based on the results of the local vulnerability detection, thereby making up for the shortcomings of the global vulnerability detection process in tracking code changes. This makes the method provided by the present invention more adaptable when dealing with code version changes and has higher accuracy in detecting vulnerabilities.

[0045] 2) Based on traditional vulnerability detection, the present invention introduces loop invariant features, uses loop invariant features to supplement semantic features, and enhances the ability of the acquired logical features to identify code logic; and fuses grammatical features and state transition diagram features to obtain structural features. This structural feature can accurately capture problems such as abnormal jumps in the code execution path, making the method provided by the present invention more accurate in complex code vulnerability detection, and can effectively reduce the false alarm rate of vulnerability detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Flow chart of the method of this invention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0048] Example 1

[0049] In order to solve the problem that in the actual code detection process, the detected code vulnerability may be the previous vulnerability code segment that has been repaired, and the traditional code detection method cannot effectively track the code changes and identify the causal relationship between the code changes and the introduction of vulnerabilities, thereby affecting the accuracy of code vulnerability detection, resulting in high omission and false alarm rates of code detection, the present invention proposes a method for global and local detection of code vulnerabilities, which strengthens the source code feature representation by fusing multimodal features such as semantics, syntax, loop invariants and state transition diagrams, and locally detects the vulnerability of the code segment with code modifications, and corrects the global detection results based on the local detection results.

[0050] The detailed process of this method is as follows Figure 1 As shown, including:

[0051] S1. Extract input data based on the source code to be detected and preprocess it to obtain semantic features, grammatical features, loop invariant features and state transition diagram features.

[0052] S11. Preprocessing: The source code segment to be detected is standardized, comments are deleted, and variable names and function names are mapped to unified symbols to obtain a standardized code segment.

[0053] S12. Perform lexical analysis and construct an abstract syntax tree (AST) on the preprocessed source code segment to be tested, extract variable declarations, function call relationships and scope information, map the AST nodes into vectors through graph embedding technology, and generate a semantic feature matrix with a dimension of code length × 1.

[0054] S13. parse the grammatical structure of the code based on the preprocessed source code segment to be detected, extract keywords (such as if, for), operators and separator sequences, convert them into one-hot encoding (One-hot) form, and obtain a grammatical feature vector with a shape of code length × 1.

[0055] S14. Extract loop invariant features:

[0056] Specifically, the loop invariant is a condition or expression that remains true in each iteration. The generation method includes:

[0057] S141. Identify loop entry, exit, and iteration path in the source code to be detected through data flow analysis to obtain a loop body, and mark variables inside and outside the loop body.

[0058] S142. Use interval analysis to determine the value ranges of variables inside and outside the loop body, and generate candidate loop invariants based on the loop conditions.

[0059] S143. Traverse all loop execution paths. If a candidate loop invariant holds true in all loop execution paths, determine that the candidate loop invariant is a loop invariant. Otherwise, eliminate the corresponding candidate loop invariant.

[0060] S144. Generate a loop invariant feature based on the loop invariant.

[0061] S15. Extract state transition graph features:

[0062] S151 , segmenting the source code to be detected according to the function structure, determining the logical relationships between and within each code segment, and constructing a control flow graph based on the logical relationships.

[0063] S152. Based on the control flow graph, the jump relationship between state nodes and special nodes, or the inclusion relationship between state nodes and special nodes, is obtained, and edges are constructed according to the jump relationship and the inclusion relationship; the state node is a function, and its attributes include ID, variables, and variable operations; the special node is a loop structure, and its attributes include loop start conditions and loop end conditions; the attributes of the edge include relationship type and variable operation.

[0064] S153: construct a state transition graph based on the state nodes, special nodes, and edges, and extract state transition graph features based on the state transition graph.

[0065] S2. Structural features are obtained by fusing syntactic features with state transition diagram features, logical features are obtained by concatenating semantic features with loop invariant features, and a vulnerability probability matrix of the source code to be detected is generated based on the structural features and logical features.

[0066] Specifically, semantic features express the functional and logical intent of the code, such as whether a code segment implements data sorting or file reading. Loop invariants, on the other hand, are properties or conditions that remain constant during loop execution, reflecting the essential characteristics of the loop. Therefore, integrating the two has the following advantages:

[0067] 1) A more comprehensive and precise understanding of the functionality and behavior of loop code. For example, when analyzing a loop that calculates the sum of array elements, the semantic features indicate a summation operation, while the loop invariant indicates that each iteration accumulates one element in the array. This fusion allows for a clearer understanding of the complete semantics of the loop code.

[0068] 2) It can more effectively detect errors in loops, such as invariants being accidentally broken or the semantics of loop execution not matching expected functionality. During debugging, by combining semantics with loop invariants, loop-related problems can be located more quickly, improving debugging efficiency.

[0069] Similarly, syntactic features reflect the structural rules of the code, such as the composition of statements and variable declaration methods, and focus on the static structure of the code. State transition diagram features dynamically display the transition relationship between different states of the program and the conditions for state transitions. Therefore, combining the two has the following advantages:

[0070] 1) It allows users to understand both the structure of the code and its dynamic behavior at runtime, thus forming a more complete understanding of the program. For example, when analyzing a finite state machine program, the syntax features describe the definition of the state machine and the syntax of related operations, while the state transition diagram features show the state changes of the state machine under different inputs. After integration, the program's working mechanism can be fully understood.

[0071] 2) It can include as much logic as possible from the existing code, so that when performing code vulnerability detection, it can more accurately identify vulnerabilities in the program's structure and behavior, reducing the probability of errors.

[0072] Based on the above reasons, the present invention performs feature-level fusion of semantic features and loop invariants, and performs feature-level fusion of grammatical features and state transition diagram features, so as to reduce the false alarm rate and missed alarm rate of code vulnerability detection.

[0073] S21. Obtain structural features: the shape of the grammatical feature is code length × 1, and the shape of the state transition diagram feature is the number of graph nodes × 1. The transpose of the grammatical feature and the state transition diagram feature is fused into a structural feature with a shape size of: code length × number of graph nodes. For example, when the code length is 100 lines and the number of loop invariants is 20, the logical feature dimension is 100×20.

[0074] S22. The shape of the loop invariant feature is the number of loop invariants × 1, and the shape of the semantic feature is the code length × 1. The transpose of the semantic feature and the loop invariant feature is fused into a logical feature of code length × the number of loop invariants.

[0075] S23. After mapping the logical features and structural features to the same dimension, dynamic weighted fusion is performed. The expression is:

[0076] g=σ[W g (F1+F2)],

[0077] F3=g⊙F1+(1-g)⊙E2,

[0078] Where g represents the gate vector; W grepresents the learnable weight matrix; F1 represents the logical feature; F2 represents the structural feature; σ[·] represents the Sigmoid function calculation; F3 represents the fusion feature;

[0079] S24. Use the Softmax function to process the fused features to obtain a vulnerability probability matrix of code length × vulnerability type; vulnerability types include logic errors, memory management errors, buffer overflows, attention vulnerabilities, and no vulnerabilities; the (i, j) element in the vulnerability probability matrix represents the probability that the i-th line of code is the j-th type of vulnerability. When the probability is greater than 50%, it indicates that the vulnerability exists.

[0080] S3. Detect code segments in the source code where code modifications exist, obtain differential features based on the code segments before and after the modification, and generate vulnerability detection results based on the differential features.

[0081] S31. Compare the code segments before and after the modification of the modified code segment, extract the differences, and encode them to obtain difference features. The difference features include semantic difference features, control flow difference features, modified variable features, and loop invariant difference features. Specifically, the step of obtaining the difference features includes:

[0082] S311. Extract the abstract syntax trees of the code segments before and after modification and align the syntax units of the codes at both ends. Based on the alignment results, mark the syntax units that have been added, deleted, or re-edited, and count the frequency of addition, deletion, or re-editing. Encode based on the marking results and frequency to obtain semantic difference features.

[0083] S312. Obtain control flow graphs of the code segments before and after modification, identify newly added paths, deleted paths, and logical correction paths based on the control flow graphs, calculate the lengths of the corresponding modified paths, normalize the calculated lengths, and obtain control flow difference features.

[0084] S313. Obtain variables with assignment changes, type changes, value range changes, and dependency changes in the code segments before and after modification, and perform one-hot encoding based on the modification types to obtain modified variable features.

[0085] S314. Obtain loop invariants that are corrected, destroyed, or newly added in the code segments before and after modification, and perform one-hot encoding based on the modification type to obtain loop invariant difference features.

[0086] S32. Project each type of difference feature to the same dimension and fuse them to obtain a fused difference feature. Use the Softmax function to obtain the vulnerability detection result on the fused difference feature. The vulnerability detection result includes the vulnerability classification and the probability of the corresponding vulnerability category. When the vulnerability probability is greater than 50%, it indicates that the code segment has a vulnerability of the corresponding category.

[0087] S4. Vulnerability detection results and vulnerability probability matrix, generate source code vulnerability detection results, the detection results include vulnerability location and danger level.

[0088] In this step, the vulnerability probability matrix needs to be corrected based on the vulnerability detection result. In detail, when the vulnerability detection result is the same as the detection result of the corresponding position in the vulnerability probability matrix, no correction is performed and step S43 is executed. When the vulnerability detection result is different from the detection result of the corresponding position in the vulnerability probability matrix, steps S42 and S43 are executed sequentially.

[0089] S42. Correct the result of the corresponding position in the vulnerability probability matrix according to the vulnerability detection result so that it is the same as the vulnerability detection result, and obtain a corrected code vulnerability matrix.

[0090] S43. Based on the vulnerability probability matrix (or the corrected code vulnerability matrix), the frequency of vulnerabilities in each code segment is obtained, and a vulnerability distribution density map is drawn based on the frequency. The danger level of code segments with high density (density greater than 80%) is marked as high risk, the danger level of code segments with low density (density less than 20%) is marked as low risk, and the danger level of code segments with density values ​​between 20% and 80% is medium risk.

[0091] Example 2

[0092] The above is an introduction to a method embodiment. This embodiment provides a source code vulnerability detection system based on multimodal feature fusion. Technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working process described can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0093] The source code vulnerability detection system provided in this embodiment includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0094] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0095] The processing unit performs the various methods and processes described above, such as methods S1 to S4. For example, in some embodiments, methods S1 to S4 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S4 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to execute methods S1 to S4 by any other appropriate means (for example, by means of firmware).

[0096] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0097] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0098] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0099] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A source code vulnerability detection method based on multimodal feature fusion, characterized in that: include: Extract input data based on the source code to be tested and preprocess it to obtain semantic features, grammatical features, loop invariant features and state transition diagram features; The grammatical features are integrated with the state transition diagram features to obtain structural features, the semantic features are combined with the loop invariant features to obtain logical features, and the vulnerability probability matrix of the source code to be detected is generated based on the structural features and the logical features; Intercepting a code segment of the source code to be detected in which code modification has occurred, obtaining difference features based on the code segment before and after the modification, and generating a vulnerability detection result based on the difference features; Based on the vulnerability detection results and the vulnerability probability matrix, a source code vulnerability detection result is generated, wherein the detection result includes the vulnerability location and the danger level.

2. A source code vulnerability detection method based on multimodal feature fusion according to claim 1, characterized in that: The loop invariant feature is a condition or expression that always holds true in each iteration, and its generation method includes: Identify loop entry, exit, and iteration path in the source code to be tested by data flow analysis to obtain the loop body, and mark the variables inside and outside the loop body; Using interval analysis to determine the value ranges of variables inside and outside the loop body, and generating candidate loop invariants in combination with loop conditions; Traversing all loop execution paths, if the candidate loop invariant is established in all loop execution paths, then determining that the candidate loop invariant is a loop invariant; otherwise, eliminating the corresponding candidate loop invariant; A loop invariant feature is generated based on the loop invariant.

3. A source code vulnerability detection method based on multimodal feature fusion according to claim 2, characterized in that: The method for obtaining the state transition diagram characteristics includes: Segmenting the source code to be detected according to the function structure, determining the logical relationships between and within each code segment, and constructing a control flow graph based on the logical relationships; Based on the control flow graph, jump relationships between state nodes and special nodes, or inclusion relationships between state nodes and special nodes, are obtained, and edges are constructed according to the jump relationships and inclusion relationships; the state nodes are functions, and their attributes include ID, variables, and variable operations; the special nodes are loop structures, and their attributes include loop start conditions and loop end conditions; and the attributes of the edges include relationship types and variable operations; A state transition graph is constructed according to the state nodes, special nodes and edges, and state transition graph features are extracted based on the state transition graph.

4. A source code vulnerability detection method based on multimodal feature fusion according to claim 3, characterized in that: The method for obtaining the structural features is: the shape of the grammatical features is code length × 1, the shape of the state transition diagram features is the number of diagram nodes × 1, and the transpose of the grammatical features and the state transition diagram features is fused into a structural feature with a shape size of: code length × number of diagram nodes.

5. The source code vulnerability detection method based on multimodal feature fusion according to claim 4 is characterized in that: The method for obtaining the logical feature is: the shape of the loop invariant feature is the number of loop invariants × 1, the shape of the semantic feature is the code length × 1, and the transpose of the semantic feature and the loop invariant feature is fused into a logical feature of code length × the number of loop invariants.

6. The source code vulnerability detection method based on multimodal feature fusion according to claim 5 is characterized in that: The method for obtaining the vulnerability probability matrix and the corresponding vulnerability types includes: After mapping the logical features and structural features to the same dimension, dynamic weighted fusion is performed, and the expression is: g=σ[W g ·(F1+F2)], F3=g⊙F1+(1-g)⊙F2, Where g represents the gate vector; W g represents the learnable weight matrix; F1 represents the logical feature; F2 represents the structural feature; σ[·] represents the Sigmoid function calculation; F3 represents the fusion feature; The fusion features are processed using a Softmax function to obtain a vulnerability probability matrix of code length × vulnerability type; the vulnerability types include logic errors, memory management errors, buffer overflows, attention vulnerabilities, and no vulnerabilities; the (i, j) element in the vulnerability probability matrix represents the probability that the i-th row of code is the j-th vulnerability type.

7. The method for detecting source code vulnerabilities based on multimodal feature fusion according to claim 1, characterized in that: Methods for obtaining the vulnerability detection results and corresponding vulnerability types include: Comparing the modified code segment before and after the modification, extracting the differences and encoding them to obtain difference features, wherein the difference features include semantic difference features, control flow difference features, modified variable features, and loop invariant difference features; Each type of difference feature is projected to the same dimension and fused to obtain a fused difference feature, and the Softmax function is used to obtain the vulnerability detection result of the fused difference feature.

8. The method for detecting source code vulnerabilities based on multimodal feature fusion according to claim 7, characterized in that: The method for obtaining the difference feature includes: Extract the abstract syntax trees of the code segments before and after the modification and align the syntax units of the two codes. Based on the alignment results, mark the syntax units that have been added, deleted, or re-edited. Count the frequency of addition, deletion, or re-editing. Encode based on the marking results and frequency to obtain semantic difference features. Obtain the control flow graphs of the code segments before and after modification, identify new paths, deleted paths, and logic correction paths based on the control flow graphs, calculate the lengths of the corresponding modified paths, and normalize the calculated lengths to obtain control flow difference features; Obtain variables with changed values, types, ranges, and dependencies in the code segments before and after modification, and perform one-hot encoding based on the modified types to obtain the modified variable features. Obtain the loop invariants that are corrected, destroyed, or newly added in the code segments before and after modification, and perform one-hot encoding based on the modification type to obtain the loop invariant difference features.

9. The method for detecting source code vulnerabilities based on multimodal feature fusion according to claim 1, characterized in that: Methods for obtaining the hazard level include: Correcting the vulnerability probability matrix based on the vulnerability detection results; When the vulnerability detection result is the same as the detection result of the corresponding position in the vulnerability probability matrix, the frequency of vulnerability occurrence in each code segment is obtained based on the vulnerability probability matrix, and a vulnerability distribution density map is drawn based on the frequency. The code segment with high density is marked as high risk, and the code segment with low density is marked as low risk. When the vulnerability detection result is different from the detection result of the corresponding position in the vulnerability probability matrix, the result of the corresponding position in the vulnerability probability matrix is ​​corrected according to the vulnerability detection result so that it is the same as the vulnerability detection result, and the corrected code vulnerability matrix is ​​obtained. The frequency of vulnerability occurrence in each code segment is obtained based on the corrected code vulnerability matrix, and a vulnerability distribution density map is drawn based on the frequency. The code segment with high density is marked as high risk, and the code segment with low density is marked as low risk.

10. A source code vulnerability detection system based on multimodal feature fusion, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Code vulnerability detection model training method and code vulnerability detection method and device

    CN118395448A