Source code deep security detection method based on artificial intelligence
By constructing the source code control flow change diagram and hierarchical structure vector, and adjusting the pruning process in combination with the CodeBERT model, the problem of inaccurate boundaries of the context capture mechanism in pruning processing is solved, and the effect of source code security detection is improved.
Patent Information
- Application Number
- CN202510884302.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing source code security detection technology in pruning processing imprecisely defines the scope boundary of the context capture mechanism, resulting in the pruning semantic structure of the source code being destroyed and reduces the security detection effect.
By converting the source code into an abstract syntax tree, a control flow change chart is constructed, the hierarchical structure vector and context semantic structure influence coefficient of the code block are obtained, and the pruning process is adjusted using the CodeBERT model to avoid excessive pruning and ensure semantic coherence.
It improves the effectiveness of source code security detection, avoids the damage to the source code semantic structure by pruning processing, and enhances the ability to identify code logical boundaries and dependencies.
Smart Images

Figure CN120387171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of source code security detection, and specifically to a deep source code security detection method based on artificial intelligence. Background Art
[0002] Performing deep security detection on source code is a core link to ensure software security and reliability. Existing source code security detection technologies mainly rely on static rule matching or dynamic fuzz testing, which have significant limitations. Traditional static analysis tools (such as SonarQube) need to pre-define vulnerability patterns and cannot identify unregistered vulnerability types. For example, logical defects in new frameworks or risk codes in context-sensitive scenarios. Such tools have a high false negative rate for dynamically generated SQL statements or indirect function call chains, and the false positive rate exceeds 40% in the scenario of mixed multi-language code.
[0003] Dynamic fuzz testing tools (such as AFL) trigger abnormal behaviors by generating random inputs, but their detection process requires the complete execution of the program and coverage of multiple paths, which is time-consuming and consumes a large amount of resources, making it difficult to integrate into the development environment for real-time operation. In addition, although existing machine learning-based detection schemes (such as DeepCode) attempt to introduce AI technology, they are mostly limited to sorting or filtering static analysis results, lacking a deep understanding of code semantics. Such models rely on manually labeled data sets, have insufficient generalization ability, and have weak parsing ability for data flow and permission context.
[0004] Therefore, in related technologies, before performing deep security detection on source code, it is usually necessary to perform pruning processing on the source code to balance semantic integrity and computational efficiency and improve the effect of security detection. However, in the pruning process, it is difficult to accurately define the scope boundary of the context capture mechanism, which affects the parameter optimization strategy in the pruning process, resulting in the system being unable to accurately identify the key logical boundaries and dependencies in the code, and further causing the semantic structure of the pruned source code to be damaged, thereby reducing the effect of source code security detection. Summary of the Invention
[0005] In order to solve the technical problem that in the pruning process, it is difficult to accurately define the scope boundary of the context capture mechanism, resulting in the semantic structure of the pruned source code being damaged and reducing the effect of source code security detection, the purpose of the present invention is to provide a deep source code security detection method based on artificial intelligence, and the specific technical solution adopted is as follows: The present invention proposes a deep source code security detection method based on artificial intelligence, and the method includes: Obtain the source code, convert the source code into an abstract syntax tree, and based on the type and hierarchical relationship of the nodes of the abstract syntax tree, obtain a control flow variation graph composed of multiple code blocks of the source code; Take any one code block as the target code block. In the control flow change graph, starting from the vertex where the target code block is located, construct the hierarchical structure vector of the target code block; according to the differences between the hierarchical structure vectors of the target code block and the lengths of the hierarchical structure vectors, obtain the context semantic structure influence coefficient of the target code block; take the sequence formed by the sequence numbers of the target code block when it is called in each other code block except the target code block as the extended order sequence of the target code block with respect to each other code block, and according to the differences between the extended order sequence of the target code block with respect to each other code block and the standard extended order sequence, as well as the context semantic structure influence coefficient of the target code block, obtain the distance dependence sensitivity coefficient of the target code block. According to the distance dependence sensitivity coefficient of the target code block, as well as the type and token of the node where the target code block is located in the abstract syntax tree, obtain the adjusted embedding structure vector of the target code block; according to the embedding structure vectors of each code block before and after adjustment, obtain the adjacency matrix of each level before and after adjustment during the pruning process; according to the differences between the adjacency matrices of the same level before and after adjustment during the pruning process, obtain the semantic coherence influence degree of each level during the pruning process. Based on the semantic coherence influence degree of each level during the pruning process, adjust the pruning process and perform security detection on the source code.
[0006] Further, the obtaining of the context semantic structure influence coefficient of the target code block includes: Take any two of the hierarchical structure vectors of the target code block as a vector group, take the cosine similarity of the two hierarchical structure vectors in each vector group as the numerator, take the sum value of the lengths of the two hierarchical structure vectors in each vector group as the denominator, and take the ratio as the similarity parameter of each vector group. Integrate the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block.
[0007] Further, the integrating of the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block includes: Normalize the average value of the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block.
[0008] Further, the obtaining of the distance dependence sensitivity coefficient of the target code block includes: Take the Manhattan distance between the extended order sequence of the target code block with respect to each other code block and the standard extended order sequence as the deviation performance value of the target code block with respect to each other code block. Use the average of the bias performance values of the target code block with respect to all other code blocks as the numerator, and the sum of the context semantic structure influence coefficient and the preset adjustment parameter of the target code block as the denominator, and perform normalization on the ratio value to obtain the distance-dependent sensitivity coefficient of the target code block.
[0009] Further, the obtaining of the adjusted embedding structure vector of the target code block includes: Based on the calculation formula of the embedding structure vector, obtain the adjusted embedding structure vector of the target code block, and the calculation formula of the embedding structure vector is: Wherein, represents the adjusted embedding structure vector of the target code block; represents the type of the node where the target code block is located in the abstract syntax tree; represents the token of the node where the target code block is located in the abstract syntax tree; represents the node type embedding function; represents the token embedding function; represents the distance-dependent sensitivity coefficient of the target code block.
[0010] Further, the obtaining of the adjacency matrix before and after adjustment for each level in the pruning process includes: Input the embedding structure vectors of all code blocks before and after adjustment into the CodeBERT model respectively, and the CodeBERT model outputs the adjacency matrix before and after adjustment for each level in the pruning process.
[0011] Further, the obtaining of the semantic coherence influence degree for each level in the pruning process includes: Take the difference between the adjacency matrices before and after adjustment for the same level in the pruning process as the differential adjacency matrix for each level in the pruning process; Perform normalization on the determinant of the differential adjacency matrix for each level in the pruning process to obtain the semantic coherence influence degree for each level in the pruning process.
[0012] Further, the security detection of the source code includes: Use the semantic coherence influence degree for each level in the pruning process as the weight of the hash contribution value for each level in the pruning process in the CodeBERT model, input the source code into the CodeBERT model, the CodeBERT model performs layer pruning operation on the source code and outputs a detection model, and input the detection model into a lightweight sandbox to quickly verify the authenticity of the vulnerability through preset test cases.
[0013] Further, the construction of the hierarchical structure vector of the target code block includes: In the control flow change diagram, taking the vertex where the target code block is located as the starting point and the vertex where any other code block except the target code block is located as the ending point, the sequence formed by the serial numbers of all vertices passed from the starting point to the ending point is used as a hierarchical structure vector of the target code block.
[0014] Furthermore, the control flow change diagram formed by multiple code blocks of the obtained source code includes: Using a lexical analyzer and a syntax analyzer, the source code is converted into an abstract syntax tree, and based on the types and hierarchical relationships of the nodes of the abstract syntax tree, the logical boundaries of the code blocks corresponding to each node are calculated, so as to obtain multiple code blocks of the source code and generate a control flow change diagram formed by each code block.
[0015] The present invention has the following beneficial effects: The present invention takes into account that in the pruning process, it is relatively difficult to accurately define the scope boundary of the context capture mechanism, resulting in the destruction of the semantic structure of the pruned source code and reducing the effect of source code security detection. Therefore, first, the source code is obtained, the source code is converted into an abstract syntax tree, and based on the types and hierarchical relationships of the nodes of the abstract syntax tree, a control flow change diagram formed by multiple code blocks of the source code is obtained. Subsequently, the dynamic changes of the control flow during the execution of each code block can be accurately described through the control flow change diagram, and a hierarchical structure vector of the target code block can be constructed. Furthermore, the correlation between the function executed by the target code block and the context source code of its location can be reflected through the obtained context semantic structure influence coefficient. Furthermore, the performance consistency of the expansion process of the target code block under the semantic structure can be reflected through the obtained distance dependence sensitivity coefficient, avoiding excessive pruning of the source code subsequently. Then, the pruning process is adjusted. First, an adjusted embedding structure vector is obtained, and then the differences between the adjacency matrices of the same level before and after the adjustment during the pruning process are analyzed. The influence degree of the pruning process on the semantic coherence of the source code is reflected through the semantic coherence influence degree. Furthermore, based on the semantic coherence influence degree, the pruning process is adjusted to avoid the destruction of the semantic structure of the source code after the pruning process, thereby improving the effect of in-depth security detection of the source code. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1Flowchart of a deep security detection method for source code based on artificial intelligence provided by an embodiment of the present invention; Figure 2 Schematic diagram of a control flow change graph provided by an embodiment of the present invention. Detailed implementation manners
[0018] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and effects of a deep security detection method for source code based on artificial intelligence proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0020] The following specifically describes the specific solution of a deep security detection method for source code based on artificial intelligence provided by the present invention with reference to the accompanying drawings.
[0021] Please refer to Figure 1 , which shows a flowchart of a deep security detection method for source code based on artificial intelligence provided by an embodiment of the present invention. The method includes: Step S1: Obtain the source code, convert the source code into an abstract syntax tree, and obtain a control flow change graph composed of multiple code blocks of the source code based on the type and hierarchical relationship of the nodes of the abstract syntax tree.
[0022] In the embodiment of the present invention, the source code is first extracted from the code library, converted into an abstract syntax tree (AST), and a control flow change graph composed of multiple code blocks of the source code is obtained based on the type and hierarchical relationship of the nodes of the abstract syntax tree. Subsequently, the dynamic changes of the control flow during the execution of each code block can be accurately described through the control flow change graph, so as to better understand the dynamic behavior of the code, optimize the code logic or discover potential problems, and improve the effect of security detection.
[0023] Preferably, in an embodiment of the present invention, the method for obtaining a control flow change graph composed of multiple code blocks of the source code specifically includes: First, use a lexical analyzer and a syntax analyzer to convert the source code into an abstract syntax tree. Each node in the abstract syntax tree represents function calls, control flow nodes (such as if-else structures, for loop structures), and variable scopes, etc. For example, when parsing Java code, the conditional expression and code block of an if statement will be mapped to the IfStatement node in the abstract syntax tree.
[0024] Then, based on the types and hierarchical relationships of the nodes in the abstract syntax tree, calculate the logical boundaries of the code blocks corresponding to each node. For example, for nested structures, depth-first traversal can be used to record the starting line number and ending line number. For loop statements, extract the iteration variables and termination conditions, and mark the code range covered by the loop body, so as to obtain multiple code blocks of the source code and generate a control flow variation graph composed of these code blocks. Please refer to Figure 2 , which shows a schematic diagram of the control flow variation graph provided by an embodiment of the present invention. Among them, the control flow variation graph is a directed graph, and each vertex can represent a code block.
[0025] So far, the parsing of the source code has been realized, and the corresponding control flow variation graph has been obtained.
[0026] Step S2: Take any code block as the target code block. In the control flow variation graph, starting from the vertex where the target code block is located, construct a hierarchical structure vector of the target code block; according to the differences between the hierarchical structure vectors of the target code block and the lengths of the hierarchical structure vectors, obtain the context semantic structure influence coefficient of the target code block; take the sequence composed of the sequence numbers when the target code block is called in each other code block except the target code block as the extended order sequence of the target code block with respect to each other code block. According to the differences between the extended order sequence of the target code block with respect to each other code block and the standard extended order sequence, and the context semantic structure influence coefficient of the target code block, obtain the distance dependence sensitivity coefficient of the target code block.
[0027] The context capture mechanism determines the local information density that the model needs to process by limiting the window range, fixed-length sliding window or dynamic expansion mechanism. During the pruning process, by analyzing the attention weight distribution, attention heads or neurons with low contribution to the key information within the window can be preferentially removed. For example, under the dynamic expansion mechanism, the model has a stronger ability to capture long-distance dependencies. At this time, pruning needs to retain the key parameters related to cross-window associations. The fixed window design retains the recent context through a circular buffer. When pruning, it is necessary to ensure that the core semantic modeling layer, such as the feed-forward network of Transformer, is not overly reduced to avoid destroying the semantic coherence across windows.
[0028] First, analyze the expansion process when a single code block is called. Take any code block as the target code block. In the control flow change graph, the vertex corresponding to the target code block after the logical boundary change is the next vertex pointed to by the vertex it corresponds to, until the entire expansion process is completed. Therefore, it is necessary to consider the semantic differences of the code block context. Here, the code block changes in the expansion process are used to represent the context language structure. By analyzing the graph structure characteristics of the control structures corresponding to different functions, that is, the different connection structures on the control flow change graph, for the target code block, first, in the control flow change graph, starting from the vertex where the target code block is located, construct the hierarchical structure vector of the target code block.
[0029] Preferably, in an embodiment of the present invention, the method for obtaining the hierarchical structure vector of the target code block specifically includes: In the control flow change graph, starting from the vertex where the target code block is located and ending at the vertex where any other code block except the target code block is located, the sequence formed by the numbers of all vertices passed from the starting point to the ending point is used as a hierarchical structure vector of the target code block. Please refer to Figure 2 , for example, vertex v1 represents the target code block. If the vertex v10 where another code block is located is selected as the ending point, a formed hierarchical structure vector is (1, 6, 4, 5, 10).
[0030] Analyze the difference degree of all hierarchical structure vectors of the target code block. Considering that the lengths of different hierarchical structure vectors of the target code block are different, therefore, according to the differences of the hierarchical structure vectors of the target code block and the lengths of each hierarchical structure vector, obtain the context semantic structure influence coefficient of the target code block. The context semantic structure influence coefficient reflects the association between the function executed by the target code block and the context source code at its location, that is, the context semantic performance with the target code block as the starting point of the graph structure. Among them, the length of the hierarchical structure vector is the number of elements included in the hierarchical structure vector.
[0031] Preferably, in an embodiment of the present invention, the method for obtaining the context semantic structure influence coefficient of the target code block specifically includes: First, take any two hierarchical structure vectors of the target code block as a vector group. Use the cosine similarity of the two hierarchical structure vectors in each vector group as the numerator, and use the sum value of the lengths of the two hierarchical structure vectors in each vector group as the denominator. Take the ratio as the similarity parameter of each vector group.
[0032] It should be noted that the lengths of the two hierarchical structure vectors in the vector group may be different. During the calculation of the cosine similarity of the two hierarchical structure vectors, for the vector group with different lengths of hierarchical structure vectors, it is necessary to perform a zero-padding operation on the hierarchical structure vector with a shorter length to make the lengths of the two hierarchical structure vectors the same. At the same time, it should be noted that the length of the hierarchical structure vector after zero-padding is still the original length.
[0033] Then, the similarity parameters of all vector groups are integrated to obtain the context semantic structure influence coefficient of the target code block.
[0034] In an embodiment of the present invention, the average value of the similarity parameters of all vector groups can be normalized, and the calculation result is limited to within a range, so as to obtain the context semantic structure influence coefficient of the target code block.
[0035] In an embodiment of the present invention, the normalization process can be, for example, the maximum-minimum normalization process, and the normalization in subsequent steps can all adopt the maximum-minimum normalization process. In other embodiments of the present invention, other normalization methods can be selected according to the specific range of values, which will not be elaborated here.
[0036] As an example, in an embodiment of the present invention, the expression of the context semantic structure influence coefficient of the target code block can be, for example: where represents the context semantic structure influence coefficient of the target code block; represents the cosine similarity of the two hierarchical structure vectors in the th vector group; and respectively represent the lengths of the two hierarchical structure vectors in the th vector group; represents the similarity parameter of the th vector group; represents the number of vector groups; represents the normalization function for normalization processing.
[0037] Since the dynamic expansion mechanism significantly improves the ability to capture long-distance dependencies in source code analysis through global attention coverage and hierarchical computing optimization, its ability to capture is reflected in taking code blocks, such as functions, loop structures, and exception handling units, as logical units, dynamically identifying logical boundaries using the abstract syntax tree and establishing cross-level associations. For example, when tracing the call of high-risk functions such as exec, the global attention mechanism can cross-functionally associate the input verification logic to ensure the integrity of the pollution propagation chain. In the scenario of multi-layer loop nesting, the model synchronously captures the dependencies between the outer configuration variables and the inner execution logic through multi-head attention, avoiding semantic breaks caused by overly deep loop levels. Then, the more the context semantic structure of the code block spans multiple syntactic structures, such as the association relationships between main and subordinate clauses and nested loops, for example, the more cross-function matches between function calls and input verification logic in the code, the more obvious the deep-level association relationships it shows. Therefore, in the embodiments of the present invention, first, the sequence composed of the sequence numbers when the target code block is called in each other code block except the target code block is used as the extension order sequence of the target code block with respect to each other code block, and based on the difference between the extension order sequence of the target code block with respect to each other code block and the standard extension order sequence, as well as the influence coefficient of the context semantic structure of the target code block, the distance-dependent sensitivity coefficient of the target code block is obtained.
[0038] Preferably, in an embodiment of the present invention, the method for obtaining the distance-dependent sensitivity coefficient of the target code block specifically includes: Taking the Manhattan distance between the extension order sequence of the target code block with respect to each other code block and the standard extension order sequence as the deviation performance value of the target code block with respect to each other code block, where the standard extension order sequence can be obtained from the general extension order performance of the code block obtained through big data and will not be elaborated here.
[0039] It should be noted that when calculating the Manhattan distance between the extension order sequence and the standard extension order sequence, for a certain extension order sequence, if the lengths of the extension order sequence and the standard extension order sequence are different, the shorter sequence needs to be padded with 0s to make the lengths of the extension order sequence and the standard extension order sequence the same.
[0040] Taking the average value of the deviation performance values of the target code block with respect to all other code blocks as the numerator, taking the sum value of the influence coefficient of the context semantic structure of the target code block and the preset adjustment parameter as the denominator, and normalizing the ratio value, and limiting the calculation result within the range, so as to obtain the distance-dependent sensitivity coefficient of the target code block.
[0041] As an example, in an embodiment of the present invention, the expression of the distance-dependent sensitivity coefficient of the target code block can be specifically, for example: Among them, represents the distance dependence sensitivity coefficient of the target code block; represents the deviation performance value of the target code block with respect to the th other code block; represents the number of other code blocks except the target code block; represents the context semantic structure influence coefficient of the target code block; represents a preset adjustment parameter used to prevent the denominator from being zero, The value range of is In one embodiment of the present invention, is set to 0.01. The specific value of can also be set by the implementer according to the specific implementation scenario and is not limited herein;
[0042] Among them, the deviation performance value of the target code block with respect to other code blocks is used to measure the deviation performance of the expansion order sequence of the target code block compared with the standard expansion order sequence. The larger the value, the more obvious the call and verification of the target code block in the expansion process in terms of the association depth. is used to reflect the performance consistency of the expansion process of the target code block under the semantic structure, that is, the confidence performance of the expansion association depth of the target code block. The larger the value, the more attention needs to be paid to the retention of the expansion relationship of its code block during the subsequent pre-training process of the artificial intelligence model, so as to avoid excessive pruning and improve the pruning effect of the subsequent source code.
[0043] Step S3: Obtain the adjusted embedded structure vector of the target code block according to the distance dependence sensitivity coefficient of the target code block, and the type and token of the node where the target code block is located in the abstract syntax tree; obtain the adjacency matrix of each level before and after adjustment during the pruning process according to the embedded structure vectors of each code block before and after adjustment; obtain the influence degree of semantic coherence of each level during the pruning process according to the difference between the adjacency matrices of the same level before and after adjustment during the pruning process.
[0044] After the code segments of the source code are tokenized and vectorized through the dynamic expansion mechanism, the analysis scope is dynamically expanded or contracted according to the code structure characteristics to ensure the capture of the complete logical chain, and the extraction of the logical structure of the source code is basically realized. Further, through the layer pruning of the pre-trained model (CodeBERT), the memory resource occupancy requirements are effectively reduced. For the pruning process, gradually removing the layers of a single-layer Transformer will cause the cross-function variable dependency relationship of the structural vector of the abstract syntax tree of the code block to be unable to be correctly captured, that is, it will affect the performance of the distance-dependent sensitivity coefficient of the code block on the adjacency matrix. Then, through the change of the pruning situation layer by layer, the impact of the pruning process on the semantic coherence of the code block can be quantified, avoiding the destruction of the semantic structure of the source code caused by pruning, thereby reducing the effect of the security detection of the source code. Therefore, in the embodiment of the present invention, first, according to the distance-dependent sensitivity coefficient of the target code block, as well as the type and token of the node where the target code block is located in the abstract syntax tree, the adjusted embedding structure vector of the target code block is obtained. Subsequently, based on the embedding structure vectors of each code block before and after adjustment, the adjacency matrices of each layer before and after adjustment in the pruning process can be obtained.
[0045] Preferably, in an embodiment of the present invention, the method for obtaining the adjusted embedding structure vector of the target code block specifically includes: Based on the calculation formula of the embedding structure vector, the adjusted embedding structure vector of the target code block is obtained, and the calculation formula of the embedding structure vector is: Wherein, represents the adjusted embedding structure vector of the target code block; represents the type of the node where the target code block is located in the abstract syntax tree; represents the token of the node where the target code block is located in the abstract syntax tree; represents the node type embedding function, which is used to map the type of the node where the target code block is located in the abstract syntax tree into a vector; represents the token embedding function, which is used to map tokens such as identifiers and operators of the node where the target code block is located in the abstract syntax tree into a vector; represents the distance-dependent sensitivity coefficient of the target code block.
[0046] Furthermore, based on the embedding structure vectors of each code block before and after adjustment, the adjacency matrices of each layer before and after adjustment in the pruning process are obtained.
[0047] Preferably, in an embodiment of the present invention, the method for obtaining the adjacency matrices of each layer before and after adjustment in the pruning process specifically includes: The embedding structure vectors of all code blocks before and after adjustment are respectively input into the CodeBERT model, and the CodeBERT model outputs the adjacency matrices of each level during the pruning process before and after adjustment. Among them, the embedding structure vector of the code block before adjustment is the vector without being adjusted by the distance dependence sensitivity coefficient. For example, for the target code block, its embedding structure vector before adjustment is , where represents the embedding structure vector of the target code block before adjustment.
[0048] The greater the difference between the adjacency matrices of the same level during the pruning process before and after adjustment, the greater the impact of the pruning on the semantic coherence of the code block. Therefore, according to the difference between the adjacency matrices of the same level during the pruning process before and after adjustment, the impact degree of semantic coherence of each level during the pruning process can be obtained, and the impact degree of the pruning process on the semantic coherence of the source code can be reflected through the impact degree of semantic coherence. Subsequently, based on the impact degree of semantic coherence, the layer pruning of the CodeBERT model can be adjusted to avoid the pruning process from damaging the semantic structure of the source code, thereby improving the effect of deep security detection.
[0049] Preferably, in an embodiment of the present invention, the method for obtaining the impact degree of semantic coherence of each level during the pruning process specifically includes: Taking the difference between the adjacency matrices of the same level during the pruning process before and after adjustment as the differential adjacency matrix of each level during the pruning process, and normalizing the determinant of the differential adjacency matrix of each level during the pruning process, and limiting the calculation result within the range, so as to obtain the impact degree of semantic coherence of each level during the pruning process.
[0050] So far, the impact degree of semantic coherence of each level during the pruning process has been obtained. Subsequently, based on the impact degree of semantic coherence, the layer pruning of the CodeBERT model can be adjusted to improve the effect of pruning the source code.
[0051] Step S4: Based on the impact degree of semantic coherence of each level during the pruning process, adjust the pruning process and perform security detection on the source code.
[0052] To avoid the pruning operation from damaging the semantic structure of the source code, the pruning process can be adjusted based on the impact degree of semantic coherence of each level during the pruning process, and security detection can be performed on the source code, thereby improving the effect of pruning the source code and the effect of performing deep security detection on the source code.
[0053] Preferably, in an embodiment of the present invention, the method for performing security detection on the source code specifically includes: Use the degree of semantic coherence impact at each level during the pruning process as the weight of the hash contribution value at each level during the pruning process in the CodeBERT model, and input the source code into the CodeBERT model. The CodeBERT model performs layer pruning operations on the source code and outputs a detection model. Input the detection model into a lightweight sandbox, and quickly verify the authenticity of vulnerabilities through preset test cases (such as SQL injection feature strings). At the same time, during the detection process, the development environment integration interface (such as a VSCode plugin) can also be used to trigger the detection process and return risk markers when the code is saved, and at the same time support clicking to view repair suggestions.
[0054] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0055] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A deep security detection method for source code based on artificial intelligence, characterized in that The method includes: Obtain the source code, convert the source code into an abstract syntax tree, and based on the types and hierarchical relationships of the nodes of the abstract syntax tree, obtain a control flow change graph composed of multiple code blocks of the source code; Take any one code block as the target code block. In the control flow change graph, starting from the vertex where the target code block is located, construct a hierarchical structure vector of the target code block; according to the differences between the hierarchical structure vectors of the target code block and the lengths of the hierarchical structure vectors, obtain the context semantic structure influence coefficient of the target code block; use the sequence formed by the sequence numbers when the target code block is called in each other code block except the target code block as the extended order sequence of the target code block with respect to each other code block, and according to the differences between the extended order sequence of the target code block with respect to each other code block and the standard extended order sequence, as well as the context semantic structure influence coefficient of the target code block, obtain the distance dependence sensitivity coefficient of the target code block; According to the distance dependence sensitivity coefficient of the target code block, as well as the type and token of the node where the target code block is located in the abstract syntax tree, obtain the adjusted embedding structure vector of the target code block; according to the embedding structure vectors of each code block before and after adjustment, obtain the adjacency matrix of each level before and after adjustment during the pruning process; according to the differences between the adjacency matrices of the same level before and after adjustment during the pruning process, obtain the semantic coherence influence degree of each level during the pruning process; Based on the semantic coherence influence degree of each level during the pruning process, adjust the pruning process and perform security detection on the source code.
2. The method for in-depth security detection of source code based on artificial intelligence according to claim 1, characterized in that The obtaining of the context semantic structure influence coefficient of the target code block includes: Take any two of the hierarchical structure vectors of the target code block as a vector group, use the cosine similarity of the two hierarchical structure vectors in each vector group as the numerator, use the sum of the lengths of the two hierarchical structure vectors in each vector group as the denominator, and use the ratio as the similarity parameter of each vector group; Integrate the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block.
3. The method for in-depth security detection of source code based on artificial intelligence according to claim 2, wherein, The integrating of the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block includes: Normalize the average value of the similarity parameters of all vector groups to obtain the context semantic structure influence coefficient of the target code block.
4. The method for deep security detection of source code based on artificial intelligence according to claim 1, characterized in that The obtaining of the distance dependence sensitivity coefficient of the target code block includes: Take the Manhattan distance between the extended order sequence of the target code block with respect to each other code block and the standard extended order sequence as the deviation performance value of the target code block with respect to each other code block; Use the average value of the deviation performance values of the target code block with respect to all other code blocks as the numerator, use the sum of the context semantic structure influence coefficient of the target code block and a preset adjustment parameter as the denominator, and normalize the ratio to obtain the distance dependence sensitivity coefficient of the target code block.
5. A deep security detection method for source code based on artificial intelligence according to claim 1, characterized in that, The obtaining of the adjusted embedding structure vector of the target code block includes: Based on the calculation formula of the embedded structure vector, obtain the adjusted embedded structure vector of the target code block. The calculation formula of the embedded structure vector is as follows: Among them, represents the adjusted embedding structure vector of the target code block; represents the type of the node where the target code block is located in the abstract syntax tree; represents the token of the node where the target code block is located in the abstract syntax tree; represents the node type embedding function; represents the token embedding function; represents the distance dependence sensitivity coefficient of the target code block.
6. The method for deep security detection of source code based on artificial intelligence according to claim 1, characterized in that The obtaining of the adjacency matrices of each level before and after adjustment in the pruning process includes: Input the embedded structure vectors of all code blocks before and after adjustment into the CodeBERT model respectively, and the CodeBERT model outputs the adjacency matrices of each level before and after adjustment in the pruning process.
7. A deep security detection method for source code based on artificial intelligence according to claim 1, characterized in that The obtaining of the influence degree of semantic coherence of each level in the pruning process includes: Take the difference between the adjacency matrices of the same level before and after adjustment in the pruning process as the differential adjacency matrix of each level in the pruning process; Normalize the determinant of the differential adjacency matrix of each level in the pruning process to obtain the influence degree of semantic coherence of each level in the pruning process.
8. An in-depth security detection method for source code based on artificial intelligence according to claim 1, characterized in that The security detection of the source code includes: Take the influence degree of semantic coherence of each level in the pruning process as the weight of the hash contribution value of each level in the pruning process in the CodeBERT model, input the source code into the CodeBERT model, the CodeBERT model performs layer pruning operation on the source code and outputs a detection model, and input the detection model into a lightweight sandbox to quickly verify the authenticity of vulnerabilities through preset test cases.
9. A deep security detection method for source code based on artificial intelligence according to claim 1, characterized in that, The construction of the hierarchical structure vector of the target code block includes: In the control flow change graph, with the vertex where the target code block is located as the starting point and the vertex where any other code block except the target code block is located as the ending point, the sequence composed of the serial numbers of all vertices passed from the starting point to the ending point is used as a hierarchical structure vector of the target code block.
10. A deep security detection method for source code based on artificial intelligence according to claim 1, characterized in that The obtaining of the control flow change graph composed of multiple code blocks of the source code includes: Use a lexical analyzer and a syntax analyzer to convert the source code into an abstract syntax tree, and based on the types and hierarchical relationships of the nodes of the abstract syntax tree, calculate the logical boundaries of the code blocks corresponding to each node, so as to obtain multiple code blocks of the source code and generate a control flow change graph composed of each code block.