Code Analysis and Detection Method and System Based on Large Language Model
By parsing the code structure diagram on the code detection platform and introducing gated constraints, the large language model is subjected to directional distillation, which solves the problems of missed detection and false positives in traditional code analysis methods, and efficient and accurate code detection is achieved.
Patent Information
- Application Number
- CN202510533603.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Traditional code analysis methods cannot fully consider the code context and multi-level detection dimensions, resulting in frequent missed detection and false alarms.
By deploying the parsing engine on the code detection platform for programming paradigm analysis, reconstructing the code structure diagram, defining the detection direction and introducing gated constraints under multi-layer gradual detection, performing directions of large language models, and performing first-order accessibility detection and second-order vulnerability detection.
It realizes accurate identification of potential vulnerabilities and optimization points in the code, and improves the accuracy and efficiency of code detection.
Smart Images

Figure CN120066935B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically to a code analysis and detection method and system based on large language models. Background Art
[0002] With the rapid expansion of the scale of software systems and the continuous evolution of diverse programming paradigms, the complexity of code and its multi-layer nested structure have increased significantly. Traditional code analysis and detection methods have gradually revealed their limitations. On the one hand, although static analysis tools can identify syntax errors and some structural defects, when faced with complex code structures across functions, modules, and even languages, they often struggle to accurately capture context dependencies, resulting in missed reports or false positives in detection results. On the other hand, manual review methods not only rely on experience and are inefficient but also difficult to meet the automation requirements in large-scale continuous integration environments. In recent years, large language models have demonstrated powerful context awareness and knowledge transfer capabilities in natural language understanding and generation, providing new possibilities for complex code understanding and semantic reasoning. Therefore, there is an urgent need for an intelligent code detection method that combines the capabilities of large language models, integrates programming semantic structures and multi-layer logical relationships, to meet the urgent needs for high-precision, wide-coverage, and high-efficiency code analysis in current software development. Summary of the Invention
[0003] This application provides a code analysis and detection method and system based on large language models, aiming to solve the technical problem that traditional code analysis methods cannot fully consider code context and multi-level detection dimensions when processing code fragments, resulting in frequent missed detections and false positives.
[0004] In the first aspect disclosed in this application, a code analysis and detection method based on large language models is provided. The method includes: deploying a parsing engine on a code detection platform to parse the programming paradigm of a code fragment, and reconstructing a code structure diagram based on the code element-context relationship; defining a detection direction according to the code structure diagram, and introducing a gating constraint based on multi-layer progressive detection, where the gating constraint is determined based on the detection dimension and code fragment granularity of the detection direction; performing directional distillation processing on the large language model according to the detection direction and the gating constraint, taking the code structure diagram as the logical guide and the code fragment as the detection target, and performing code analysis and detection to determine the code detection result; where the detection step includes first-order reachability detection and second-order vulnerability detection.
[0005] Another aspect disclosed in this application provides a code analysis and detection system based on a large language model. The system includes: a code parsing module: deploying a parsing engine on a code detection platform to parse the programming paradigm of code snippets, and reconstructing a code structure diagram based on the code element-context relationship; a constraint introduction module: defining a detection direction according to the code structure diagram and introducing a gated constraint based on multi-level progressive detection, where the gated constraint is determined based on the detection dimension and code fragment granularity in the detection direction; a code detection module: performing directional distillation processing on the large language model according to the detection direction and the gated constraint, using the code structure diagram as the logical guide and the code snippet as the detection target, and performing code analysis and detection to determine the code detection result; where the detection steps include first-order reachability detection and second-order vulnerability detection.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0007] The above code analysis and detection method based on a large language model first deploys a parsing engine on a code detection platform to parse the programming paradigm of code snippets, and reconstructs the code structure diagram by analyzing the relationship between code elements and context. Subsequently, according to this structure diagram, a detection direction is defined, and a gated constraint based on multi-level progressive detection is introduced. The setting of the gated constraint is based on the detection dimension and the granularity of code fragments. Then, the detection direction and the gated constraint are used to perform directional distillation on the large language model to optimize the model to meet the requirements of code analysis and detection. Finally, taking the code structure diagram as the guide and combining the code snippets, an analysis including first-order reachability detection and second-order vulnerability detection is performed to determine the code detection result, so as to accurately identify potential vulnerabilities and optimization points in the code, and improve the accuracy and efficiency of code detection.
[0008] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the specific embodiments of this application are hereby specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 It is a schematic flowchart of a code analysis and detection method based on a large language model in an embodiment.
[0011] Figure 2 It is an architecture diagram of a code analysis and detection system based on a large language model in an embodiment.
[0012] Explanation of reference numerals: Code parsing module 11, Constraint introduction module 12, Code detection module 13. Specific implementation manner
[0013] In the embodiments of the present application, by providing a code analysis and detection method and system based on a large language model, the technical problem that traditional code analysis methods cannot fully consider code context and multi-level detection dimensions when processing code fragments, resulting in frequent missed detections and false alarms, is solved.
[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0015] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0016] Embodiment 1, as Figure 1 shown, the present application provides a code analysis and detection method based on a large language model, and the method includes:
[0017] Deploy a parsing engine on the code detection platform to perform programming paradigm parsing on the code fragment, and reconstruct the code structure diagram based on the code element-context relationship.
[0018] In the embodiments of the present application, on the code detection platform, a parsing engine is deployed. This engine is responsible for performing programming paradigm parsing on the input code fragment. By analyzing the basic elements of the code (such as variables, functions, classes, etc.) and their context relationships, the parsing engine can understand the overall structure and relationships of the code, and determine the parsing structure. Finally, based on the determined parsing structure, the code structure diagram is reconstructed. This structure diagram will show the various components of the code and their mutual relationships. This process helps to comprehensively understand the architecture and semantics of the code, laying a foundation for subsequent code analysis and detection.
[0019] Furthermore, the present application provides a reconstructed code structure diagram, including:
[0020] Connect to the programming database, where the programming database includes a low-code sub-library and a programming language library; use the low-code sub-library to perform first-order matching and parsing on the code snippet to determine the first parsing structure; use the programming language library to perform second-order matching and parsing on the code snippet to determine the second parsing structure; fuse the first parsing structure and the second parsing structure to reconstruct the code structure diagram.
[0021] Preferably, first, establish a connection with the programming database built into the code detection platform. This built-in programming database includes a low-code sub-library and a programming language library. Among them, the low-code sub-library is oriented to general business logic and is used to process standardized and componentized code patterns; the programming language library covers various programming grammars and structural rules and is used to adapt to multiple development languages. Subsequently, use the low-code sub-library to perform first-order matching and parsing on the code snippet to be analyzed. This step mainly identifies whether there are structures similar to common business templates or standard components in the code and generates the corresponding first parsing structure, such as basic frameworks for form control calls and process node controls. Then, use the programming language library to perform second-order matching and parsing on the same code snippet, focusing on parsing details such as specific syntax implementation, function logic, and nested calls, so as to form the second parsing structure, which represents the underlying syntax and logical organization method of the code, such as variable definition, function call, control statement, scope, etc. Then, based on the unique identifier of the code snippet (such as function name, variable name, or statement position information), establish a mapping relationship between the first parsing structure and the second parsing structure to ensure that the parsing results at different granularities can be uniformly mapped to the same code area. Then, through the graph fusion algorithm, align the business semantic nodes in the first parsing structure with the syntax logic nodes in the second parsing structure and merge the boundaries. For example, for two nodes representing the same code function, merge them and mark their dual semantics (such as business logic + syntax call); for complementary nodes, establish cross-layer connections (such as the call edge between a low-code control and its underlying event binding function). After completing the node fusion, generate a complete code structure diagram according to the semantic dependency relationship and control / data flow logic. Each node in the diagram not only contains structure information but also is associated with context semantics, execution path, and other labels, which can be used as a unified logical guiding diagram for code understanding, problem location, and vulnerability detection, providing an accurate semantic basis for subsequent detection.
[0022] Furthermore, the present application provides matching and parsing of the code snippet, including:
[0023] Determine the code paradigm framework based on the programming paradigm as the matching criterion; according to the code paradigm framework, locate the code elements through the parsing engine to determine the code entities; according to the code entities, through the parsing engine, perform context relationship parsing under the constraints of the code paradigm framework to determine the code relationships; according to the code paradigm framework, code entities and code relationships, determine the parsing structure.
[0024] Optionally, when performing the matching and parsing of code snippets, whether it is first-order matching and parsing or second-order matching and parsing, it is necessary to determine the code paradigm framework, code entities, and code relationships according to the programming paradigm and the parsing engine. Taking the first-order matching and parsing of code snippets based on the low-code sub-library as an example, after accessing the low-code sub-library, first use the built-in low-code programming paradigm in it as the matching criterion. For example, form-driven paradigm, event response paradigm, process logic paradigm, etc., to identify the code paradigm framework corresponding to the code to be detected. This identification process usually depends on key module names, component call syntax, and their business semantic tags, etc. By keyword matching and syntax template comparison, determine the low-code paradigm category to which the code segment belongs. After the code paradigm framework is determined, based on the rules of the code paradigm, use the parsing engine to perform structural positioning on the code snippet, identify the key code elements in it, such as form fields, binding logic, control components, trigger conditions, interface references, etc., and map these elements to structured code entities, and each entity is marked with its function, type, and position in the paradigm structure. Subsequently, use the identified code entities to perform context relationship parsing under the structural constraint rules of the selected paradigm framework. In this process, the parsing engine will analyze the explicit or implicit connection relationships between entities, such as "binding of fields and controls", "triggering of buttons and events", "mapping of data sources and interface calls", etc., so as to determine the complete code relationships, and attach the corresponding business semantics or call path information to each pair of relationships. Finally, integrate the extracted code paradigm framework, the set of code entities, and the context relationships between entities, and organize them into a graphical structure or a tree structure, so as to construct the first parsing structure of the code snippet. This structure not only reflects the module attribution and functional logic of the code in the low-code system, but also provides a standardized input basis for subsequent structure fusion, gating constraint introduction, and model analysis.
[0025] According to the code structure diagram, define the detection direction, and introduce the gating constraint based on multi-layer progressive detection, where the gating constraint is determined according to the detection dimension and code fragment granularity based on the detection direction.
[0026] In one embodiment, first, the reconstructed code structure diagram is analyzed. According to the information shown in the code structure diagram and the actual business requirements, determine which dimensions in the code need to be key detected, and define the detection direction of the code based on this to ensure the comprehensiveness and accuracy of the detection. Among them, the code structure diagram shows information such as the logical structure of the code, module relationships, data flow, and control flow; the detection direction refers to the analysis direction or dimension in the code detection process, which can be different detection levels in the code, such as the syntax level, semantic level, data flow level, etc. Subsequently, introduce the gating constraint based on multi-layer progressive detection. The gating constraint is to gradually filter unnecessary consumption of computing resources between different detection stages and only perform a more in-depth analysis on the parts that may have problems. By setting the gating constraint, the analysis intensity of the code can be flexibly adjusted according to different dimensions of the detection direction. For example, if the complexity of a certain module in the code is high (such as deeply nested loops or recursive functions), different detection granularities, that is, code fragment granularities, will be set according to the complexity, so as to break down the complex parts into smaller units for analysis. In this way, the analysis of code fragments can be gradually deepened, thereby improving efficiency and reducing false positives. Generally speaking, the determined detection direction defines the direction and focus of the analysis, and the determined gating constraint controls the depth and scope of the analysis through a multi-layer progressive strategy. These two pieces of data can make the detection process both comprehensive and efficient.
[0027] Furthermore, the present application provides a gating constraint introduced under multi-layer progressive detection, including:
[0028] Set the detection dimension, where the detection dimension at least includes syntax - semantics - data flow - control flow; according to the code structure diagram, evaluate the complexity of the code fragment, and locate the graph nodes with progressive layering requirements, where the evaluation factors at least include cross-modal, logical complexity, and syntax nesting complexity; for the graph nodes, set the gating constraint.
[0029] Preferably, according to the information shown in the code structure diagram and the actual business requirements, determine the detection dimensions. These dimensions define different levels of code analysis and detection, including at least four dimensions: syntax, semantics, data flow, and control flow. Among them, the syntax dimension mainly checks the syntax structure of the code and can identify common syntax errors, including basic syntax problems such as identifier mismatches and unbalanced parentheses; the semantics dimension focuses on verifying the logical and semantic relationships of the code to ensure that the code logically conforms to the expected design and avoids potential logical loopholes or semantic inconsistencies; the data flow dimension pays attention to the data flow in the code to ensure that there are no errors in the data transfer process and that the definition and use of variables meet the expectations. For example, check whether variables are correctly initialized or whether there are unused redundant variables; the control flow dimension focuses on analyzing the control structure of the program, including key logics such as conditional statements and loop structures, to ensure that there are no infinite loops, invalid execution paths, or unnecessary branch judgments in the control flow logic. Subsequently, starting from the detection dimensions, evaluate the complexity of the code structure diagram according to the pre-set evaluation factors. Among them, the evaluation factors include cross-modal, logical complexity, and syntax nesting complexity. When performing cross-modal evaluation on the four dimensions of syntax, semantics, data flow, and control flow, three aspects of cross-modal, logical complexity, and syntax nesting complexity will be evaluated respectively. For the cross-modal evaluation, cross-modal refers to the interactions across modules and functions in the code fragment. For example, function calls may be located between different files or modules. Such interactions usually increase the complexity of code analysis because it is necessary to handle the dependencies between modules, trace the data flow and control flow. If the code fragment involves cross-file or cross-function interactions (such as a function calling a function in another file or cross-module data transfer), mark these cross-modal data and trace the data flow from the code structure diagram, record the depth involved in this data, and evaluate the cross-modal complexity of the code fragment. For the evaluation of logical complexity, logical complexity refers to the control flow complexity of the code fragment, especially the complexity of structures such as multiple conditional judgments, recursive calls, and parallel logical paths. In code with high logical complexity, it is judged whether it contains multiple nested conditional judgments, loops, or recursions, etc. Code with high logical complexity often has more complex semantics. By analyzing the semantic structure of the code, understand the number of conditional judgments, the nesting level of loops, as well as the number of recursion levels and nesting depths, etc. By weighting these data, quantify the logical complexity. For syntax nesting complexity, syntax nesting complexity refers to the nesting depth of the syntax structure in the code, such as multiple nested loops, conditional judgments, etc. Code with high nesting complexity is usually more difficult to understand and maintain and is prone to hiding potential errors. By calculating the nesting level of the syntax, such as the loop nesting depth and the level of conditional judgments, evaluate the syntax nesting complexity.After evaluating the complexity of the code, the cross-modal, logical complexity, and syntactic nesting complexity of each node in the code structure diagram are judged. If any one of them is greater than or equal to the corresponding threshold, it indicates that the code segment corresponding to the node is relatively complex. At this time, the node will be regarded as a graph node that needs to be gradually layered. These graph nodes are the parts of the code that require more in-depth analysis and may need higher-level inspections due to high complexity, cross-module interactions, or other factors. For these nodes, gate constraints are set through the analysis of complexity elements and complexity coefficients. The gate constraints are used to control the depth and scope of the analysis to ensure that the most important code parts can be focused on. Finally, through these set gate constraints, targeted analysis is performed on each part of the code to ensure the accuracy and efficiency of the detection process.
[0030] Furthermore, the present application provides a method for setting the gate constraints for the graph nodes, including:
[0031] Setting the detection layering dimension according to the complexity elements; setting the layering fragment granularity according to the complexity coefficient; performing setting and coupling based on the detection layering dimension and the layering fragment granularity for the graph nodes to determine the node gate constraints; and marking the code structure diagram according to the node gate constraints.
[0032] Optionally, when evaluating the complexity of code snippets, it is first necessary to set the hierarchical dimensions of detection through complexity factors, which include but are not limited to cross-modal, logical complexity, and syntactic nesting complexity. Based on these complexity factors, multiple detection hierarchical dimensions are defined, such as the high-level abstraction layer, the middle-level functional layer, and the low-level syntax layer. Among them, the high-level abstraction layer focuses on the logical flow of the entire code, the functions and structures of the main modules; the middle-level functional layer focuses on the control flow and data flow relationships within specific function blocks or modules; the low-level syntax layer delves into syntactic details, function calls, and the use of local variables, etc. Subsequently, the complexity factors of each graph node are weighted and summed to obtain a complexity coefficient. According to the calculated complexity coefficient, the hierarchical fragment granularity of different code segments is set. For nodes with a higher complexity coefficient, small-grain analysis is set, that is, these nodes are further split into smaller analysis units. For example, nested functions, complex loops, etc. are inspected step by step with a finer granularity. For relatively simple nodes, large-grain analysis is used, that is, some redundant checks are skipped to quickly locate potential problems. The specific setting can be determined according to the complexity-granularity mapping table. After obtaining the detection hierarchical dimensions and hierarchical fragment granularity, the detection hierarchical dimensions and hierarchical fragment granularity are coupled according to the graph nodes, that is, those belonging to the same graph node are added to a set and set as the node gating constraint of the graph node to ensure that different-grain detections can be performed on nodes with different complexities in the subsequent process, and potential problems can be accurately captured. Finally, the specific positions of each graph node are matched in the code structure diagram, and the node gating constraint is marked at the corresponding positions. Through this marking, the system can effectively control the analysis scope and depth of each node according to this marked information during the subsequent detection process, ensuring the accuracy and efficiency of code analysis.
[0033] According to the said detection direction and the gating constraint, perform directional distillation processing on the large language model, take the said code structure diagram as the logical guide, take the said code snippet as the detection target, and execute code analysis detection to determine the code detection result; wherein, the detection steps include first-order reachability detection and second-order vulnerability detection.
[0034] In one embodiment, according to the detection direction and gating constraints, the large language model is subjected to directional distillation processing. The large language model is trained through a large number of samples, usually by training on a large number of public datasets and code snippets in multiple programming languages, and learns extensive language understanding capabilities and code reasoning capabilities. Therefore, the preliminarily trained large model has universality and completeness, can handle multiple languages and a wide range of code scenarios, and has powerful capabilities such as grammar understanding, semantic analysis, and problem reasoning. However, this generality cannot fully meet the requirements of specific programming languages or project needs. When faced with a specific programming language (such as Python, Java, C++ etc.) or the code of a certain project, the model may have deficiencies, such as inaccurate understanding of certain special syntax, libraries, frameworks, or business logics. Therefore, the large model is fine-tuned or screened through distillation training to make the model adapt to the requirements of specific tasks and programming languages. Directional distillation refers to fine-tuning the pre-trained large language model to make it adapt to the requirements of specific tasks. In this process, according to the detection dimensions and gating constraints specified by the detection direction, the large language model is optimized to ensure that the model can analyze the code more effectively. Subsequently, taking the code structure diagram as the logical guide, the large language model is based on this structure diagram, focuses on the key areas in the code, and gradually executes the analysis and detection tasks to identify potential problems. This enables the large language model to accurately identify the nodes that may hide problems when understanding the overall structure and relationships of the code, ensuring the comprehensiveness of the analysis. When performing the analysis and detection, the code snippet is used as the detection target, and the code analysis and detection are performed through the large language model. The detection process includes first-order reachability detection and second-order vulnerability detection. First-order reachability detection is to check the execution reachability between different parts of the code to ensure that the code path can be correctly executed as expected. If a certain path cannot be accessed, it will be marked. Second-order vulnerability detection is to perform deeper detection to identify potential risks such as logical vulnerabilities, data leakage, and security hazards. Through the above progressive analysis, it can be ensured that all potential problems can be captured, thereby improving the quality and security of the code and ensuring the high adaptability of the large language model in a specific programming environment.
[0035] Furthermore, the present application provides directional distillation processing for the large language model, including:
[0036] According to the detection direction and the gating constraints, screen the architecture and logic of the large language model to determine the detection requirement orientation; according to the detection requirement orientation, perform directional distillation training on the large language model to determine the lightweight model.
[0037] Preferably, the large language model is analyzed according to the defined detection direction and gating constraints. The detection direction refers to the direction of analyzing the code at different levels, such as syntax, semantics, data flow, control flow, etc. The gating constraints are the limits on the depth and complexity of the analysis process, which are used to accurately control which parts of the code need deeper analysis and which parts can be skipped. The detection direction and gating constraints are mapped to the architecture and logic of the large language model, and the specific functions and modules in the model are determined to ensure that the model can give priority to the code parts related to the detection task when processing specific code fragments. For example, if a task focuses on detecting control flow problems, the model will pay more attention to the control structure and ignore the problems at the grammatical level. Through this screening, the architecture and logic of the model will be adjusted to the detection demand orientation, that is, according to the actual needs, the processing focus and analysis direction of the model are determined, thereby providing a clear direction and goal for the subsequent distillation training. After the detection demand orientation is screened out, the large language model is subjected to directional distillation training based on the orientation. Distillation training refers to fine-tuning the pre-trained large language model through a small amount of specific sample data to make it more adaptable to the needs of specific tasks. During the directional distillation training process, training samples are extracted from a small amount of specific data related to the task according to the code analysis areas that the model needs to focus on (such as logical vulnerabilities, data flow problems, etc.). For example, for the code vulnerability detection task, code snippets containing potential vulnerabilities are selected as the fine-tuning dataset. Subsequently, the pre-trained model is trained using the selected task samples, and the model is adjusted through a very small amount of specific data to optimize the model's weights, making it more suitable for specific task requirements and strengthening the model's capabilities in the target field. Through directional distillation training, unnecessary calculations and redundant parts in the model can be gradually eliminated, retaining functional modules that are highly relevant to the target task. This process achieves lightweight by reducing the amount of calculation and parameters of the model, ensuring that the model can run in a resource-constrained environment while maintaining efficient detection performance. Finally, the model trained with directional distillation is optimized into a lightweight model that not only has high analysis accuracy, but also has better operating efficiency, and can quickly respond to and complete specific code detection tasks.
[0038] Furthermore, the present application provides a method for performing a first-order detection based on code reachability and a second-order detection based on code vulnerabilities on the code fragment according to the lightweight model, with the graph structure of the code structure diagram as the logical framework and the gating constraints as the directional logic requirements, to determine the code detection result.
[0039] Optionally, after obtaining a lightweight model through directional distillation training, use the code structure diagram as the logical framework. This diagram structure shows the relationships between modules, functions, data flows, and control flows in the code, providing clear structural guidance for code analysis. Then, use gating constraints as the directional logical requirements. The gating constraints limit the analysis depth and scope of the model, ensuring that it does not perform redundant comprehensive analysis on all code, but instead focuses on those code parts with higher complexity or greater detection risks. Under the constraint of the directional logical requirements, use the lightweight model to perform first-order reachability detection on the code snippet referred to by the current logical framework, that is, check whether the execution paths between modules in the code are reachable, such as whether the branches of the control flow are correct, and whether functions or conditional statements can be correctly executed. If a certain code path is unreachable, a bad state mark will be made for this code path, indicating that this part is in an unreachable state and further checking for potential vulnerabilities is required. Otherwise, a reachability coefficient will be quantified for it using the lightweight model to evaluate whether the driving effect of this code path meets the requirements. For code paths with bad state marks, second-order vulnerability detection will be performed to deeply analyze more complex vulnerability problems that may exist in the code, such as logical vulnerabilities, potential security risks, memory leaks, data leaks, etc. The second-order detection not only checks the execution paths of the code, but also comprehensively analyzes in combination with data flows, function calls, variable passing, etc. to find possible vulnerabilities or errors, and then merges the results of the second-order detection with the results of the first-order reachability detection to determine the code detection result of the current code snippet. This detection result will form the final report and be provided to developers for code repair or optimization.
[0040] Furthermore, the present application provides a first-order detection based on code reachability, including:
[0041] According to the code structure diagram, perform reachability analysis from bottom to top in a multi-layer progressive manner based on the pointing relationship between the control flow and the data flow; among them, if it is in a reachable state, evaluate the reachability coefficient under the driving effect; if it is in an unreachable state, make a bad state mark.
[0042] Optionally, first, based on the trained lightweight model, a bottom-up reachability analysis is performed on the code structure diagram in a multi-layer progressive manner. The model starts from the smallest unit of the code block, i.e., the smallest function, conditional statement, or loop structure, and conducts a preliminary analysis on it. According to the pointing relationships of the control flow and data flow of these basic units, it is determined whether each unit can reach other modules or logical blocks to ensure that all parts of the code can be correctly executed and interconnected. This way of gradually pushing up and evaluating larger code blocks and modules layer by layer based on the analysis results of the smallest units can ensure that reachability issues from the smallest unit to the entire system can be effectively identified. When performing reachability analysis, if the path of a certain code part is reachable, that is, it can be executed as expected to that part, the lightweight model will further analyze the driving effect of that part. The driving effect refers to whether the actual result after the execution of that part of the code meets the expectations, such as whether it can complete the task on time or whether the set performance requirements are met. To evaluate the driving effect, the model calculates a reachability coefficient based on the execution duration or other performance requirements, and this coefficient reflects the quality and effectiveness of the code execution path. If the reachability coefficient is lower than the reachability threshold, it indicates that the code path has low execution efficiency or cannot be executed (i.e., is in an unreachable state). At this time, a poor state mark is applied to this code path for subsequent targeted vulnerability scanning, thereby improving the quality of the code.
[0043] Furthermore, the present application provides for identifying poor elimination marks, performing vulnerability scanning at the execution layer dimension to determine code vulnerabilities; using the reachability coefficient and the code vulnerabilities as layer detection results; and performing top-down recursive detection based on the inter-layer relationship and the layer detection results to determine the code detection result.
[0044] Optionally, during the analysis process, those graph nodes marked as eliminated will be deeply inspected first. The elimination mark refers to the potential quality problems found in certain code segments during the reachability analysis, such as low efficiency, poor execution effect, and unreachable situations. Based on these marks, vulnerability scanning will be performed in the hierarchical dimension, that is, the code will be analyzed layer by layer for potential vulnerabilities. For example, low-level scanning is to detect basic syntax errors, resource leaks, uninitialized variables, etc.; mid-level scanning is to check for potential vulnerabilities in the data flow, such as unvalidated inputs, incorrect function calls, data out-of-bounds, etc.; high-level scanning is to analyze cross-module and cross-file dependencies and find cross-level logical vulnerabilities, permission issues, etc. In each level, potential vulnerabilities in the code will be gradually identified, including but not limited to logical vulnerabilities (such as incorrect conditional judgments, incorrect loop nesting), security vulnerabilities (such as SQL injection, cross-site scripting attacks), and performance vulnerabilities (such as infinite loops, inefficient algorithms). After the vulnerability scanning is performed, the detection results for each layer will be generated based on the reachability coefficient and code vulnerabilities. The reachability coefficient indicates whether the code path achieves the expected effect, and the code vulnerabilities record the vulnerabilities in each layer. By aggregating the reachability coefficient and the code vulnerabilities of the corresponding layer into a set, the layer detection results are obtained, providing basic data for the subsequent top-down detection. Once the detection results of each layer are determined, the top-down detection will be performed according to the inter-layer relationship between the layers. This process is achieved by gradually pushing the low-level detection results to the upper layer to help the higher-level code detection system understand the overall problem. Specifically, the dependencies and interaction relationships between each layer will be analyzed. For example, a low-level functional module may affect multiple upper-level modules, or the data of multiple upper-level modules depends on the execution results of the lower-level module. By gradually passing the low-level detection results to the upper layer, potential problems in the higher layer can be identified. For example, the performance bottleneck of a low-level module may affect the execution efficiency of the entire application, and a mid-level logical vulnerability may cause the functions of multiple upper-level modules to malfunction. Through recursion, these upper-level errors can be traced and fixed. The recursive process ensures that problems can be gradually discovered and solved from the bottom layer to the upper layer of the code, ultimately obtaining the complete code detection results, and optimizing the code according to the recursive results to eliminate potential vulnerabilities and defects, thereby improving the code quality and execution efficiency.
[0045] In summary, the embodiments of the present application have at least the following technical effects:
[0046] In the embodiments of the present application, a parsing engine is first deployed on a code detection platform to parse the programming paradigm of code snippets, reconstruct a code structure diagram based on the code element-context relationship; subsequently, according to the code structure diagram, a detection direction is defined, and a gating constraint based on multi-layer progressive detection is introduced, wherein the gating constraint is determined based on the detection dimension of the detection direction and the code fragment granularity; then, according to the detection direction and the gating constraint, directional distillation processing is performed on the large language model, taking the code structure diagram as the logical guide and the code snippet as the detection target, and code analysis detection is executed to determine the code detection result; among them, the detection steps include first-order reachability detection and second-order vulnerability detection. These technical effects jointly solve the technical problem that traditional code analysis methods cannot fully consider the code context and multi-level detection dimensions when processing code snippets, resulting in frequent missed detections and false alarms, and realize the technical effects of accurately identifying potential vulnerabilities and optimization points in the code and improving the accuracy and efficiency of code detection by introducing directional distillation processing and gating constraint mechanism based on the large language model.
[0047] Embodiment 2, based on the same inventive concept as the code analysis detection method based on the large language model in the foregoing embodiment, as Figure 2 shown, the present application provides a code analysis detection system based on the large language model, and the system includes: a code parsing module 11: deploying a parsing engine on a code detection platform to parse the programming paradigm of code snippets, and reconstructing a code structure diagram based on the code element-context relationship; a constraint introduction module 12: defining a detection direction according to the code structure diagram, and introducing a gating constraint based on multi-layer progressive detection, wherein the gating constraint is determined based on the detection dimension of the detection direction and the code fragment granularity; a code detection module 13: performing directional distillation processing on the large language model according to the detection direction and the gating constraint, taking the code structure diagram as the logical guide and the code snippet as the detection target, and executing code analysis detection to determine the code detection result; among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
[0048] Further, the code parsing module 11 is further used to execute the following method:
[0049] Connect to a programming database, where the programming database includes a low-code sub-library and a programming language library; perform first-order matching and parsing on the code snippet with the low-code sub-library to determine a first parsing structure; perform second-order matching and parsing on the code snippet with the programming language library to determine a second parsing structure; fuse the first parsing structure and the second parsing structure to reconstruct the code structure diagram.
[0050] Further, the code parsing module 11 is further used to execute the following method:
[0051] Determine the code paradigm framework based on the programming paradigm as the matching basis; according to the code paradigm framework, locate code elements through the parsing engine to determine code entities; according to the code entities, perform context relationship parsing under the constraints of the code paradigm framework through the parsing engine to determine code relationships; determine the parsing structure according to the code paradigm framework, code entities, and code relationships.
[0052] Furthermore, the constraint introduction module 12 is also used to execute the following method:
[0053] Set the detection dimension, where the detection dimension at least includes syntax-semantics-data flow-control flow; according to the code structure diagram, evaluate the complexity of the code snippet, and locate the graph nodes with progressive stratification requirements, where the evaluation elements at least include cross-modal, logical complexity, and syntax nesting complexity;
[0054] For the graph nodes, set the gating constraints.
[0055] Furthermore, the constraint introduction module 12 is also used to execute the following method:
[0056] Set the detection stratification dimension according to the complexity elements; set the stratification fragment granularity according to the complexity coefficient; for the graph nodes, perform setting and coupling based on the detection stratification dimension and the stratification fragment granularity to determine the node gating constraints; mark the code structure diagram according to the node gating constraints.
[0057] Furthermore, the code detection module 13 is also used to execute the following method:
[0058] Screen the architecture and logic of the large language model according to the detection direction and the gating constraints to determine the detection requirement orientation; perform directional distillation training on the large language model according to the detection requirement orientation to determine the lightweight model.
[0059] Furthermore, the code detection module 13 is also used to execute the following method:
[0060] Based on the lightweight model, use the graph structure of the code structure diagram as the logical framework and the gating constraints as the directional logical requirements to perform first-order detection based on code reachability and second-order detection based on code vulnerabilities on the code snippet to determine the code detection result.
[0061] Furthermore, the code detection module 13 is also used to execute the following method:
[0062] According to the code structure diagram, perform reachability analysis from bottom to top in a multi-layer progressive manner based on the pointing relationship between the control flow and the data flow; among them, if it is in a reachable state, evaluate the reachability coefficient under the driving effect; if it is in an unreachable state, perform a bad state mark.
[0063] Furthermore, the code detection module 13 is further configured to execute the following method:
[0064] Identify the elimination mark, perform vulnerability scanning in the hierarchical dimension to determine code vulnerabilities; use the reachability coefficient and the code vulnerabilities as the layer detection results; based on the inter-layer relationship and the layer detection results, perform upper-level recursive detection to determine the code detection results.
[0065] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes a specific embodiment of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired result. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0066] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0067] This specification and the drawings are only exemplary descriptions of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A code analysis and detection method based on a large language model, characterized in that, The method includes: Deploy a parsing engine on the code detection platform to parse the programming paradigm of the code snippet, and reconstruct the code structure diagram with the code element-context relationship; According to the code structure diagram, define the detection direction, and introduce a gating constraint based on multi-layer progressive detection, where the gating constraint is determined by the detection dimension and code fragment granularity based on the detection direction; According to the detection direction and the gating constraint, perform directional distillation processing on the large language model, taking the code structure diagram as the logical guide and the code snippet as the detection target, and execute code analysis detection to determine the code detection result; Among them, the detection steps include first-order reachability detection and second-order vulnerability detection; Introducing a gating constraint based on multi-layer progressive detection includes: Set the detection dimension, where the detection dimension at least includes syntax-semantics-data flow-control flow; According to the code structure diagram, evaluate the complexity of the code snippet, and locate the graph nodes with progressive stratification requirements, where the evaluation elements at least include cross-modal, logical complexity, and syntax nesting complexity; For the graph nodes, set the gating constraint, including: Set the detection stratification dimension according to the complexity elements; Set the stratification fragment granularity according to the complexity coefficient; For the graph nodes, perform setting and coupling based on the detection stratification dimension and the stratification fragment granularity to determine the node gating constraint; Mark the code structure diagram according to the node gating constraint.
2. The code analysis and detection method based on a large language model according to claim 1, characterized in that, Reconstructing the code structure diagram includes: Connect to the programming database, where the programming database includes a low-code sub-library and a programming language library; Use the low-code sub-library to perform first-order matching and parsing on the code snippet to determine the first parsing structure; Use the programming language library to perform second-order matching and parsing on the code snippet to determine the second parsing structure; Fuse the first parsing structure and the second parsing structure to reconstruct the code structure diagram.
3. The code analysis and detection method based on the large language model according to claim 2, wherein Performing matching and parsing on the code snippet includes: Determine the code paradigm framework based on the programming paradigm as the matching basis; According to the code paradigm framework, locate the code elements through the parsing engine to determine the code entities; According to the code entities, perform context relationship parsing under the constraint of the code paradigm framework through the parsing engine to determine the code relationships; Determine the parsing structure according to the code paradigm framework, code entities, and code relationships.
4. The code analysis and detection method based on a large language model according to claim 1, wherein Performing directional distillation processing on the large language model includes: According to the detection direction and the gating constraint, screen the architecture and logic of the large language model to determine the detection requirement orientation; According to the detection requirement orientation, perform directional distillation training on the large language model to determine the lightweight model.
5. The code analysis and detection method based on a large language model according to claim 4, wherein According to the lightweight model, taking the graph structure of the code structure diagram as the logical framework and the gating constraint as the directional logical requirement, perform first-order detection based on code reachability and second-order detection based on code vulnerabilities on the code snippet to determine the code detection result.
6. The code analysis and detection method based on the large language model according to claim 5, wherein Performing first-order detection based on code reachability includes: According to the code structure diagram, from bottom to top under multi-layer progression, perform reachability analysis based on the pointing relationship between the control flow and the data flow; Among them, if it is a reachable state, evaluate the reachability coefficient under the driving effect; If it is an unreachable state, perform a bad state mark.
7. The code analysis and detection method based on a large language model according to claim 6, characterized in that Identify the elimination mark, perform vulnerability scanning under the hierarchical dimension, and determine the code vulnerability; Use the reachability coefficient and the code vulnerability as the layer detection result; Based on the inter-layer relationship and the layer detection result, perform upper-level recursive detection to determine the code detection result.
8. The code analysis and detection system based on the large language model is characterized in that, The system is used to execute the code analysis and detection method based on the large language model according to any one of claims 1-7, including: Code parsing module: Deploy a parsing engine on the code detection platform, parse the programming paradigm of the code fragment, and reconstruct the code structure diagram with the code element-context relationship; Constraint introduction module: According to the code structure diagram, define the detection direction, and introduce the gating constraint based on multi-layer progressive detection, where the gating constraint is determined by the detection dimension and code fragment granularity based on the detection direction; Code detection module: According to the detection direction and the gating constraint, perform directional distillation processing on the large language model, use the code structure diagram as the logical guide, use the code fragment as the detection target, and perform code analysis and detection to determine the code detection result; Among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
Citation Information
Patent Citations
Vulnerability positioning system and method based on deep learning
CN117150507A
Vulnerability detection method based on source code structure and detection system thereof
CN119066667A