Code analysis detection method and system based on large language model
By analyzing and structural reconstruction of code snippets on the code detection platform, and combining with the directional distillation processing of large language models, the problem that traditional code analysis methods cannot fully consider the code context and multi-level detection dimensions are solved, and high-precision and high-efficiency code detection are achieved.
Patent Information
- Application Number
- CN202510533603.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When handling code snippets, traditional code analysis methods cannot fully consider the code context and multi-level detection dimensions, resulting in frequent missed detection and false alarms.
By deploying a parsing engine on the code detection platform, the code fragments are analyzed programmatically, the code structure diagram is reconstructed, and the gated constraints for multi-layer progressive detection are introduced according to the structure diagram, and the large language model is directionally distilled to perform code analysis and detection.
It realizes accurate identification of potential vulnerabilities and optimization points in the code, improves the accuracy and efficiency of code detection, and reduces missed detection and false alarms.
Smart Images

Figure CN120066935A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a code analysis and detection method and system based on large language models. Background Art
[0002] With the rapid expansion of the scale of software systems and the continuous evolution of diverse programming paradigms, the complexity of code and its multi-layer nested structure have increased significantly. Traditional code analysis and detection methods have gradually revealed their limitations. On the one hand, although static analysis tools can identify syntax errors and some structural defects, they often struggle to accurately capture context dependencies when faced with complex code structures across functions, modules, and even languages, resulting in missed or false alarms in detection results. On the other hand, manual review methods not only rely on experience and are inefficient but also difficult to meet the automation requirements in large-scale continuous integration environments. In recent years, large language models have demonstrated powerful context awareness and knowledge transfer capabilities in natural language understanding and generation, providing new possibilities for complex code understanding and semantic reasoning. Therefore, there is an urgent need for an intelligent code detection method that combines the capabilities of large language models, integrates programming semantic structures and multi-layer logical relationships to meet the urgent needs of high-precision, wide-coverage, and high-efficiency code analysis in current software development. Summary of the Invention
[0003] This application provides a code analysis and detection method and system based on large language models, aiming to solve the technical problem that traditional code analysis methods cannot fully consider code context and multi-level detection dimensions when processing code fragments, resulting in frequent missed detections and false alarms.
[0004] In the first aspect disclosed in this application, a code analysis and detection method based on large language models is provided. The method includes: deploying a parsing engine on a code detection platform to parse the programming paradigm of a code fragment, and reconstructing a code structure diagram based on the code element-context relationship; defining a detection direction according to the code structure diagram, and introducing a gating constraint based on multi-layer progressive detection, where the gating constraint is determined based on the detection dimension and code fragment granularity of the detection direction; performing directional distillation processing on the large language model according to the detection direction and the gating constraint, using the code structure diagram as the logical guide and the code fragment as the detection target, and performing code analysis and detection to determine the code detection result; where the detection step includes first-order reachability detection and second-order vulnerability detection.
[0005] Another aspect disclosed in this application provides a code analysis and detection system based on a large language model. The system includes: a code parsing module: deploying a parsing engine on a code detection platform to parse the programming paradigm of code snippets and reconstruct a code structure diagram based on the code element-context relationship; a constraint introduction module: defining a detection direction according to the code structure diagram and introducing a gated constraint based on multi-level progressive detection, where the gated constraint is determined based on the detection dimension and code fragment granularity in the detection direction; a code detection module: performing directional distillation processing on the large language model according to the detection direction and the gated constraint, using the code structure diagram as the logical guide and the code snippet as the detection target, and executing code analysis and detection to determine the code detection result; where the detection steps include first-order reachability detection and second-order vulnerability detection.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: The above-mentioned code analysis and detection method based on a large language model first deploys a parsing engine on a code detection platform to parse the programming paradigm of code snippets and reconstructs a code structure diagram by analyzing the relationship between code elements and context. Subsequently, according to this structure diagram, a detection direction is defined, and a gated constraint based on multi-level progressive detection is introduced. The setting of the gated constraint is based on the detection dimension and the granularity of code fragments. Then, the detection direction and the gated constraint are used to perform directional distillation on the large language model to optimize the model to meet the requirements of code analysis and detection. Finally, guided by the code structure diagram and combined with the code snippet, an analysis including first-order reachability detection and second-order vulnerability detection is performed to determine the code detection result, so as to accurately identify potential vulnerabilities and optimization points in the code and improve the accuracy and efficiency of code detection.
[0007] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically presents the specific embodiments of this application. Brief Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0009] Figure 1 It is a schematic flow chart of a code analysis and detection method based on a large language model in an embodiment.
[0010] Figure 2It is an architecture diagram of a code analysis and detection system based on a large language model in an embodiment.
[0011] Explanation of reference numerals: Code parsing module 11, Constraint introduction module 12, Code detection module 13. Detailed implementation manners
[0012] In the embodiments of the present application, by providing a code analysis and detection method and system based on a large language model, the technical problem that traditional code analysis methods cannot fully consider code context and multi-level detection dimensions when processing code fragments, resulting in frequent missed detections and false alarms, is solved.
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0014] It should be noted that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0015] Embodiment 1, as Figure 1 shown, the present application provides a code analysis and detection method based on a large language model, and the method includes: Deploy a parsing engine on the code detection platform to parse the programming paradigm of the code fragment, and reconstruct the code structure diagram based on the code element-context relationship.
[0016] In the embodiments of the present application, on the code detection platform, a parsing engine is deployed. This engine is responsible for parsing the programming paradigm of the input code fragment. By analyzing the basic elements of the code (such as variables, functions, classes, etc.) and their context relationships, the parsing engine can understand the overall structure and relationships of the code, and determine the parsing structure. Finally, based on the determined parsing structure, the code structure diagram is reconstructed. This structure diagram will show the various components of the code and their mutual relationships. This process helps to comprehensively understand the architecture and semantics of the code, laying a foundation for subsequent code analysis and detection.
[0017] Further, the present application provides a reconstructed code structure diagram, including: Connect to the programming database, where the programming database includes a low-code sub-library and a programming language library; use the low-code sub-library to perform first-order matching and parsing on the code snippet to determine the first parsing structure; use the programming language library to perform second-order matching and parsing on the code snippet to determine the second parsing structure; fuse the first parsing structure and the second parsing structure to reconstruct the code structure diagram.
[0018] Preferably, first, establish a connection with the programming database built into the code detection platform. This built-in programming database includes a low-code sub-library and a programming language library. Among them, the low-code sub-library is oriented to general business logic and is used to process standardized and componentized code patterns; the programming language library covers various programming syntaxes and structural rules and is used to adapt to multiple development languages. Subsequently, use the low-code sub-library to perform first-order matching and parsing on the code snippet to be analyzed. This step mainly identifies whether there are structures similar to common business templates or standard components in the code and generates the corresponding first parsing structure, such as basic frameworks for form control calls and process node controls. Then, use the programming language library to perform second-order matching and parsing on the same code snippet, focusing on parsing details such as specific syntax implementation, function logic, and nested calls to form the second parsing structure, which represents the underlying syntax and logical organization method of the code, such as variable definition, function call, control statement, and scope. Then, based on the unique identifier of the code snippet (such as function name, variable name, or statement position information), establish a mapping relationship between the first parsing structure and the second parsing structure to ensure that the parsing results at different granularities can be uniformly mapped to the same code area. Then, through the graph fusion algorithm, align the business semantic nodes in the first parsing structure with the syntax logic nodes in the second parsing structure and merge the boundaries. For example, for two nodes representing the same code function, merge them and label their dual semantics (such as business logic + syntax call); for complementary nodes, establish cross-layer connections (such as the call edge between a low-code control and its underlying event binding function). After completing the node fusion, generate a complete code structure diagram according to the semantic dependency relationship and control / data flow logic. Each node in the diagram not only contains structural information but also is associated with context semantics, execution path, and other labels, which can be used as a unified logical guiding diagram for code understanding, problem location, and vulnerability detection, providing an accurate semantic basis for subsequent detection.
[0019] Furthermore, the present application provides matching and parsing of the code snippet, including: Determine the code paradigm framework based on the programming paradigm as the matching basis; according to the code paradigm framework, locate the code elements through the parsing engine to determine the code entities; according to the code entities, perform context relationship parsing under the constraints of the code paradigm framework through the parsing engine to determine the code relationships; determine the parsing structure according to the code paradigm framework, code entities, and code relationships.
[0020] Optionally, when matching and parsing code snippets, whether it is first-order or second-order matching and parsing, it is necessary to determine the code paradigm framework, code entities, and code relationships based on the programming paradigm and parsing engine. Taking the first-order matching and parsing of code snippets based on the low-code sub-library as an example, after accessing the low-code sub-library, first use the built-in low-code programming paradigms in it as the matching basis. For example, form-driven paradigm, event response paradigm, process logic paradigm, etc., to identify the code paradigm framework corresponding to the code to be detected. This identification process usually depends on key module names, component call syntax, and their business semantic tags, etc. By keyword matching and syntax template comparison, determine the low-code paradigm category to which the code segment belongs. After the code paradigm framework is determined, based on the rules of this code paradigm, use the parsing engine to perform structural positioning on the code snippet, identify the key code elements in it, such as form fields, binding logic, control components, trigger conditions, interface references, etc., and map these elements to structured code entities, each entity is marked with its function, type, and position in the paradigm structure. Subsequently, use the identified code entities to perform context relationship parsing under the structural constraint rules of the selected paradigm framework. In this process, the parsing engine will analyze the explicit or implicit connection relationships between entities, such as "field and control binding", "button and event triggering", "data source and interface call mapping", etc., so as to determine the complete code relationship, and attach the corresponding business semantics or call path information to each pair of relationships. Finally, integrate the extracted code paradigm framework, the set of code entities, and the context relationships between entities, and organize them into a graphical structure or a tree structure, so as to construct the first parsing structure of the code snippet. This structure not only reflects the module attribution and functional logic of the code in the low-code system, but also provides a standardized input basis for subsequent structural fusion, gating constraint introduction, and model analysis.
[0021] According to the code structure diagram, define the detection direction, and introduce gating constraints based on multi-layer progressive detection, where the gating constraints are determined based on the detection dimension and code fragment granularity of the detection direction.
[0022] In one embodiment, first, the reconstructed code structure diagram is analyzed. According to the information shown in the code structure diagram and the actual business requirements, determine which dimensions in the code need to be key detected, and define the detection direction of the code based on this to ensure the comprehensiveness and accuracy of the detection. Among them, the code structure diagram shows information such as the logical structure of the code, module relationships, data flow, and control flow; the detection direction refers to the analysis direction or dimension in the code detection process, which can be different detection levels in the code, such as the syntax level, semantic level, data flow level, etc. Subsequently, introduce the gating constraint based on multi-layer progressive detection. The gating constraint is to gradually filter unnecessary consumption of computing resources between different detection stages and only perform a more in-depth analysis on the parts that may have problems. By setting the gating constraint, the analysis intensity of the code can be flexibly adjusted according to different dimensions of the detection direction. For example, if the complexity of a certain module in the code is high (such as deeply nested loops or recursive functions), different detection granularities, that is, code fragment granularities, will be set according to the complexity, so as to break down the complex parts into smaller units for analysis. In this way, the analysis of code fragments can be gradually deepened, thereby improving efficiency and reducing false alarms. Generally speaking, the determined detection direction defines the direction and key points of the analysis, and the determined gating constraint controls the depth and scope of the analysis through a multi-layer progressive strategy. These two pieces of data can make the detection process both comprehensive and efficient.
[0023] Furthermore, the present application provides a gating constraint introduced under multi-layer progressive detection, including: Set the detection dimension, where the detection dimension at least includes syntax - semantics - data flow - control flow; according to the code structure diagram, evaluate the complexity of the code fragment, and locate the graph nodes with progressive hierarchical requirements, where the evaluation elements at least include cross-modal, logical complexity, and syntax nesting complexity; for the graph nodes, set the gating constraint.
[0024] Preferably, according to the information shown in the code structure diagram and the actual business requirements, determine the detection dimensions. These dimensions define different levels of code analysis and detection, including at least four dimensions: syntax, semantics, data flow, and control flow. Among them, the syntax dimension mainly checks the syntax structure of the code and can identify common syntax errors, including basic syntax problems such as identifier mismatches and unbalanced parentheses; the semantics dimension focuses on verifying the logical and semantic relationships of the code to ensure that the code logically conforms to the expected design and avoid potential logical loopholes or semantic inconsistencies; the data flow dimension pays attention to the data flow in the code to ensure that there are no errors in the data transfer process and that the definition and use of variables conform to the expectations. For example, check whether variables are correctly initialized or whether there are unused redundant variables; the control flow dimension focuses on analyzing the control structure of the program, including key logics such as conditional statements and loop structures, to ensure that there are no infinite loops, invalid execution paths, or unnecessary branch judgments in the control flow logic. Subsequently, starting from the detection dimensions, evaluate the complexity of the code structure diagram according to the pre-set evaluation factors. Among them, the evaluation factors include cross-modal, logical complexity, and syntax nesting complexity. When performing cross-modal evaluation on the four dimensions of syntax, semantics, data flow, and control flow, evaluations will be carried out in three aspects: cross-modal, logical complexity, and syntax nesting complexity. For the cross-modal evaluation, cross-modal refers to the interactions across modules and functions in the code fragment. For example, function calls may be located between different files or modules. Such interactions usually increase the complexity of code analysis because it is necessary to handle dependencies between modules, track data flow and control flow. If the code fragment involves cross-file or cross-function interactions (such as a function calling a function in another file or cross-module data transfer), mark these cross-modal data and track the data flow from the code structure diagram, record the depth involved by this data, and evaluate the cross-modal complexity of the code fragment. For the evaluation of logical complexity, logical complexity refers to the control flow complexity of the code fragment, especially the complexity of structures such as multiple conditional judgments, recursive calls, and parallel logical paths. In code with high logical complexity, it is judged whether it contains multiple nested conditional judgments, loops, or recursions, etc. Code with high logical complexity often has more complex semantics. By analyzing the semantic structure of the code, understanding the number of conditional judgments and the nesting level of loops, as well as the number of recursion levels and nesting depth, etc., and weighting these data, the logical complexity is quantified. For syntax nesting complexity, syntax nesting complexity refers to the nesting depth of the syntax structure in the code, such as multi-layer nested loops, conditional judgments, etc. Code with high nesting complexity is usually more difficult to understand and maintain and is prone to hiding potential errors. By calculating the nesting level of the syntax, such as the loop nesting depth and the level of conditional judgments, the syntax nesting complexity is evaluated.After evaluating the complexity of the code, the cross-modal, logical complexity, and syntactic nesting complexity of each node in the code structure diagram are judged. If any one of them is greater than or equal to the corresponding threshold, it means that the code segment corresponding to this node is relatively complex. At this time, this node will be used as a graph node that needs to be gradually layered. These graph nodes are the parts of the code that need to be analyzed more deeply and may require higher-level inspections due to high complexity, cross-module interactions, or other factors. For these nodes, gate constraints are set through the analysis of complexity elements and complexity coefficients. The gate constraints are used to control the depth and scope of the analysis to ensure that the most important code parts can be focused on. Finally, through these set gate constraints, targeted analysis is performed on each part of the code to ensure the accuracy and efficiency of the detection process.
[0025] Furthermore, the present application provides a method for setting the gate constraints for the graph nodes, including: Setting the detection layering dimension according to the complexity elements; setting the layering fragment granularity according to the complexity coefficient; performing setting and coupling based on the detection layering dimension and the layering fragment granularity for the graph nodes to determine the node gate constraints; and marking the code structure diagram according to the node gate constraints.
[0026] Optionally, when evaluating the complexity of code snippets, it is first necessary to set the hierarchical dimensions of detection through complexity factors, which include but are not limited to cross-modal, logical complexity, and syntactic nesting complexity. Based on these complexity factors, multiple detection hierarchical dimensions are defined, such as the high-level abstraction layer, the middle-level function layer, and the low-level syntax layer. Among them, the high-level abstraction layer focuses on the logical flow of the entire code, the functions and structures of the main modules; the middle-level function layer focuses on the control flow and data flow relationships within specific function blocks or modules; the low-level syntax layer delves into syntactic details, function calls, and the use of local variables, etc. Subsequently, the complexity factors of each graph node are weighted and summed to obtain a complexity coefficient. Based on the calculated complexity coefficient, the hierarchical fragmentation granularity of different code segments is set. For nodes with a higher complexity coefficient, small-grained analysis will be set, that is, these nodes will be further split into smaller analysis units. For example, more fine-grained step-by-step inspections will be carried out on nested functions, complex loops, etc. For relatively simple nodes, large-grained analysis will be used, that is, some redundant checks will be skipped to quickly locate potential problems. The specific setting can be determined according to the complexity-granularity mapping table. After obtaining the detection hierarchical dimensions and hierarchical fragmentation granularity, the detection hierarchical dimensions and hierarchical fragmentation granularity will be coupled according to the graph nodes, that is, those belonging to the same graph node will be added to a set and set as the node gating constraint of that graph node to ensure that different-grained detections can be performed on nodes with different complexities in the follow-up and potential problems can be accurately captured. Finally, the specific positions of each graph node are matched in the code structure diagram, and the node gating constraint is marked at the corresponding positions. Through this marking, the system can effectively control the analysis scope and depth of each node based on this marking information during the subsequent detection process, ensuring the accuracy and efficiency of code analysis.
[0027] According to the detection direction and the gating constraint, perform directional distillation processing on the large language model. Using the code structure diagram as the logical guide and the code snippet as the detection target, execute code analysis detection to determine the code detection result; among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
[0028] In one embodiment, according to the detection direction and gating constraints, the large language model is subjected to directional distillation processing. Among them, the large language model is trained through a large number of samples, usually by training on a large number of public data sets and code snippets in various programming languages, learning a wide range of language understanding capabilities and code reasoning capabilities. Therefore, the preliminarily trained large model has universality and completeness, can handle multiple languages and a wide range of code scenarios, and has powerful grammar understanding, semantic analysis, problem reasoning and other capabilities. However, this generality cannot fully meet the needs of specific programming languages or project requirements. When faced with a specific programming language (such as Python, Java, C++, etc.) or the code of a certain project, the model may have deficiencies, such as inaccurate understanding of certain special grammars, libraries, frameworks or business logics. Therefore, the large model is fine-tuned or screened through distillation training to make the model adapt to the needs of specific tasks and programming languages. Directional distillation refers to fine-tuning the pre-trained large language model to make it adapt to the needs of specific tasks. In this process, according to the detection dimensions and gating constraints specified by the detection direction, the large language model is optimized to ensure that the model can analyze the code more effectively. Subsequently, taking the code structure diagram as the logical guide, the large language model takes this structure diagram as the basis, focuses on the key areas in the code, and gradually executes the analysis and detection tasks to identify potential problems. This enables the large language model to accurately identify the nodes that may hide problems when understanding the overall structure and relationships of the code, ensuring the comprehensiveness of the analysis. When performing analysis and detection, the code snippet is used as the detection target, and the code analysis and detection are performed through the large language model. The detection process includes first-order reachability detection and second-order vulnerability detection. First-order reachability detection is to check the execution reachability between different parts of the code to ensure that the code path can be correctly executed as expected. If a certain path cannot be accessed, it will be marked. Second-order vulnerability detection is to perform deeper detection to identify potential risks such as logical vulnerabilities, data leaks, and security hazards. Through the above progressive analysis, it can be ensured that all potential problems can be captured, thereby improving the quality and security of the code and ensuring the high adaptability of the large language model in a specific programming environment.
[0029] Further, the present application provides a method for performing directional distillation processing on a large language model, including: According to the detection direction and the gating constraints, screen the architecture and logic of the large language model to determine the detection requirement orientation; according to the detection requirement orientation, perform directional distillation training on the large language model to determine a lightweight model.
[0030] Preferably, the large language model is analyzed according to the defined detection direction and gating constraints. The detection direction refers to the direction of analyzing the code at different levels, such as syntax, semantics, data flow, control flow, etc. The gating constraints are the limits on the depth and complexity of the analysis process, which are used to accurately control which parts of the code need deeper analysis and which parts can be skipped. The detection direction and gating constraints are mapped to the architecture and logic of the large language model, and the specific functions and modules in the model are determined to ensure that the model can give priority to the code parts related to the detection task when processing specific code fragments. For example, if a task focuses on detecting control flow problems, the model will pay more attention to the control structure and ignore the problems at the grammatical level. Through this screening, the architecture and logic of the model will be adjusted to the detection demand orientation, that is, according to the actual needs, the processing focus and analysis direction of the model are determined, thereby providing a clear direction and goal for the subsequent distillation training. After the detection demand orientation is screened out, the large language model is subjected to directional distillation training based on the orientation. Distillation training refers to fine-tuning the pre-trained large language model through a small amount of specific sample data to make it more adaptable to the needs of specific tasks. During the directional distillation training process, training samples are extracted from a small amount of specific data related to the task according to the code analysis areas that the model needs to focus on (such as logical vulnerabilities, data flow problems, etc.). For example, for the code vulnerability detection task, code snippets containing potential vulnerabilities are selected as the fine-tuning dataset. Subsequently, the pre-trained model is trained using the selected task samples, and the model is adjusted through a very small amount of specific data to optimize the model's weights, making it more suitable for specific task requirements and strengthening the model's capabilities in the target field. Through directional distillation training, unnecessary calculations and redundant parts in the model can be gradually eliminated, retaining functional modules that are highly relevant to the target task. This process achieves lightweight by reducing the amount of calculation and parameters of the model, ensuring that the model can run in a resource-constrained environment while maintaining efficient detection performance. Finally, the model trained with directional distillation is optimized into a lightweight model that not only has high analysis accuracy, but also has better operating efficiency, and can quickly respond to and complete specific code detection tasks.
[0031] Furthermore, the present application provides a method for performing a first-order detection based on code reachability and a second-order detection based on code vulnerabilities on the code fragment according to the lightweight model, with the graph structure of the code structure diagram as the logical framework and the gating constraints as the directional logic requirements, to determine the code detection result.
[0032] Optionally, after obtaining the lightweight model through directed distillation training, using the code structure diagram as the logical framework, which shows the relationships between modules, functions, data flows, and control flows in the code, providing clear structural guidance for code analysis. Then, using the gating constraint as the directed logical requirement, the gating constraint limits the analysis depth and scope of the model, ensuring that it does not perform redundant comprehensive analysis on all code, but instead focuses on code sections with higher complexity or greater detection risks. Under the constraint of the directed logical requirement, the lightweight model is used to perform first-order reachability detection on the code snippet referred to by the current logical framework, that is, to check whether the execution paths between modules in the code are reachable, such as whether the branches of the control flow are correct, and whether functions or conditional statements can be correctly executed. If a certain code path is unreachable, a bad state mark will be made for this code path, indicating that this part is in an unreachable state and further checking for vulnerabilities is required. Otherwise, the lightweight model will quantify a reachability coefficient for it to evaluate whether the driving effect of this code path meets the requirements. For code paths with bad state marks, second-order vulnerability detection will be performed to deeply analyze more complex vulnerability problems that may exist in the code, such as logical vulnerabilities, potential security risks, memory leaks, data leaks, etc. The second-order detection not only checks the execution paths of the code, but also comprehensively analyzes in combination with data flows, function calls, variable passing, etc. to find possible vulnerabilities or errors, and then merges the results of the second-order detection with the results of the first-order reachability detection to determine the code detection result of the current code snippet. This detection result will form the final report and be provided to the developer for code repair or optimization.
[0033] Furthermore, the present application provides a first-order detection based on code reachability, including: According to the code structure diagram, from bottom to top in a multi-layer progressive manner, perform reachability analysis based on the pointing relationship between the control flow and the data flow; among them, if it is in a reachable state, evaluate the reachability coefficient under the driving effect; if it is in an unreachable state, make a bad state mark.
[0034] Optionally, first, perform a bottom-up reachability analysis on the code structure diagram in a multi-layer progressive manner according to the trained lightweight model. The model starts from the smallest unit of the code block, that is, the smallest function, conditional statement, or loop structure, and conducts a preliminary analysis on it. According to the pointing relationships of the control flow and data flow of these basic units, determine whether each unit can reach other modules or logical blocks to ensure that all parts of the code can be correctly executed and interconnected. This way of gradually pushing up and evaluating larger code blocks and modules layer by layer based on the analysis results of the smallest units can ensure that reachability issues from the smallest unit to the entire system can be effectively identified. When performing reachability analysis, if the path of a certain code part is reachable, that is, it can be executed as expected to that part, the lightweight model will further analyze the driving effect of that part. The driving effect refers to whether the actual result after the execution of that part of the code meets the expectations, such as whether the task can be completed on time or whether the set performance requirements are met. To evaluate the driving effect, the model will calculate a reachability coefficient based on the execution duration or other performance requirements, and this coefficient reflects the quality and effectiveness of the code execution path. If the reachability coefficient is lower than the reachability threshold, it indicates that the code path has low execution efficiency or cannot be executed (i.e., in an unreachable state). At this time, a poor state mark will be applied to this code path for subsequent targeted vulnerability scanning, thereby improving the quality of the code.
[0035] Furthermore, the present application provides for identifying inferior elimination marks, performing vulnerability scanning at the execution layer dimension to determine code vulnerabilities; using the reachability coefficient and the code vulnerabilities as layer detection results; and performing an upper-level recursive detection based on the inter-layer relationship and the layer detection results to determine the code detection result.
[0036] Optionally, during the analysis process, those graph nodes marked as eliminated will be deeply inspected first. The elimination mark refers to the situation where, during the reachability analysis, potential quality problems are found in certain code segments, such as low efficiency, poor execution effect, and unreachability. Based on these marks, vulnerability scanning will be carried out in the hierarchical dimension, that is, the code will be analyzed layer by layer for possible vulnerabilities. For example, the low-level scan is to detect basic syntax errors, resource leaks, uninitialized variables, etc.; the middle-level scan is to check for potential vulnerabilities in the data flow, such as unvalidated inputs, incorrect function calls, data out-of-bounds, etc.; the high-level scan is to analyze cross-module and cross-file dependencies and find cross-level logical vulnerabilities, permission issues, etc. In each level, potential vulnerabilities in the code will be gradually identified, including but not limited to logical vulnerabilities (such as incorrect conditional judgments, incorrect loop nesting), security vulnerabilities (such as SQL injection, cross-site scripting attacks), and performance vulnerabilities (such as infinite loops, inefficient algorithms). After the vulnerability scan is executed, the detection results for each level will be generated based on the reachability coefficient and code vulnerabilities. The reachability coefficient indicates whether the code path achieves the expected effect, and the code vulnerabilities record the vulnerabilities in each level. By aggregating the reachability coefficient and the code vulnerabilities of the corresponding level into a set, the layer detection results are obtained, providing basic data for the subsequent top-down detection. Once the detection results of each level are determined, the top-down detection will be executed according to the inter-layer relationships between the levels. This process is achieved by gradually pushing the low-level detection results to the upper layer to help the higher-level code detection system understand the overall problem. Specifically, the dependencies and interaction relationships between each level will be analyzed. For example, a low-level functional module may affect multiple upper-level modules, or the data of multiple upper-level modules depends on the execution results of the lower-level module. By gradually passing the low-level detection results to the upper layer, potential problems in the higher level can be identified. For example, the performance bottleneck of a low-level module may affect the execution efficiency of the entire application, and a logical vulnerability in the middle level may cause the functions of multiple upper-level modules to malfunction. Through recursion, these upper-level errors can be traced and fixed. The recursive process ensures that problems can be gradually discovered and solved from the bottom layer to the upper layer of the code, ultimately obtaining the complete code detection results, and optimizing the code according to the recursive results to eliminate potential vulnerabilities and defects, thereby improving the code quality and execution efficiency.
[0037] In summary, the embodiments of the present application at least have the following technical effects: In the embodiments of the present application, a parsing engine is first deployed on a code detection platform to parse the programming paradigm of code snippets, and reconstruct a code structure diagram based on the code element-context relationship. Subsequently, according to the code structure diagram, a detection direction is defined, and a gating constraint based on multi-layer progressive detection is introduced, where the gating constraint is determined based on the detection dimension and code fragment granularity of the detection direction. Then, according to the detection direction and the gating constraint, directional distillation processing is performed on the large language model, using the code structure diagram as the logical guide and the code snippet as the detection target, and code analysis and detection are executed to determine the code detection result. Among them, the detection steps include first-order reachability detection and second-order vulnerability detection. These technical effects jointly solve the technical problem that traditional code analysis methods cannot fully consider the code context and multi-level detection dimensions when processing code snippets, resulting in frequent missed detections and false alarms, and achieve the technical effects of accurately identifying potential vulnerabilities and optimization points in the code and improving the accuracy and efficiency of code detection by introducing directional distillation processing and gating constraint mechanism based on the large language model.
[0038] Embodiment 2, based on the same inventive concept as the code analysis and detection method based on the large language model in the foregoing embodiment, as Figure 2 shown, the present application provides a code analysis and detection system based on the large language model. The system includes: a code parsing module 11: deploying a parsing engine on a code detection platform to parse the programming paradigm of code snippets, and reconstructing a code structure diagram based on the code element-context relationship; a constraint introduction module 12: defining a detection direction according to the code structure diagram, and introducing a gating constraint based on multi-layer progressive detection, where the gating constraint is determined based on the detection dimension and code fragment granularity of the detection direction; a code detection module 13: performing directional distillation processing on the large language model according to the detection direction and the gating constraint, using the code structure diagram as the logical guide and the code snippet as the detection target, and executing code analysis and detection to determine the code detection result. Among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
[0039] Furthermore, the code parsing module 11 is further configured to execute the following method: Connect to a programming database, where the programming database includes a low-code sub-library and a programming language library; perform first-order matching and parsing on the code snippet using the low-code sub-library to determine a first parsing structure; perform second-order matching and parsing on the code snippet using the programming language library to determine a second parsing structure; fuse the first parsing structure and the second parsing structure to reconstruct the code structure diagram.
[0040] Furthermore, the code parsing module 11 is further configured to execute the following method: Determine the code paradigm framework based on the programming paradigm as the matching basis; according to the code paradigm framework, locate code elements through the parsing engine to determine code entities; according to the code entities, perform context relationship parsing under the constraints of the code paradigm framework through the parsing engine to determine code relationships; according to the code paradigm framework, code entities and code relationships, determine the parsing structure.
[0041] Further, the constraint introduction module 12 is further configured to execute the following method: Set the detection dimension, where the detection dimension at least includes syntax-semantics-data flow-control flow; according to the code structure diagram, evaluate the complexity of the code fragment, and locate the graph nodes with progressive stratification requirements, where the evaluation elements at least include cross-modal, logical complexity, and syntax nesting complexity; Set the gating constraint for the graph node.
[0042] Further, the constraint introduction module 12 is further configured to execute the following method: Set the detection stratification dimension according to the complexity elements; set the stratification fragment granularity according to the complexity coefficient; perform setting and coupling based on the detection stratification dimension and the stratification fragment granularity for the graph node to determine the node gating constraint; mark the code structure diagram according to the node gating constraint.
[0043] Further, the code detection module 13 is further configured to execute the following method: Screen the architecture and logic of the large language model according to the detection direction and the gating constraint to determine the detection requirement orientation; perform directed distillation training on the large language model according to the detection requirement orientation to determine the lightweight model.
[0044] Further, the code detection module 13 is further configured to execute the following method: According to the lightweight model, use the graph structure of the code structure diagram as the logical framework and the gating constraint as the directed logical requirement to perform first-order detection based on code reachability and second-order detection based on code vulnerabilities on the code fragment to determine the code detection result.
[0045] Further, the code detection module 13 is further configured to execute the following method: According to the code structure diagram, perform reachability analysis from bottom to top under multi-layer progression, based on the pointing relationship between the control flow and the data flow; where, if it is in a reachable state, evaluate the reachability coefficient under the driving effect; if it is in an unreachable state, perform a bad state mark.
[0046] Further, the code detection module 13 is further configured to execute the following method: Identify the elimination marks, perform vulnerability scanning at the hierarchical dimension, and determine the code vulnerabilities; use the reachability coefficient and the code vulnerabilities as the layer detection results; use the inter-layer relationship and the layer detection results to perform upper-level recursive detection to determine the code detection results.
[0047] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0048] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0049] This specification and the drawings are only exemplary descriptions of the present application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A code analysis and detection method based on a large language model, characterized in that: The method comprises: Deploy a parsing engine on the code detection platform to parse the code snippets for programming paradigms and reconstruct the code structure diagram based on the code element-context relationship; According to the code structure diagram, a detection direction is defined, and a gating constraint based on multi-layer progressive detection is introduced, wherein the gating constraint is determined by a detection dimension based on the detection direction and a code fragment granularity; According to the detection direction and the gating constraint, a directional distillation process is performed on the large language model, the code structure diagram is used as a logical guide, the code snippet is used as a detection target, and code analysis detection is performed to determine a code detection result; Among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
2. The code analysis and detection method based on a large language model as claimed in claim 1, characterized in that: Refactor the code structure diagram, including: Connecting to a programming database, wherein the programming database includes a low-code sub-library and a programming language library; Using the low-code sub-library, performing first-order matching and parsing on the code snippet to determine a first parsing structure; Using the programming language library, performing second-order matching and parsing on the code snippet to determine a second parsing structure; The first parsing structure and the second parsing structure are integrated to reconstruct a code structure graph.
3. The code analysis and detection method based on a large language model as claimed in claim 2, characterized in that: Matching and parsing the code snippet includes: Based on the matching of programming paradigm, determine the code paradigm framework; According to the code paradigm framework, the code elements are located and the code entities are determined by the parsing engine; According to the code entity, the parsing engine performs contextual relationship parsing under the constraints of the code paradigm framework to determine the code relationship; The parsing structure is determined according to the code paradigm framework, code entities and code relationships.
4. The code analysis and detection method based on a large language model as claimed in claim 1, characterized in that: Introduce gating constraints based on multi-layer progressive detection, including: Setting detection dimensions, wherein the detection dimensions at least include syntax-semantics-data flow-control flow; According to the code structure diagram, the code snippet is evaluated for complexity, and the graph nodes with progressive layering requirements are located, wherein the evaluation factors at least include cross-modality, logic complexity, and syntax nesting complexity; The gating constraint is set for the graph node.
5. The code analysis and detection method based on a large language model as claimed in claim 4, characterized in that: For the graph node, setting the gating constraint includes: According to the complexity factors, set the detection layering dimensions; According to the complexity coefficient, set the granularity of layered fragments; For the graph nodes, setting and coupling based on the detection layer dimension and the layer fragment granularity are performed to determine the node gating constraint; The code structure diagram is marked according to the node gating constraint.
6. The code analysis and detection method based on a large language model as claimed in claim 1, characterized in that: Directed distillation of large language models, including: According to the detection direction and the gating constraint, the architecture and logic of the large language model are screened to determine the detection requirement orientation; According to the detection requirement orientation, the large language model is subjected to directional distillation training to determine a lightweight model.
7. The code analysis and detection method based on a large language model as claimed in claim 6, characterized in that: According to the lightweight model, taking the graph structure of the code structure diagram as the logical framework and the gating constraints as the directional logic requirements, the code fragment is subjected to a first-order detection based on code reachability and a second-order detection based on code vulnerabilities to determine the code detection result.
8. The code analysis and detection method based on a large language model as claimed in claim 7, characterized in that: Perform first-order detection based on code reachability, including: According to the code structure diagram, reachability analysis is performed from bottom to top in a multi-layer progressive manner based on the directional relationship between control flow and data flow; Among them, if it is a reachable state, the reachability coefficient under the driving effect is evaluated; If it is in an unreachable state, mark it as inferior.
9. The code analysis and detection method based on a large language model as claimed in claim 8, characterized in that: Identify inferior elimination markers, perform vulnerability scans in hierarchical dimensions, and identify code vulnerabilities; Using the reachability coefficient and the code vulnerability as a layer detection result; Based on the inter-layer relationship and the layer detection result, upper recursive detection is performed to determine the code detection result.
10. A code analysis and detection system based on a large language model, characterized in that: The system is used to execute the code analysis and detection method based on a large language model according to any one of claims 1 to 9, comprising: Code parsing module: deploys a parsing engine on the code detection platform to parse the code snippets for programming paradigms and reconstructs the code structure diagram based on the code element-context relationship; Constraint introduction module: defines the detection direction according to the code structure diagram, and introduces the gating constraints based on multi-layer progressive detection, wherein the gating constraints are determined by the detection dimension based on the detection direction and the code fragment granularity; Code detection module: according to the detection direction and the gating constraint, the large language model is subjected to directional distillation processing, the code structure diagram is used as a logical guide, the code snippet is used as a detection target, and code analysis detection is performed to determine the code detection result; Among them, the detection steps include first-order reachability detection and second-order vulnerability detection.
Citation Information
Patent Citations
Vulnerability positioning system and method based on deep learning
CN117150507A
Vulnerability detection method based on source code structure and detection system thereof
CN119066667A
Source code detection method
CN119442240A
Python code vulnerability detection method and system based on deep learning
CN119885195A
Methods and systems for identifying control flow patterns and dataflow constraints in software code to detect software anomalies
US12259805B1
Cited By
Automatic code auditing method and device, computer equipment and storage medium
CN121009550A