Code analysis method, device, storage medium and program product

CN122838245APending Publication Date: 2026-09-29ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371880.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但是,随着软件开发技术的发展,单个应用程序的功能越来越复杂,其对应的源代码的数量也越来越多,源代码涉及的开发语言的类型也越来越多,通过人工的方式对应用程序的源代码进行代码分析变得日益困难

Benefits of technology

[0024]本申请实施例提供的代码分析方案中,针对某一应用程序,首先,获取应用程序对应的程序功能描述信息,以及待分析的程序源代码。之后,根据程序功能描述信息和程序源代码,通过缺陷代码识别大语言模型,对待分析的程序源代码进行整体的代码分析,以确定待分析的程序源代码中存在的潜在缺陷代码,以及潜在缺陷代码对应的代码位置信息和第一缺陷描述信息。为了保证代码分析结果的准确性,进一步地,对潜在缺陷代码进行缺陷检验,具体地,根据潜在缺陷代码的代码位置信息,从待分析的程序源代码中获取包含潜在缺陷代码的目标代码片段,其中,目标代码片段所包含的源代码具有完整的语义信息;然后,根据第一缺陷描述信息和目标代码片段,确定目标代码片段对应的目标形式化模型中存在的目标缺陷位置,其中,目标形式化模型用于反映目标代码片段的代码执行流程和代码状态转换,目标缺陷位置对应于第一缺陷描述信息所描述的缺陷状态;最后,根据目标形式化模型和目标缺陷位置,确定程序源代码中存在的目标缺陷代码。本方案中,先通过缺陷代码识别大语言模型,整体识别待分析的程序源代码中的潜在缺陷代码,再针对潜在缺陷代码对应目标代码片段进行缺陷检验,从而确定程序源代码中真正存在缺陷的目标缺陷代码,通过模型识别加自动化检验的代码分析方式,既能提高应用程序对应的程序源代码的分析效率,又能提升程序源代码中缺陷代码识别的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838245A_ABST
    Figure CN122838245A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a code analysis method, device, storage medium and program product, comprising: obtaining program function description information corresponding to an application program and program source code to be analyzed; identifying code location information and first defect description information corresponding to potential defect code in the program source code as a whole through a defect code recognition large language model; obtaining a target code segment containing the potential defect code from the program source code to be analyzed according to the code location information, and performing defect testing on the target code segment based on the first defect description information, and determining the target defect code in the program source code that actually exists defects according to the target defect position existing in the target formalized model corresponding to the determined target code segment. The code analysis method of model recognition plus automatic testing can improve the analysis efficiency of the program source code corresponding to the application program, and improve the accuracy of defect code recognition in the program source code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a code analysis method, device, storage medium, and program product. Background Technology

[0002] In the field of software development, analyzing application source code, identifying and correcting defects, helps improve the reliability and stability of application operation. However, with the advancement of software development technology, the functionality of individual applications is becoming increasingly complex, resulting in a larger volume of source code and a wider variety of programming languages ​​involved. This makes manual code analysis of application source code increasingly difficult. Summary of the Invention

[0003] This application provides a code analysis method, device, storage medium, and program product to improve the efficiency and accuracy of analyzing application source code.

[0004] In a first aspect, embodiments of this application provide a code analysis method, the method comprising:

[0005] Obtain the program function description information corresponding to the application, as well as the program source code to be analyzed;

[0006] Based on the program function description information and the program source code, the code location information and first defect description information corresponding to the potential defect code in the program source code are determined by using the defect code identification large language model.

[0007] Based on the code location information, a target code segment containing the potential defective code is obtained from the program source code, wherein the source code contained in the target code segment has complete semantic information;

[0008] Based on the first defect description information and the target code fragment, the location of the target defect in the target formal model corresponding to the target code fragment is determined. The target formal model is used to reflect the code execution flow and code state transition of the target code fragment. The location of the target defect corresponds to the defect state described by the first defect description information.

[0009] Based on the formal model of the target and the location of the target defect, the target defect code existing in the program source code is determined.

[0010] Secondly, embodiments of this application provide a code analysis apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire program function description information corresponding to the application, as well as the program source code to be analyzed;

[0012] The processing module is used to determine the code location information and first defect description information corresponding to the potential defect code in the program source code by using the defect code identification large language model based on the program function description information and the program source code;

[0013] The determination module is configured to: obtain a target code segment containing the potential defective code from the program source code based on the code location information, wherein the source code contained in the target code segment has complete semantic information; determine the location of the target defect in the target formal model corresponding to the target code segment based on the first defect description information and the target code segment, wherein the target formal model is used to reflect the code execution flow and code state transition of the target code segment, and the location of the target defect corresponds to the defect state described by the first defect description information; and determine the target defective code existing in the program source code based on the target formal model and the location of the target defect.

[0014] Thirdly, embodiments of this application provide a code analysis method, the method comprising:

[0015] Receive a request triggered by a client device by calling the code analysis service provided by the server device. The request includes program function description information corresponding to the application and the program source code to be analyzed.

[0016] Based on the program function description information and the program source code, the code location information and the first defect description information corresponding to the first potential defect code in the program source code are determined by the defect code identification large language model, and the defect code identification large language model is a large language model.

[0017] Based on the code location information, a target code segment containing the potential defective code is obtained from the program source code, wherein the source code contained in the target code segment has complete semantic information;

[0018] Based on the first defect description information and the target code fragment, the location of the target defect in the target formal model corresponding to the target code fragment is determined. The target formal model is used to reflect the code execution flow and code state transition of the target code fragment. The location of the target defect corresponds to the defect state described by the first defect description information.

[0019] Based on the target formal model and the target defect location, the target defect code existing in the program source code is determined;

[0020] The target defect code is then fed back to the client device for display.

[0021] Fourthly, embodiments of this application provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores a computer program, and when the computer program is executed by the processor, the processor is able to implement at least the code analysis method described in the first or third aspect.

[0022] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor of an electronic device, enables the processor to at least implement the code analysis method as described in the first or third aspect.

[0023] In a sixth aspect, embodiments of this application provide a computer program product, including: a computer program or instructions, which, when executed by a processor of an electronic device, enable the processor to at least implement the code analysis method as described in the first or third aspect.

[0024] In the code analysis scheme provided in this application embodiment, for a specific application, firstly, the program function description information corresponding to the application and the source code of the program to be analyzed are obtained. Then, based on the program function description information and the program source code, a comprehensive code analysis is performed on the source code of the program to be analyzed using a large language model for defect code identification, to determine the potential defective code present in the source code of the program to be analyzed, as well as the code location information and first defect description information corresponding to the potential defective code. To ensure the accuracy of the code analysis results, further defect verification is performed on the potential defective code. Specifically, based on the code location information of the potential defective code, a target code segment containing the potential defective code is obtained from the source code of the program to be analyzed, wherein the source code contained in the target code segment has complete semantic information; then, based on the first defect description information and the target code segment, the location of the target defect in the target formal model corresponding to the target code segment is determined, wherein the target formal model is used to reflect the code execution flow and code state transition of the target code segment, and the location of the target defect corresponds to the defect state described by the first defect description information; finally, based on the target formal model and the location of the target defect, the target defective code present in the program source code is determined. In this solution, a large language model for defect code identification is first used to identify potential defective codes in the source code of the program to be analyzed. Then, defect inspection is performed on the target code segments corresponding to the potential defective codes to determine the target defective codes that actually exist in the source code. By combining model identification with automated inspection, this code analysis method can improve both the efficiency of analyzing the source code of the application and the accuracy of identifying defective codes in the source code. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A schematic diagram of the hardware execution environment for a code analysis method provided in an embodiment of this application;

[0027] Figure 2 This is an application diagram of a cloud computing environment provided in an embodiment of this application;

[0028] Figure 3 A flowchart illustrating a code analysis method provided in this application embodiment;

[0029] Figure 4 A schematic diagram illustrating a code analysis method provided in an embodiment of this application;

[0030] Figure 5 A schematic diagram of an abstract syntax tree provided for the implementation of this application;

[0031] Figure 6 A flowchart illustrating a potential defect code determination process provided in this application embodiment;

[0032] Figure 7 A flowchart illustrating another potential defect code determination process provided for embodiments of this application;

[0033] Figure 8 A flowchart illustrating yet another potential defect code determination process provided in this application embodiment;

[0034] Figure 9 This is a schematic diagram of the structure of a code analysis device provided in an embodiment of this application;

[0035] Figure 10 To and Figure 9 The illustrated embodiment provides a schematic diagram of the electronic device corresponding to the code analysis device. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0038] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0039] The relevant concepts involved in the embodiments of this application will be explained below.

[0040] An application is a computer program or software system designed to meet specific user needs or solve specific problems. Optionally, an application can be a single, independent program or a software system composed of multiple components.

[0041] Application function description information refers to the information provided by the application documentation that details the application's functions and operating mechanisms, including but not limited to: the application's entry function or script, the components included in the application and the functions implemented by those components, the concurrent execution involved in the application (e.g., the execution details of multithreading and asynchronous operations), and the strategy for deploying application information on multiple nodes in multiple copies.

[0042] Application source code refers to the set of programming instructions written by developers using one or more programming languages ​​such as Python, Java, and C++, which describes the application's logic, functionality, and behavior. The physical files used to store the source code are called source files. The source code of different components of the application can be stored in different source files, and multiple different source files together constitute the whole application.

[0043] An Abstract Syntax Tree (AST) is a tree-like representation of the abstract syntactic structure of source code. Each node in an AST represents a structure in the source code, such as variable declarations, function declarations, function definitions, logical expressions, assignment statements, etc.

[0044] Formalization refers to the precise description of a concept, system, or process using rigorous mathematical language, symbols, and logical rules, employing mathematical and logical methods. In software development, formal methods can be used to formally model the code execution flow and state transitions of application source code, thereby verifying the correctness of the source code. Formal models can be constructed using methods such as Petri nets, Pi calculus, and the Unified Modeling Language (UML).

[0045] Large Language Models (LLMs) are large-scale parametric language models learned using deep learning frameworks and large-scale data corpora. They can be used to handle various natural language tasks such as text classification, question answering, and dialogue. In practice, LLMs understand user intent based on input prompts and output predictions that match that intent.

[0046] Defect code identification large language model, which is a large language model used to identify defective code in source code and determine the code location information and defect description information of defective code.

[0047] Formal modeling of large language models refers to the formal modeling of the code execution flow and code state transitions in source code to generate a large language model corresponding to the formal model of the source code.

[0048] Optionally, in the embodiments of this application, the defect code identification large language model and the formal modeling large language model can be implemented as the same or different large language models. Different large language models can have different parameter scales. In specific implementation, the appropriate model can be flexibly selected based on actual needs.

[0049] Defective code refers to code in the source code of an application that affects the correct and stable operation of the application. Optionally, in this embodiment, defective code includes, but is not limited to, code exhibiting any of the following defect types: syntax errors, logical errors, runtime errors, boundary condition errors, uninitialized variables, resource leaks, concurrency issues, performance issues, security vulnerabilities, maintainability issues, and dependency issues.

[0050] In the field of software development, analyzing the source code of an application, identifying and correcting defective code is an important task in the software development process. This helps improve the quality of the source code and enhance the reliability and stability of the application.

[0051] However, with the development of software development technology, applications are constantly iterating and updating, the functions of individual applications are becoming more and more complex, the amount of their corresponding source code is also increasing, and the types of development languages ​​involved in the source code are also increasing, making it increasingly difficult to manually analyze the source code of applications.

[0052] To address the challenge of manually analyzing source code, this application provides a solution. The basic idea is as follows: First, obtain the program function description information corresponding to the application and the source code to be analyzed. Then, using a large language model for defect code identification, identify potential defective codes in the source code as a whole, and determine the code location information and first defect description information corresponding to the potential defective codes. Next, based on the code location information, obtain the target code segment containing the potential defective code from the source code to be analyzed. Using traditional formal modeling verification methods, perform defect verification on the target code segment based on the first defect description information to identify the target defective code that actually exists in the source code. This solution, through a code analysis approach combining model recognition and automated verification, can improve both the efficiency of analyzing the application's source code and the accuracy of identifying defective codes in the source code.

[0053] The code analysis scheme provided in the embodiments of this application will be described below. Figure 1 A schematic diagram of the hardware execution environment for a code analysis method provided in this application embodiment is shown below. Figure 1 As shown, the hardware execution environment for this code analysis method can consist of a client device 101 and a server device 102, with the client device 101 and server device 102 communicating with each other. The server device 102 can be a cloud server from a cloud service provider. The client device 101 can be a laptop, tablet, PC, robot, etc.

[0054] In an optional embodiment, the execution process of the above code analysis method may be as follows: Client device 101 sends a code analysis request to server device 102 by calling the code analysis service provided by server device 102. The request includes program function description information corresponding to the application and the program source code to be analyzed. Server device 102 first determines the code location information and first defect description information corresponding to the potential defect code in the program source code based on the program function description information and the defect code identification large language model. Then, based on the code location information, it obtains the target code segment containing the potential defect code from the program source code. The source code contained in the target code segment has complete semantic information. Next, based on the first defect description information and the target code segment, it determines the target defect location in the target formal model corresponding to the target code segment. The target formal model is used to reflect the code execution flow and code state transition of the target code segment. The target defect location corresponds to the defect state described by the first defect description information. Based on the target formal model and the target defect location, it determines the target defect code in the program source code. Finally, it feeds back the target defect code to client device 101, so that client device 101 can correct the program source code to be analyzed that is actually input by the user based on the target defect code.

[0055] In practical applications, the aforementioned server-side device 102 can be a cloud server maintained by a cloud service provider—referred to as a computing node. In cases such as... Figure 2 The cloud computing environment shown may include several distributed deployments. Figure 2 The diagram illustrates compute nodes (201-1, 201-2, ...), each possessing processing resources such as computing and storage. In a cloud computing environment, multiple compute nodes can be organized to provide a specific service; conversely, a single compute node can provide one or more services. Figure 2 The diagram illustrates services A, B, C, and D. In a cloud computing environment, these services can be provided via external service interfaces, which client device 101 calls to use the corresponding services. Service interfaces can take the form of Software Development Kits (SDKs) or Application Programming Interfaces (APIs).

[0056] The services described above are deployed using various virtualization technologies supported by cloud computing environments, such as virtual machine-based and container-based virtualization technologies. Taking container-based virtualization technology as an example, several containers corresponding to a service can be assembled into a container group (pod). For example... Figure 2The illustrated service B can be configured with one or more pods, and each pod can include a proxy and one or more containers. The one or more containers in the pod are used to handle requests related to one or more corresponding functions of the service, and the proxy in the pod is used to control network functions related to the service, such as routing and load balancing.

[0057] During operation, executing a request from client device 101 may require invoking one or more services in the cloud computing environment, and executing one or more functions of one service may require invoking one or more functions of another service. For example... Figure 2 As shown, after receiving a request from client device 101, service A can call service B. Service B can request service D to perform one or more functions. In this embodiment of the application, cloud services for code analysis can be deployed on one or more computing nodes.

[0058] The execution process of the code analysis method provided in this application embodiment is described in detail below with reference to the accompanying drawings. This code analysis method can be executed by a computing node in the aforementioned cloud computing environment.

[0059] Figure 3 A flowchart of a code analysis method provided in this application embodiment is shown below. Figure 3 As shown, it may include the following steps:

[0060] 301. Obtain the program function description information corresponding to the application, as well as the program source code to be analyzed.

[0061] 302. Based on the program function description information and the program source code, the code location information and first defect description information corresponding to the potential defect codes in the program source code are determined through the defect code identification large language model.

[0062] 303. Based on the code location information, obtain the target code segment containing potentially defective code from the program source code. The source code contained in the target code segment has complete semantic information.

[0063] 304. Based on the first defect description information and the target code fragment, determine the location of the target defect in the target formal model corresponding to the target code fragment. The target formal model is used to reflect the code execution flow and code state transition of the target code fragment. The location of the target defect corresponds to the defect state described by the first defect description information.

[0064] 305. Based on the formal model of the target and the location of the target defect, determine the target defect code existing in the program source code.

[0065] In this embodiment of the application, the source code to be analyzed can be the complete source code of the application or a portion of the complete source code of the application. For example, when a global analysis of the application's source code is required, the source code to be analyzed is the complete source code of the application; or, for example, when a targeted analysis of the application's source code is required, such as analyzing the source code of a certain component or function of the application, the source code to be analyzed is the source code corresponding to that component.

[0066] For ease of description, in the following embodiments, the program source code to be analyzed will be referred to as program source code. Unless otherwise specified, the program source code mentioned refers to the program source code to be analyzed.

[0067] Optionally, whether the program source code is complete or partial, it can be analyzed either all at once or sequentially. For example, when the program source code is complete, it can be analyzed all at once, or different source codes within the complete source code can be analyzed sequentially according to a preset analysis strategy. For instance, it can be analyzed sequentially based on the dependencies between the source files storing the source code. Specifically, it can first analyze the source code in the application's entry point source file (i.e., the source code file executed first when the application starts), then analyze the source code in other source files that depend on the entry point source file, and so on, until all source code in all source files has been analyzed. When the program source code is partial, the corresponding one-time or sequential code analysis process is similar to that of the complete source code, and will not be elaborated further here.

[0068] In this embodiment, the method of analyzing the program source code is not limited and can be customized. For example, when the amount of code corresponding to the program source code is small, the program source code can be analyzed all at once; when the amount of code corresponding to the program source code is large, the program source code can be divided into different code blocks, and then each code block can be analyzed sequentially.

[0069] It is worth noting that regardless of whether the program source code is complete or partial, and whether the code analysis is performed in a one-time or sequential manner, the processing procedure for each code analysis is the same. Therefore, in the following embodiments, when describing the code analysis process of the program source code, the code type (i.e., complete or partial source code) and the code analysis method (i.e., one-time or sequential code analysis) will no longer be distinguished.

[0070] The following combination Figure 4The code analysis process provided in the embodiments of this application will be described.

[0071] Figure 4 This is a schematic diagram of a code analysis method provided in an embodiment of this application, such as... Figure 4 As shown, in summary, the code analysis process of the program source code in this application embodiment can be divided into two stages:

[0072] The first stage, the initial identification of defective codes: Based on the application's functional description information, the large language model for defective code identification is used to identify potential defective codes in the overall source code, and to determine the code location information and the first defect description information corresponding to the potential defective codes.

[0073] The second stage, the secondary defect code verification stage, involves first obtaining the target code segment containing the potential defect code from the program source code based on the code location information of the potential defect code; then, based on the first defect description information and the target code segment, determining the location of the target defect in the target formal model corresponding to the target code segment; finally, performing defect verification based on the target formal model and the target defect location to determine the target defect code present in the program source code.

[0074] In detail, during the initial defect code identification stage, in the process of determining the code location information and first defect description information of potential defect codes in the program source code based on the program function description information of the application and through the defect code identification language model, optionally, prompt words can be generated based on the obtained program function description information and program source code, and these prompt words can be input into the defect code identification language model. Through prompt word engineering, the defect code identification language model can identify potential defect codes in the program source code under the guidance of the prompt words, and determine the code location information and first defect description information of the potential defect codes.

[0075] Among them, program function description information refers to the information provided by the application documentation that details the application's functions and operating mechanisms, which can provide the necessary context for defect code identification and large language model analysis of the program source code.

[0076] The code location information corresponding to the potential defect code refers to the specific location of the potential defect code in the program source code, including but not limited to the file name of the source file where the potential defect code is located, the line number of the code line corresponding to the potential defect code, etc.

[0077] The first defect description information corresponding to the potential defect code includes, but is not limited to: defect type description information, such as: concurrency problem, dependency problem, atomicity problem, etc.; defect impact description information matching the defect type description information, such as: when the defect type is a concurrency problem, the defect impact description information is that it may cause race conditions under multiple threads or processes; when the defect type is an atomicity problem, the defect impact description information is that the entire update process is not atomic, etc.; defect severity description information, such as: severe, moderate, no impact, etc. Optionally, the defect severity can be customized according to the degree of impact of a certain type of defect on the stability of application operation.

[0078] Optionally, after identifying potential defective codes, the large language model for defect code identification can also generate code correction suggestions or best practices for those codes. The best practices for potential defective codes refer to validated and effective solutions and guidelines, typically derived from industry experience, standards, and successful case studies. Based on the code correction suggestions provided by the large language model, the potential defective codes can be located using their location information, and then corrected according to the suggested corrections.

[0079] In this embodiment, the defect code recognition large language model, as a large language model, has good programming language understanding and text processing capabilities. Therefore, it can perform code analysis on the source code of programs in various programming languages ​​contained in the input prompt words, effectively solving the problem of code analysis difficulties when the source code of a program uses more than one programming language, effectively dealing with complex scenarios in software development where multiple programming languages ​​are used in combination, and effectively improving the applicability and efficiency of code analysis.

[0080] In practical applications, large language models for defect code recognition may experience model illusions during use, meaning that the identified potential defect codes may be somewhat inaccurate, which in turn affects the accuracy of the code location information and the first defect description information corresponding to the output potential defect codes.

[0081] In an optional embodiment, in order to improve the accuracy of the output information of the large language model for defect code recognition, the large language model for defect code recognition can output target prompt information during the execution of the analysis task corresponding to the program source code. The target prompt information is used to prompt relevant personnel to confirm or clarify certain issues in the code analysis process, such as: confirming whether there is really a certain type of defect problem.

[0082] Optionally, by configuring corresponding interaction mechanisms in the prompts of the large language model for defect code recognition, the large language model for defect code recognition can be guided to output target prompt information during the execution of the analysis task corresponding to the program source code, and to perform further code analysis based on the feedback information after obtaining the feedback information corresponding to the target prompt information.

[0083] In this solution, through the aforementioned interaction mechanism, the large language model for defect code identification can obtain relevant feedback information in a timely manner and conduct further analysis based on this feedback. This effectively prevents the model analysis results of the large language model for defect code identification from deviating from the actual correct results, reduces model illusion, and improves the accuracy of code analysis.

[0084] In another optional embodiment, in order to improve the accuracy of the output information of the large language model for defect code identification, a secondary defect code inspection can be performed on the potential defect code based on the code location information and the first defect description information corresponding to the potential defect code output by the large language model for defect code identification, so as to determine the code that actually has defects in the program source code.

[0085] Understandably, potentially defective code is source code that contains defects. It may be one or more lines in the program's source code. It may not be possible to know the execution context information of the source code through the potentially defective code alone, making it difficult to verify whether the potentially defective code actually contains defects.

[0086] Therefore, in the secondary defect code verification stage, target code segments containing potential defective code can be obtained from the program source code. Then, based on the obtained target code segments, a secondary defect code verification is performed on the potential defective code to identify the target defective code that actually contains defects in the program source code.

[0087] In practical applications, inaccurate target code snippet location may occur when acquiring target code snippets. For example, the target code snippet may lack execution context information, or it may cross code block boundaries. Specifically, a target code snippet crossing a code block boundary means that a code snippet that should logically be considered a whole (i.e., a code block) has been segmented during target code snippet identification, and the target code snippet only contains a portion of that whole code snippet. This inaccurate target code snippet location leads to the source code contained in the target code snippet lacking complete semantic information, making it difficult to understand and affecting the accuracy of secondary verification of defective code.

[0088] To address the issue of inaccurate target code segment location and improve the accuracy of secondary verification of potentially defective code, in this embodiment, optionally, an abstract syntax tree is used to accurately locate the target code segment containing the potentially defective code, starting from the code location corresponding to the potentially defective code, and ensuring that the target code segment is syntactically complete and that the contained source code has complete semantic information.

[0089] In the specific implementation process, the abstract syntax tree (AST) corresponding to the program source code can be determined first. Then, based on the code location information of potential defective codes output by the large language model for defect code identification, the target node corresponding to the potential defective code in the AST can be determined. Next, starting from the target node, the target parent node closest to the target node and / or the target child node referencing the target node in the AST can be determined; the target parent node is the function definition node. Finally, based on the first source code corresponding to the nodes between the target parent node and the target node, and / or the second source code corresponding to the nodes between the target node and the target child node, the target code fragment can be determined. The specific implementation process of generating the AST corresponding to a source code segment can be found in relevant technologies and will not be elaborated upon in this embodiment.

[0090] For ease of understanding, the following is combined with Figure 5 An illustrative explanation is provided of the process of determining target code fragments based on abstract syntax trees. Figure 5 A schematic diagram of an abstract syntax tree provided for the implementation of this application, such as... Figure 5 As shown, the root node of the abstract syntax tree corresponding to the program source code is the source file of the program source code. Under the root node are multiple different nodes corresponding to different lines of source code, such as node 1, node 2, etc. Assuming the code location information corresponding to the potential defective code is line n in the program source code, the node corresponding to line n in the abstract syntax tree is called the target node. Figure 5 Node 6 in the middle.

[0091] In the process of determining the target code snippet, such as Figure 5 As shown, on the one hand, the target code fragment can be determined from the dimension of the source code implementation corresponding to the target node. Specifically, starting from the target node, a search is performed towards the root node to determine the target parent node in the abstract syntax tree that is closest to the target node and is a function definition node (that is, the source code corresponding to the target parent node is used to define the function that implements the source code of line n). The source code corresponding to the nodes between the target parent node and the target node is then determined as the first source code used to construct the target code fragment. Here, the target parent node corresponds to... Figure 5 Node 4 in the middle, based on this, in Figure 5 In the illustrated scenario, the first source code is the source code corresponding to nodes 4, 5, and 6.

[0092] On the other hand, the target code fragment can be determined from the dimension of the source code referenced by the target node. Specifically, starting from the target node, a search is performed towards the leaf nodes to identify the target child nodes that reference the target node, and the source code corresponding to the nodes between the target node and the target child nodes is determined as the second source code used to construct the target code fragment. Here, the target child node corresponds to... Figure 5 Based on nodes 7 and 8 in the data, in Figure 5 In the illustrated scenario, the second source code is the source code corresponding to nodes 6, 7, and 8.

[0093] Optionally, the target code fragment may contain only the first source code or the second source code, or it may contain both the first source code and the second source code. For example, when the target node corresponding to the potential defective code is a function definition node, the target code fragment may contain only the second source code; when the target node corresponding to the potential defective code is a leaf node (i.e., not referenced by any child nodes), the target code fragment may contain only the first source code.

[0094] It should be noted that, Figure 5 The node distribution in the abstract syntax tree shown is for illustrative purposes only and is not intended to be limiting.

[0095] Optionally, after determining the target code segment based on the abstract syntax tree, it is further possible to obtain code description information corresponding to other source code segments besides the target code segment in the program's source code, such as... Figure 5 The diagram shows the code descriptions for nodes such as Node 1 and Node 3. These descriptions are natural language descriptions of the code's functionality within the source code. The code descriptions for other source code segments provide contextual information about the target code segment's execution during secondary verification of potentially defective code. This helps in understanding the target code segment's execution environment, behavior, and intent, thereby improving the accuracy of the secondary verification.

[0096] This solution utilizes the abstract syntax tree corresponding to the program source code to identify the code location information of potential defective codes output by the large language model based on defective code, accurately extracting semantically complete and independent target code fragments from the program source code. Since these target code fragments not only contain potential defective codes but also retain semantically complete contextual information, they provide an accurate analytical foundation for subsequent secondary defect code verification, thereby improving the accuracy of effectively ensuring secondary defect code verification.

[0097] It's worth noting that in the field of software development, formal methods for formally modeling the code execution flow and state transitions of application source code, and then verifying the correctness of the source code based on these formal models, is an efficient and accurate method for source code verification. The essence of formal modeling is to represent the code execution flow and state transitions of the source code using a mathematical model. Optionally, formal models include, but are not limited to, Petri nets, Pi calculus, and the Unified Modeling Language (UML).

[0098] Based on this, after obtaining the target code fragment, the target formal model corresponding to the target code fragment and the target code fragment are further determined based on the first defect description information and the target code fragment corresponding to the potential defect code output by the large language model of defect code identification. The target defect location in the target formal model is also determined.

[0099] The target formal model is used to reflect the code execution flow and code state transition of the target code segment; the target defect location corresponds to the defect state described by the first defect description information.

[0100] In practical applications, a formal model can be understood as a directed graph, using nodes and directed edges to represent the code states in the target code fragment and the transitions between states in the code execution flow. The target defect location in the formal model is, in other words, the node location in the formal model used to constitute the defect state described by the first defect description information.

[0101] Optionally, the target defect location can be a single defect location (or node location) in the target formal model, or it can be a combination of multiple defect locations in the target formal model. When the target defect location is a combination of multiple defect locations, it indicates that when these multiple defect locations simultaneously exhibit a certain state, a certain type of code defect will occur, affecting the operation of the application.

[0102] Finally, based on the target formal model and the location of the target defect, a formal model verification tool is used to perform defect verification to identify the target defect code present in the program source code. In practical applications, the target defect code may be contained within the potential defect code or may be distinct from the potential defect code.

[0103] In practice, if the potential defect code includes the target defect code, it indicates that the target defect code has been identified as having a code defect through both the large language model for defect code identification and the secondary verification method. If the potential defect code does not include the target defect code, it indicates that the target defect code is a newly identified defect code in the secondary defect verification stage.

[0104] Optionally, in order to improve the accuracy of the target defect code, in response to the fact that the potential defect code does not include the target defect code, a prompt message corresponding to the target defect code can be output. This prompt message is used to prompt relevant personnel to confirm whether the target defect code is a defect code that actually exists in the program source code. In response to the confirmation operation triggered by the relevant personnel based on the prompt message, the second defect code is confirmed to be the target defect code corresponding to the program source code.

[0105] Optionally, for defective codes that are not detected through secondary defect code verification in potential defective codes, it can be directly determined that they are not defective codes existing in the program source code. Alternatively, for defective codes that are not detected in potential defective codes, a prompt message can be output to prompt relevant personnel to determine whether the undetected defective code is a defective code. In response to relevant personnel confirming that the undetected defective code is a defective code based on the prompt message, the undetected defective code is determined to be the target defective code corresponding to the program source code; otherwise, the undetected defective code is determined not to be the target defective code corresponding to the program source code.

[0106] In summary, in this embodiment, after obtaining the program function description information corresponding to the application and the program source code to be analyzed, the potential defective codes in the program source code are first identified through a large language model for defect code identification, and the code location information and first defect description information corresponding to the potential defective codes are determined. Then, based on the code location information, target code segments containing potential defective codes are obtained from the program source code, and based on the first defect description information and the target code segments, the location of the target defect in the target formal model corresponding to the target code segments is determined. Finally, using traditional formal modeling verification methods, formal verification is performed based on the target formal model and the target defect location to determine the target defective codes present in the program source code. This multi-stage code analysis method, which combines large language model identification with automated secondary verification, improves both the analysis efficiency of the application's program source code and the accuracy of defective code identification in the program source code. The overall process of the code analysis method provided in this embodiment has been described above.

[0107] Next, the specific process for determining potential defect codes and target defect codes will be explained in detail.

[0108] The process of identifying potential defect codes will be explained below.

[0109] As described above, the code analysis method provided in this application identifies potential defective code in the program source code through a large language model for defective code recognition, and determines the code location information and first defect description information corresponding to the potential defective code. The large language model for defective code recognition, acting as a large language model, performs corresponding tasks guided by prompt words.

[0110] In this embodiment of the application, in the application scenario of defect code analysis, the first prompt word used to input the large language model for defect code recognition includes at least the following elements: input information, output requirements, and task description.

[0111] The input information refers to the basic data or contextual content provided to the large language model for defect code recognition, including but not limited to: the source code of the program to be analyzed, and the program function description information of the application. It is easy to understand that the more accurate the information used for defect code recognition in the input information, the better the large language model for defect code recognition can understand the defect code recognition requirements, and thus more accurately identify the defect code.

[0112] Output requirements refer to the specific expectations and specifications for the output results of the large language model for defect code recognition, used to clarify the content type, format, etc., that the large language model for defect code recognition should generate. For example, in this embodiment of the application, the output requirements corresponding to the large language model for defect code recognition may be: the output results include code location information and first defect description information corresponding to the potential defect code, wherein the code location information includes, but is not limited to, the file name of the source file where the potential defect code is located, the line number of the potential defect code, etc.; the first defect description information includes, but is not limited to, defect type description information, defect impact description information, defect severity description information, etc. Optionally, the output requirements may also include information guiding the large language model for defect code recognition to output code correction suggestions or practical references corresponding to the potential defect code.

[0113] The task description refers to a detailed explanation of the defect code recognition task that the large language model is required to perform. For example, it might describe how to identify defective code in the program's source code based on the program's functional description information in the input. It's easy to understand that the more accurate the task description, the better the large language model can identify defective code in the program's source code.

[0114] Based on the above description of the first prompt word, in practical applications, optionally, the accuracy of potential defective codes contained in the source code of the large language model recognition program can be improved by flexibly configuring elements such as input information, output requirements, and task descriptions in the first prompt word. This ensures the accuracy of the code location information and the first defect description information corresponding to the potential defective code. Several specific embodiments are provided below for illustrative purposes.

[0115] Figure 6 A flowchart of a potential defect code determination process provided for embodiments of this application is shown below. Figure 6 As shown, it can include the following steps:

[0116] 601. Based on the application's functional description information, identify the large language model of defect code to determine the program execution mode corresponding to the source code of the program to be analyzed. Different program execution modes correspond to different types of code defects.

[0117] 602. Based on the code defects corresponding to the program execution mode, identify potential defective codes in the program source code through the defective code identification large language model, and determine the code location information and first defect description information corresponding to the potential defective codes.

[0118] In practical applications, source code may be designed with different program execution modes based on the different application usage requirements. Optionally, the program execution mode includes, but is not limited to, concurrent execution, asynchronous execution, distributed execution, and repeated invocation.

[0119] It's easy to understand that the types of potential code defects differ depending on the program execution mode. For example, for concurrent execution, the potential code defect type might be concurrency conflicts; for asynchronous execution, the potential code defect types might be race conditions, resource leaks, unhandled exceptions, etc.

[0120] In practical applications, if the first prompt word used in the large language model for input defect code recognition only roughly describes the need to identify defective codes in the program source code based on the program function description information in the input information, and to determine the code location information and first defect description information corresponding to the potential defective code, then the large language model for defect code recognition may ignore certain types of code defects during the process of identifying potential defective codes under the guidance of the first prompt word, resulting in low accuracy of the identification results of potential defective codes.

[0121] In this embodiment of the application, in order to improve the accuracy of the large language model for defect code identification in identifying potential defective codes, the task description in the first prompt word can be refined according to the program execution mode and the potential code defect type corresponding to the program execution mode. For example, the task description of the first prompt word may be as follows: 1) Determine the program execution mode corresponding to the program source code based on the program function description information; 2) Identify the potential defective code in the program source code according to the potential code defect type corresponding to different program execution modes; 3) Determine the code location information and the first defect description information corresponding to the potential defective code.

[0122] In practical applications, after inputting a first prompt word containing the program execution mode and the corresponding potential code defect type into the defect code recognition language model, the defect code recognition language model first determines the program execution mode corresponding to the program source code based on the program function description information of the application; then, based on the code defect type corresponding to the program execution mode, it identifies the potential defect code in the program source code; finally, it determines the code location information and first defect description information corresponding to the potential defect code.

[0123] In this solution, by adding information about the program execution mode and the corresponding code defect type to the first prompt word, the large language model for defect code identification can be guided to identify defective codes from the dimension of program execution mode, and to focus on the potential code defect type of the program execution model corresponding to the program source code. This can effectively improve the accuracy of potential defect code identification.

[0124] Figure 7 A flowchart of another potential defect code determination process provided for embodiments of this application is shown below. Figure 7 As shown, it can include the following steps:

[0125] 701. Based on the program function description information of the application, identify the component function description information associated with the target component corresponding to the program source code to be analyzed by using the defect code identification large language model. The component function description information includes at least one of the following: the component name of the target component, the dependency relationship between the target component and other components contained in the application, and the definition information of the interface associated with the target component.

[0126] 702. Based on the component function description information associated with the target component, identify potential defective codes in the program source code through the defect code identification large language model, and determine the code location information and first defect description information corresponding to the potential defective codes.

[0127] In this embodiment of the application, the target component is one of the functional components included in the application, and the program function description information of the application includes the component function description information associated with the target component.

[0128] It's easy to understand that program function description information is the complete functional description information corresponding to the application, which often contains a large amount of information. When the first prompt only contains program function description information, the large language model for defect code identification may lack targeted analysis, ignoring some details related to the program source code in the program function description information, and thus the accuracy of identifying potential defect codes may be poor.

[0129] Since the component function description information of the target component corresponding to the program source code has less information and is more closely related to the program source code, in this embodiment of the application, the component function description information of the target component can be used as one of the input information to generate the first prompt word for the large language model for inputting defect code recognition.

[0130] Optionally, the component functional description information includes at least one of the following: the component name of the target component, the dependency relationship between the target component and other components included in the application, and the definition information of the interface associated with the target component.

[0131] In practice, when the component function description information contained in the first prompt word is the component name of the target component, the first prompt word can be used to guide the large language model for defect code identification to focus on information related to the target component in the program function description information when identifying defect code. For example, it can guide the large language model for defect code identification to obtain information related to the target component from the program function description information, including but not limited to the functions that the target component is to implement and the way in which it implements them. Based on the information related to the target component obtained, the large language model for defect code identification further determines whether the functions implemented by the program source code are the same as the functions of the target component, whether the ways in which they are implemented are the same, etc., thereby identifying potential defect code in the program source code.

[0132] When the component function description information contained in the first prompt word is the definition information of the interface associated with the target component, the first prompt word can be used to guide the large language model for defect code identification to identify potential defective code in the program source code from the dimension of interface definition. For example, based on whether the input and output data types corresponding to a certain interface in the program source code are consistent with the input and output data types defined in the interface definition information, or whether the corresponding interface in the program source code is consistent with the interface provided by the target component, potential defective code can be identified from the program source code.

[0133] When the component function description information contained in the first prompt word indicates a dependency relationship between the target component and other components included in the application, this first prompt word can guide the defect code identification language model to obtain information related to the target component and other components from the program function description information. This information includes, but is not limited to: the implementation of the Application Programming Interface (API) called between the target component and other components, the function or method call chain between the target component and other components, and the class dependencies between the target component and other components. This information provides the defect code identification language model with targeted contextual information related to the program source code. Based on this contextual information, the defect code identification language model can more accurately identify potential defective code from the program source code.

[0134] In this solution, by adding the component function description information associated with the target component corresponding to the program source code to the first prompt word, the large language model for defect code identification can be guided to focus on the information related to the target component in the program function description information when identifying defect code, thereby improving the targeting of defect code identification and thus improving the accuracy of potential defect code identification.

[0135] Optionally, Figure 6 and Figure 7 The method for determining potential defective codes in the illustrated embodiment can be executed individually or in combination. For example, based on program function description information and program source code, the method uses a large language model for defective code identification to determine the program execution mode corresponding to the program source code, as well as the component function description information associated with the target component corresponding to the program source code; based on the code defects corresponding to the program execution mode and the component function description information, the method uses a large language model for defective code identification to determine the potential defective codes present in the program source code, as well as the code location information and first defect description information corresponding to the potential defective codes.

[0136] Figure 8 A flowchart of another potential defect code determination process provided for embodiments of this application is shown below. Figure 8 As shown, it can include the following steps:

[0137] 801. Based on the program function description information and the program source code, identify other program source code related to the code implementation or code call of the program source code through the large language model of defect code identification.

[0138] 802. Based on the program function description information, program source code, and other program source code, use the defect code identification large language model to determine the code location information and first defect description information corresponding to the potential defect codes in the program source code.

[0139] Optionally, in this embodiment of the application, the defect code identification large language model has the ability to call static code analysis tools. The static code analysis tools are used to: determine the location information of the code implementations that the source code depends on, or the location information of the code calls that the source code depends on.

[0140] In identifying potential defective code in program source code, the large language model for defect code identification can determine the location information of the code implementations / calls that the program source code depends on by calling static code analysis tools. Then, based on the location information of the code implementations / calls that the program source code depends on, other program source codes that match the location information of the code implementations / calls that the program source code depends on are obtained from the full source code of the application. These other program source codes have complete semantic information. Finally, based on the program function description information, the program source code, and the other program source codes, the code location information and the first defect description information corresponding to the potential defective code in the program source code are determined.

[0141] Understandably, other program source code can provide more code information related to the program source code for the large language model for defect code identification. More complete code information helps improve the accuracy of identifying potential defective code.

[0142] Optionally, after obtaining the source code of other programs, the defect code identification language model can further determine the location information of the code implementation / code call on which the source code of other programs depends by calling static code analysis tools. Then, similar to the process of obtaining the source code of other programs, the target program source code that matches the location information of the code implementation / code call on which the source code of other programs depends can be obtained from the full source code of the application. Based on the program function description information, the program source code, the source code of other programs, and the target program source code, the code location information and the first defect description information corresponding to the potential defect code in the program source code can be determined.

[0143] Similarly, the large language model for defect code identification can continuously acquire source code related to the program's source code by repeatedly calling static code analysis tools, and then combine the acquired source code to identify potential defect codes in the program's source code.

[0144] Optionally, when identifying potential defective codes based on acquired sources such as other program source code, the defective code recognition language model can also output the code location information and defect description information of potential defective codes in other program source code, so as to facilitate timely correction of other program source code.

[0145] Optionally, by configuring the corresponding interaction mechanism in the first prompt word, the large language model for defect code identification can obtain the source code of the aforementioned other programs during the process of identifying potential defect codes.

[0146] The process described above, which involves obtaining the source code of other programs from the full source code of the application based on the location information of the code implementations / code calls that the source code of other programs depends on, is similar to... Figure 5 The process of obtaining the target code fragment based on the code location information of the potential defective code, as shown in the example, is similar and will not be described again here.

[0147] In this solution, the large language model for defect code identification continuously acquires source code related to the program's source code during the process of identifying potential defect codes. By combining the acquired source code with the source code to identify potential defect codes, it can perform a more comprehensive analysis of the program's source code based on richer contextual information of the code implementation, thereby improving the accuracy of potential defect code identification.

[0148] Optionally, when the program source code is not the complete source code corresponding to the application, the first prompt can also include the location information of the program source code within the complete source code corresponding to the application, such as: which lines of code in which source file corresponding to the application the program source code is located in. Based on the location information corresponding to the program source code, the large language model for defect code recognition can easily and quickly obtain information related to the program source code from the program function description information when performing defect code recognition, accurately obtain the source code that the program source code depends on, or the source code that depends on the program source code, etc., thereby obtaining targeted and comprehensive contextual information and improving the accuracy of potential defect code recognition.

[0149] Optionally, the first prompt may also include the type and version of the programming language used in the program source code (e.g., Python 2.7, Java SE8, etc.), as well as the custom functions corresponding to the programming language used in the program source code.

[0150] As mentioned above, in this embodiment, the ability of a large language model to understand different programming languages ​​is utilized to identify potential defective code in the program source code. Therefore, the prerequisite for defective code identification in program source code is to identify the programming language type corresponding to the program source code.

[0151] However, in practical applications, there may be certain similarities between different programming languages, or between different versions of the same programming language. Therefore, the large language model for defect code identification carries the risk of incorrectly identifying the programming language type corresponding to the program source code.

[0152] It is understandable that different programming languages, or even different versions of the same programming language, may have certain differences in syntax or functionality. Therefore, if the large language model for defect code identification incorrectly identifies the programming language category of the program's source code, the final identification of potential defect code may also be inaccurate.

[0153] In this solution, by specifying the programming language type used in the program source code in the first prompt word, as well as the custom functions corresponding to that programming language (excluding known functions), accurate information related to the programming language of the program source code can be provided to the large language model for defect code identification, thereby effectively improving the accuracy of potential defect code identification.

[0154] The process for identifying potential defect codes has been explained above. The process for identifying target defect codes will be explained next.

[0155] As mentioned above, in this embodiment of the application, the target defect code in the program source code is determined by formally modeling the target code fragment corresponding to the potential defect code and formally verifying the established formal model.

[0156] In practical applications, formal modeling via manual methods requires not only that modelers understand formal knowledge and code execution logic, but also that they are familiar with the language characteristics of different programming languages. These requirements make formal modeling of source code particularly difficult. Furthermore, the results of formal modeling may contain some inaccuracies due to the influence of the modeler's personal subjective bias.

[0157] To address at least one of the aforementioned technical problems, embodiments of this application utilize the programming language understanding and formal model generation capabilities of formal modeling large language models to generate target formal models corresponding to target code fragments. This allows for defect code verification based on the target formal model, thereby identifying target defective code present in the program source code.

[0158] Specifically, based on the first defect description information, the target code fragment, and the model definition of the formal model to be modeled, the formal modeling large language model is used to output the target formal model corresponding to the target code fragment, as well as the location of the target defect in the target formal model; then, based on the target formal model and the location of the target defect, the target defect code in the program source code is determined.

[0159] The embodiments of this application do not limit the specific type of the target formal model. For example, it can be a Petri net or other types of formal models.

[0160] Similar to the defect code recognition large language model in the aforementioned embodiments, the formal modeling large language model, as a large language model, can also perform corresponding tasks through the guidance of prompt words.

[0161] In practical applications, optionally, a second prompt word can be generated based on the first defect description information, the target code snippet, and the description information of the formal model to be modeled. This prompt word allows the formal model to generate the target formal model corresponding to the target code snippet based on the description information of the formal model, and to determine the location of the target defect in the target formal model based on the first defect description information. The second prompt word must contain at least the following elements: input information, output requirements, and task description.

[0162] The model definition of the formal model to be modeled includes at least one of the following: descriptive information of the formal model to be modeled, modeling constraints of formal modeling, etc. The model definition of the formal model to be modeled is used to provide corresponding formal model information for the formal modeling large language model, so as to guide it to generate a formal model that meets the expectations.

[0163] The description information of the formal model to be modeled is the description of the formal model to be constructed. For example, when it is expected that a formal model corresponding to a target code fragment will be constructed using Petri nets, the description of the formal model to be modeled can be: the Petri net is a quintuple PN = {P, T, F, W, M0}, where P represents the set of positions, used to represent the code states in the source code; T represents the set of transitions, used to represent events or operations in the code execution flow; F represents the set of flow relations, used to represent the connection relationship between P and T, indicating the flow of markers between P and T, represented by directed edges; W represents the weight function, used to represent the number of markers that can be transmitted by each directed edge; and M0 represents the initial markers corresponding to each position in P at the initial time. Based on the description information of the Petri net in the second prompt word, the formal modeling large language model can generate the Petri net corresponding to the target code fragment.

[0164] Formal modeling constraints are typically matched to the subsequent verification process. For example, to achieve efficient and accurate formal verification, constraints can be placed on the number of nodes in the formal model and whether cycles exist within it. A cycle refers to a cyclic dependency relationship between certain nodes in the formal model. By using modeling constraints, the problem of the target formal model generated from a large formal model being unable to undergo formal verification can be effectively avoided.

[0165] In the application scenario of formal model construction, optionally, the input information contained in the second prompt word may include, in addition to the target code fragment, the model definition of the formal model to be modeled, and the first defect description information, at least one of the following: the location information of the target code fragment in all the source code corresponding to the application, the type and version of the programming language used by the target code fragment, the custom functions corresponding to the programming language used by the target code fragment, and information related to the program execution mode corresponding to the target code fragment, such as concurrent execution information, distributed execution information, etc.

[0166] The second prompt word provides information on the location of the target code fragment within the entire source code of the application. For example, it indicates which lines of code in which source file the target code fragment belongs to. This information guides the formal modeling language model in generating the target formal model. Based on the location information of the target code fragment, the model can quickly and easily obtain contextual information related to the target code fragment from the program's functional description information, thereby improving the accuracy of the generated formal model.

[0167] The type and version of the programming language used in the target code fragment, the custom functions corresponding to the programming language used in the target code fragment, etc. in the second prompt are used to specify the type of programming language used in the target code fragment for the formal modeling large language model, as well as the custom functions corresponding to the programming language of this type, excluding the known functions. This enables the formal modeling large language model to adopt the analysis method that matches the specified programming language and construct the target formal model corresponding to the target code fragment.

[0168] The information related to the program execution mode corresponding to the target code fragment in the second prompt word is used to guide the formal modeling language model to generate a more concise and less complex formal model. For example, when the information related to the program execution mode corresponding to the target code fragment is concurrent execution information, this concurrent execution information can describe the master-slave relationship (or dependency relationship) between different code segments in the target code fragment. Based on the master-slave relationship in the concurrent execution information, the formal modeling language model can constrain the state transition relationships between different nodes in the generated target formal model, thereby avoiding listing too many possible state transition relationships in the target formal model and simplifying the complexity of the target formal model.

[0169] Optionally, the output requirements included in the second prompt may, in addition to describing that the output contains the target formal model and the location of the target defect within the target formal model, also describe that the output contains second defect description information corresponding to the target defect location. This second defect description information is obtained based on the first defect description information through further formal modeling and defect code analysis. Therefore, it can include more granular defect description information, including but not limited to: the conditions that caused the defect, the impact of the defect on application operation, and suggestions for resolving the defect.

[0170] Optionally, the task description included in the second prompt can describe specific formal modeling details, such as: formally modeling each statement of the source code in the target code snippet, paying attention to control flow structures such as loop statements and conditional statements, as well as function calls; identifying the program execution mode corresponding to the target code snippet, and focusing on the code state transitions of preset key variables corresponding to this program execution mode. Here, key variables refer to variables that may have a significant impact on the application's operation under the program execution mode; for example, in concurrent execution, key variables can be variables shared by multiple concurrent threads. These formal modeling details in the task description can better guide the formal modeling of large language models to generate formal models.

[0171] In practical applications, the execution of the source code corresponding to the target code snippet may involve interactions with resources other than the entire source code of the application, such as database operations. To ensure the accuracy of formal analysis based on a formal model, the interactions of the target code snippet with other resources can optionally be appropriately represented in the formal model, for example, indicating which information from external databases the target code snippet calls.

[0172] After generating the target formal model corresponding to the target code fragment by formally modeling the large language model, and determining the target defect location in the target formal model, further formal analysis is performed on the target formal model to determine whether the defect state corresponding to the target defect location will actually occur. When it is determined that it will actually occur, the source code corresponding to the target defect location is obtained as the target defect code in the target code fragment.

[0173] It is worth noting that the formal model obtained by transforming the target code fragment is a mathematical model with unified specifications. It is not affected by the programming language type. Therefore, when performing formal analysis on the formal model, a unified analysis method can be used for efficient and fast analysis.

[0174] Optionally, it can be determined whether the defect state corresponding to the target defect location will actually occur by judging whether there is a state transition path from the initial state of the target formal model to the defect state corresponding to the target defect location. Specifically, if there is a state transition path from the initial state of the target formal model to the defect state corresponding to the target defect location, then it is determined that the defect state corresponding to the target defect location will actually occur; if there is no state transition path from the initial state of the target formal model to the defect state corresponding to the target defect location, then it is determined that the defect state corresponding to the target defect location will not actually occur.

[0175] In practical applications, the formal model can first be converted into a reachability graph. Then, for the node corresponding to the target defect location in the reachability graph, the search begins from the starting node. If a state transition path exists in the reachability graph from the starting node to the node corresponding to the target defect location, the defect state corresponding to the target defect location is determined to be reachable. That is, by executing the source code corresponding to the node included in the state transition path and performing the corresponding code state transition, the defect state corresponding to the target defect location will be generated. If no state transition path exists in the reachability graph from the starting node to the node corresponding to the target defect location, the defect state corresponding to the target defect location is determined to be unreachable, that is, it will not occur.

[0176] Furthermore, if there exists a state transition path in the target formal model from the initial state of the target formal model to the defect state corresponding to the target defect location, then the source code corresponding to the target defect location is determined to be the target defect code corresponding to the program source code.

[0177] In practical applications, the purpose of identifying defective codes is to correct them. Therefore, based on the first defect description information, the target code fragment, and the description information of the formal model to be modeled, a second defect description information corresponding to the target defect location can be generated through formal modeling of a large language model. Then, in response to determining that the target defect code corresponding to the source code at the target defect location is the target defect code corresponding to the program source code, the target defect code and the second defect description information are output. This allows others to correct the target defect code based on the second defect description information.

[0178] In summary, this embodiment of the application, by automatically constructing the target formal model corresponding to the target code fragment and identifying the target defect location in the target formal model through formal modeling of a large language model, can effectively avoid the complexity of formal modeling brought about by the characteristics of programming languages, reduce the dependence on professional formal modeling knowledge, minimize the uncertainty of manual formal modeling, and improve the efficiency of formal modeling. After formal modeling, formal verification of the target formal model can further determine whether the target defect location identified by the formal modeling of the large language model will actually cause the corresponding type of defect, thereby ensuring the accuracy of the secondary verification results of the defect code. In this solution, by combining formal modeling of a large language model with formal verification, the flexibility of the large language model is maintained, while the accuracy and reliability of the target defect code determination results are improved through rigorous formal verification.

[0179] The code analysis apparatus of one or more embodiments of this application will be described in detail below. Those skilled in the art will understand that these apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.

[0180] Figure 9 This is a schematic diagram of the structure of a code analysis device provided in an embodiment of this application, as shown below. Figure 9 As shown, the device includes: an acquisition module 11, a processing module 12, and a determination module 13.

[0181] The acquisition module 11 is used to acquire the program function description information corresponding to the application and the program source code to be analyzed.

[0182] Processing module 12 is used to determine the code location information and first defect description information corresponding to potential defect codes in the program source code by using a defect code identification large language model based on the program function description information and the program source code.

[0183] The determining module 13 is configured to: obtain a target code segment containing the potential defective code from the program source code based on the code location information, wherein the source code contained in the target code segment has complete semantic information; determine the location of the target defect in the target formal model corresponding to the target code segment based on the first defect description information and the target code segment, wherein the target formal model is used to reflect the code execution flow and code state transition of the target code segment, and the location of the target defect corresponds to the defect state described by the first defect description information; and determine the target defective code existing in the program source code based on the target formal model and the location of the target defect.

[0184] In an optional embodiment, when determining the location of the target defect in the target formal model corresponding to the target code fragment based on the first defect description information and the target code fragment, the determining module 13 is specifically used to: based on the first defect description information, the target code fragment and the model definition of the formal model to be modeled, output the target formal model corresponding to the target code fragment and the location of the target defect in the target formal model by formal modeling a large language model.

[0185] In an optional embodiment, when determining the target defect code existing in the program source code based on the target formal model and the target defect location, the determining module 13 is specifically used to: if there is a state transition path in the target formal model from the initial state of the target formal model to the defect state corresponding to the target defect location, then determine that the defect state corresponding to the target defect location exists; determine that the source code corresponding to the target defect location is the target defect code existing in the program source code.

[0186] In an optional embodiment, when the determining module 13 outputs the target formal model corresponding to the target code fragment and the target defect location in the target formal model through a formal modeling large language model based on the model definition of the first defect description information, the target code fragment, and the formal model to be modeled, it is specifically used to: output the second defect description information of the target formal model corresponding to the target code fragment, the target defect location in the target formal model, and the defect state corresponding to the target defect location through a formal modeling large language model based on the description information of the first defect description information, the target code fragment, and the formal model to be modeled.

[0187] Correspondingly, the determining module 13 is further configured to: in response to determining that the source code corresponding to the target defect location is the target defect code existing in the program source code, output the code location information of the target defect code and the second defect description information.

[0188] In an optional embodiment, when the processing module 12 determines the code location information and first defect description information corresponding to the potential defect code in the program source code based on the program function description information and the program source code through the defect code identification large language model, it is specifically configured to: determine the program execution mode corresponding to the program source code and / or the component function description information associated with the target component corresponding to the program source code based on the program function description information and the program source code through the defect code identification large language model; wherein, different program execution modes correspond to different types of code defects, and the component function description information includes at least one of the following: the component name of the target component, the dependency relationship between the target component and other components included in the application, and the definition information of the interface associated with the target component; and determine the potential defect code in the program source code, as well as the code location information and first defect description information corresponding to the potential defect code, based on the code defect corresponding to the program execution mode and / or the component function description information through the defect code identification large language model.

[0189] In an optional embodiment, when the processing module 12 determines the code location information and first defect description information corresponding to the potential defect code in the program source code based on the program function description information and the program source code through the defect code identification large language model, it is specifically configured to: obtain other program source code related to the code implementation or code call of the program source code based on the program function description information and the program source code through the defect code identification large language model; and determine the code location information and first defect description information corresponding to the potential defect code in the program source code based on the program function description information, the program source code, and the other program source code through the defect code identification large language model.

[0190] In an optional embodiment, when the determining module 13 obtains the target code segment containing the potentially defective code from the program source code based on the code location information, it is specifically configured to: determine the abstract syntax tree corresponding to the program source code; determine the target node corresponding to the potentially defective code in the abstract syntax tree based on the code location information; determine the target parent node closest to the target node in the abstract syntax tree, and / or the target child node referencing the target node, with the target parent node being a function definition node; and determine the target code segment based on the first source code corresponding to the node between the target parent node and the target node, and / or the second source code corresponding to the node between the target node and the target child node.

[0191] Figure 9The device shown can perform the steps described in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.

[0192] In one possible design, the above Figure 9 The structure of the code analysis device shown can be implemented as an electronic device, such as... Figure 10 As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores a computer program, which, when executed by the processor 22, enables the processor 22 to at least implement the code analysis method provided in the foregoing embodiments.

[0193] The aforementioned memory 21 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0194] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.

[0195] Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable code analysis device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable code analysis device can be implemented as a means to implement the corresponding functions in the above method embodiments.

[0196] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. This application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0198] Finally, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0199] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A code analysis method, characterized in that, include: Obtain the program function description information corresponding to the application, as well as the source code of the program to be analyzed; Based on the program function description information and the program source code, the code location information and first defect description information corresponding to the potential defect code in the program source code are determined by using the defect code identification large language model. Based on the code location information, a target code segment containing the potential defective code is obtained from the program source code, wherein the source code contained in the target code segment has complete semantic information; Based on the first defect description information and the target code fragment, the location of the target defect in the target formal model corresponding to the target code fragment is determined. The target formal model is used to reflect the code execution flow and code state transition of the target code fragment. The location of the target defect corresponds to the defect state described by the first defect description information. Based on the formal model of the target and the location of the target defect, the target defect code existing in the program source code is determined.

2. The method according to claim 1, characterized in that, The step of determining the location of the target defect in the target formal model corresponding to the target code fragment based on the first defect description information and the target code fragment includes: Based on the first defect description information, the target code snippet, and the model definition of the formal model to be modeled, the target formal model corresponding to the target code snippet, as well as the location of the target defect in the target formal model, are output through formal modeling of a large language model.

3. The method according to claim 1, characterized in that, The step of determining the target defect code present in the program source code based on the target formal model and the target defect location includes: If there exists a state transition path in the target formal model from the initial state of the target formal model to the defect state corresponding to the target defect location, then it is determined that the defect state corresponding to the target defect location exists. The source code corresponding to the location of the target defect is determined to be the target defect code existing in the program source code.

4. The method according to claim 2, characterized in that, The step of defining the model based on the first defect description information, the target code snippet, and the formal model to be modeled, and outputting the target formal model corresponding to the target code snippet, as well as the location of the target defect in the target formal model, through formal modeling of a large language model, includes: Based on the first defect description information, the target code fragment, and the description information of the formal model to be modeled, the second defect description information is output through formal modeling of a large language model, including the target formal model corresponding to the target code fragment, the target defect location in the target formal model, and the defect state corresponding to the target defect location. The method further includes: In response to determining that the source code corresponding to the target defect location is the target defect code existing in the program source code, the code location information of the target defect code and the second defect description information are output.

5. The method according to claim 1, characterized in that, The step of determining the code location information and first defect description information corresponding to potential defect codes in the program source code through a defect code identification large language model based on the program function description information and the program source code includes: Based on the program function description information and the program source code, the program execution mode corresponding to the program source code is determined through the defect code identification large language model, and / or the component function description information associated with the target component corresponding to the program source code; wherein, different program execution modes correspond to different types of code defects, and the component function description information includes at least one of the following: the component name of the target component, the dependency relationship between the target component and other components included in the application, and the definition information of the interface associated with the target component; Based on the code defects corresponding to the program execution mode, and / or the component function description information, the potential defective code in the program source code is determined through the defective code identification large language model, as well as the code location information and first defect description information corresponding to the potential defective code.

6. The method according to claim 1, characterized in that, The step of determining the code location information and first defect description information corresponding to potential defect codes in the program source code through a defect code identification large language model based on the program function description information and the program source code includes: Based on the program function description information and the program source code, other program source code related to the code implementation or code call of the program source code is obtained through the defect code identification large language model. Based on the program function description information, the program source code, and the other program source codes, the code location information and first defect description information corresponding to the potential defect codes in the program source code are determined through the defect code identification large language model.

7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the target code segment containing the potentially defective code from the program source code based on the code location information includes: Determine the abstract syntax tree corresponding to the program source code; Based on the code location information, determine the target node in the abstract syntax tree corresponding to the potentially defective code; Starting from the target node, determine the target parent node in the abstract syntax tree that is closest to the target node, and / or the target child node that references the target node, wherein the target parent node is a function definition node; The target code fragment is determined based on the first source code corresponding to the nodes between the target parent node and the target node, and / or the second source code corresponding to the nodes between the target node and the target child node.

8. A code analysis method, characterized in that, include: Receive a request triggered by a client device by calling the code analysis service provided by the server device. The request includes program function description information corresponding to the application and the program source code to be analyzed. Based on the program function description information and the program source code, the code location information and the first defect description information corresponding to the first potential defect code in the program source code are determined by using the defect code identification large language model. Based on the code location information, a target code segment containing the potential defective code is obtained from the program source code, wherein the source code contained in the target code segment has complete semantic information; Based on the first defect description information and the target code fragment, the location of the target defect in the target formal model corresponding to the target code fragment is determined. The target formal model is used to reflect the code execution flow and code state transition of the target code fragment. The location of the target defect corresponds to the defect state described by the first defect description information. Based on the target formal model and the target defect location, the target defect code existing in the program source code is determined; The target defect code is then fed back to the client device for display.

9. An electronic device, characterized in that, include: The system includes a memory, a processor, and a communication interface; wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the code analysis method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor of an electronic device, causes the processor to perform the code analysis method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, include: A computer program or instruction that, when executed by a processor of an electronic device, causes the processor to perform the code analysis method as described in any one of claims 1 to 8.