A code extraction and analysis method

By extracting and classifying code features and handing them over to the large language model for processing, the problem that large language models are difficult to deal with huge code bases is solved, and the code reading and analysis efficiency is improved.

CN119781777BActive Publication Date: 2025-05-09ZHOUPU DATA TECH NANJING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510267595.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-09
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In the prior art, large language models are difficult to deal with large-scale and complex code bases, resulting in inefficient code reading, analysis and troubleshooting, especially when team personnel change frequently.

Method used

By traversing the code base, code features are extracted, including class information, control layer information, mapping layer information and service layer information, and classified and stored, and finally the extracted code is handed over to the large language model for processing.

Benefits of technology

The code scale reduction is achieved, allowing large language models to complete code reading, analysis and document tasks more efficiently, thereby improving software development efficiency and reducing the burden of manual reading and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781777B_ABST
    Figure CN119781777B_ABST
Patent Text Reader

Abstract

The present invention provides a code extraction and analysis method, comprising the following steps: S1: traversing a code base and copying all codes; S2: extracting code features in a corresponding manner; S3: storing the extracted code files; S4: extracting preset codes from the code files; S5: extracting referenced codes from the code files; S6: assembling results and sending them to a model; S7: calling the model and completing the output; using a large language model instead of manual reading and analysis of codes can greatly improve the efficiency of code reading and analysis; accurately extracting the required key codes, so that the large language model can complete the code analysis work within the range of logical complexity that it is best at; extracting codes across projects and services, so that code analysis is no longer an isolated reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code extraction and analysis, and in particular to a code extraction and analysis method. Background Art

[0002] Large language models (LLMs) are good at handling small-scale code processing tasks. However, in reality, most system structures are complex, often involving dozens or even hundreds of components and millions or even tens of millions of lines of code. If this information is handed over to large language models for processing, it will be expensive and often unable to successfully complete the related tasks.

[0003] During the design and coding phases of software development, more than half of the time is spent reading, analyzing, and debugging legacy code. This loss is exacerbated by frequent team turnover. If this code is distributed across multiple components, reading it becomes inherently challenging. The functionality that requires tweaking in daily development tasks is often not complex; it's the sheer volume of legacy code that creates the difficulty.

[0004] To reduce this loss, existing technologies have begun to incorporate large language models to process this portion of code. However, the difficulty encountered in existing technologies is that the code base is too large and involves too much content, making it impossible for large language models to process it smoothly. Therefore, it is necessary to extract the code related to the processing task and reduce its size until the large language model can complete tasks such as reading, debugging, and documenting this scale of code, thereby saving a lot of time for software developers. Summary of the Invention

[0005] In order to solve the defects and deficiencies in the prior art, the present invention provides a code extraction and analysis method.

[0006] The specific solution provided by the present invention is:

[0007] A code extraction and analysis method, characterized in that it includes the following steps:

[0008] S1: Go through the codebase and copy all the code;

[0009] S2: Use the corresponding method to extract code features;

[0010] S3: stores extracted code files;

[0011] S4: extracting preset codes from code files;

[0012] S5: Extract reference code from code files;

[0013] S6: Assemble the results and send them to the model;

[0014] S7: Call the model and complete the output.

[0015] As a further preferred embodiment of the present invention, in step S1, a timer is used to actively and regularly trigger a preset script to traverse the code library and copy all codes.

[0016] As a further preferred embodiment of the present invention, in step S1, all codes are copied only from the preset branch.

[0017] As a further preferred embodiment of the present invention, in step S2, the extracted code features at least include code features in class information, control layer information, mapping layer information and service layer information.

[0018] As a further preferred embodiment of the present invention, step S2 includes the following steps:

[0019] S2.1: Extract code features of class information: Traverse all Java files and extract each Java file into compilation units. Extract class information and interface information from the compilation units to generate storage files for the corresponding modules.

[0020] S2.2: Extract code features of interface information: Determine whether the class information is service layer information from its code features, and then obtain the code features of the corresponding interface through the interface implementation information and the import package corresponding to the interface;

[0021] S2.3: Extract code features of framework information: first parse the markup language to obtain the preset space in the mapping layer information, and find the interface information corresponding to the mapping layer information through the preset space, then parse the preset tag to obtain the code features corresponding to the preset tag, and introduce code features with preset fields; if the code features corresponding to the preset tag are not obtained, use the preset statements in the markup language to supplement the code features of the mapping layer information.

[0022] As a further preferred embodiment of the present invention, in step S3, the extracted code features are classified and stored in the corresponding modules in a preset file manner.

[0023] As a further preferred embodiment of the present invention, in step S4, the preset code at least includes java code, sql code and sql code in markup language.

[0024] As a further preferred embodiment of the present invention, in step S4,

[0025] When extracting Java code, extract the Java code with the same method name and parameter list as the preset method from the code file;

[0026] When extracting SQL code, extract the SQL code with preset tags from the code file;

[0027] When extracting SQL code from markup language, first extract data with the same ID as the interface information from the mapping layer information, then extract data with preset tags from the data, and introduce code features with preset fields.

[0028] As a further preferred embodiment of the present invention, step S5 includes the following steps:

[0029] S5.1: Delete redundant symbols and unify code format;

[0030] S5.2: Extract static reference code using regular expressions;

[0031] S5.3: Use regular expressions to extract local reference codes.

[0032] As a further preferred embodiment of the present invention, in step S6, when assembling the results, the entry method is placed at the front, and the codes are sorted according to the call chain and the call frequency.

[0033] Compared with the existing technology, the present invention can achieve the following technical effects:

[0034] 1) The present invention provides a code extraction and analysis method that uses a large language model to replace manual code reading and analysis, which can greatly improve the efficiency of code reading and analysis.

[0035] 2) The present invention provides a code extraction and analysis method that accurately extracts the required key codes, enabling the large language model to complete the code analysis work within the range of logical complexity that it is best at.

[0036] 3) The present invention provides a code extraction and analysis method, which extracts code across projects and services, making code analysis no longer an isolated reading.

[0037] 4) The present invention provides a code extraction and analysis method with strong scalability. Data states can be added later. By extracting data in a configuration state, problem analysis is more accurate. Other types of large language models can also be used at any time.

[0038] 5) The present invention provides a code extraction and analysis method with low cost. The task can be successfully completed by using only cloud services without the need for special training models, saving time and effort.

[0039] 6) The present invention provides a code extraction and analysis method, which uses corresponding different methods to extract code features of different information, and then classifies and stores them in the form of different code files, thereby improving the extraction efficiency and accuracy while facilitating the subsequent calling process of the code file; at the same time, by using different methods to extract the required different codes, the effectiveness and reliability of code extraction and subsequent code analysis are further improved, and the efficiency and accuracy of code analysis are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 Shown is a flow chart of the method steps provided by the present invention;

[0041] Figure 2 Shown is a logic flow chart of code extraction of the present invention;

[0042] Figure 3 The figure shows a schematic diagram of the code for extracting class information and interface information in step S2.1 of the present invention;

[0043] Figure 4 The figure shows a schematic diagram of extracting only some types of codes that affect code extraction in step S2.1 of the present invention;

[0044] Figure 5 The figure shows a schematic diagram of the header definition code of a typical service code in step S2.2 of the present invention;

[0045] Figure 6 The following is a schematic diagram of code for extracting code features from interface information in step S2.2 of the present invention:

[0046] Figure 7 The figure shows a code diagram of the markup language xml of the typical framework information mybatis in step S2.3 of the present invention;

[0047] Figure 8 The figure shows a schematic diagram of code for extracting code features from framework information in step S2.3 of the present invention;

[0048] Figure 9 The figure shows a schematic diagram of the code in which the SQL code is written in the declaration of the interface class in step S2.3 of the present invention;

[0049] Figure 10 The figure shows a schematic diagram of the code for extracting Java code in step S4 of the present invention;

[0050] Figure 11 The figure shows a schematic diagram of the code for extracting the SQL code in step S4 of the present invention;

[0051] Figure 12 The figure shows a schematic diagram of code for extracting the SQL code in the markup language in step S4 of the present invention;

[0052] Figure 13 The figure shows a schematic diagram of code for extracting static reference codes using regular expressions in step S5.2 of the present invention;

[0053] Figure 14 The figure shows a code diagram of extracting local reference codes using regular expressions in step S5.3 of the present invention;

[0054] Figure 15 Shown is a schematic diagram of the final output result of the present invention. DETAILED DESCRIPTION

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0056] In the description of the present invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front end," "rear end," "both ends," "one end," "the other end," and the like, indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limiting the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0057] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "provided with," "connected," etc., should be understood in a broad sense. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection or an electrical connection; it may refer to a direct connection or an indirect connection through an intermediate medium; it may refer to internal communication between two components. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0058] [First embodiment]

[0059] Since the core business of the current springboot framework and mybatis framework are concentrated in sql code and java code, take the process of obtaining sql code and java code from the framework that includes both springboot and mybatis as an example, Figure 1The first embodiment of the present invention is shown, which provides a code extraction and analysis method, comprising the following steps:

[0060] S1: Traverse the code base and copy all the code; Figure 2 As shown, a timer can be used to actively trigger a preset script. In this embodiment, the code base can be selected from, for example, a git database, and the preset script corresponds to a git script. The triggering frequency can be set according to actual needs, such as actively triggering once every day or every week, etc., to traverse the code base and copy all the codes.

[0061] It is worth noting that in this embodiment, all codes are copied only from the preset branch, which can be set to, for example, the master branch, because the master branch is the main branch, which is used for all subsequent work, so it can be set to copy all codes only from the master branch.

[0062] S2: extract code features in a corresponding manner; the extracted code features at least include code features in class information, controller information of the control layer, mapper information of the mapping layer, and service information of the service layer.

[0063] Since the target code is a Java framework project based on Spring Boot, although the basic component unit of Java is the class, Spring Boot has also added many other unit concepts, such as the controller layer, service layer, component layer Component API, etc., which can essentially be regarded as a class. Therefore, when designing the code extraction logic, they are uniformly processed in a class manner.

[0064] According to this classification method, there are still many unit concepts, such as Controller Rest, Controller Request, Mapping Get, and Mapping Post. Based on this, they are simplified into the following three types:

[0065] 1) Single class type: such as static method, Component method, etc.;

[0066] 2) Interface type: such as API Service, etc.

[0067] 3) Framework type: for example, the interface type and XML implementation type in MyBatis;

[0068] Those skilled in the art know that, for existing code types such as aspects and declarations, since they are usually codes related to the code framework, they are relatively rarely used in business and usually do not change the logical meaning of the business code. Therefore, they are not considered in this embodiment.

[0069] The code features extracted in this embodiment are essentially the code features corresponding to the above type information;

[0070] Therefore, in step S2, the following steps are included:

[0071] S2.1: Extracting code features of class information: Since the basic unit of Java code is class, we use the com.github.javaparser class library to traverse all Java files and extract each Java file into a compilation unit, namely CompilationUnit. Using compilation units can effectively reduce complexity. After entering the compilation unit, we can extract only the core code according to the compilation rules, while ignoring a large number of other irrelevant details. For example, in this embodiment, since there is a lot of information in the compilation unit, we only need to extract the class information class and interface information interface, such as Figure 3 As shown, to generate the storage file of the corresponding module;

[0072] It is worth noting that in order to avoid the tedious process and time consumption that may result from processing numerous node types, in this embodiment, only some types that affect code extraction are extracted for processing, such as Figure 4 As shown, the remaining types that do not affect code extraction are handled by the larger language model that is better at it.

[0073] S2.2: Extract code features of interface information: Determine whether the class information is service layer information from its code features, and then obtain the code features of the corresponding interface through the interface implementation information and the import package corresponding to the interface;

[0074] After extracting the interface and class, the class declaration will include the service layer Service label, such as Figure 5 As shown, by judging whether the code feature of the class information contains the service layer Service label, it can be determined whether it is service layer service information. Since the service layer information can implement one or more interfaces, and the interface must be imported through the corresponding import package, the code feature of the corresponding interface can be obtained through the interface implementation information and the import package corresponding to the interface, as shown in Figure 6 As shown;

[0075] S2.3: Extract code features of framework information:

[0076] Since the typical framework information mybatis markup language xml has an obvious format, for example: Figure 7 As shown, it must be = <mapper namespoace="”xxxxxx.xxx.x.xx”">This will facilitate the identification of whether it is the mapping layer mapper information, and it must contain <insert> 、 <delete>etc.

[0077] Therefore, we first parse the markup language XML to obtain the preset space namespace in the mapping layer mapper information, and then find the interface information corresponding to the mapping layer mapper information through the preset space namespace, and then parse the preset tag (for example <insert> 、 <delete>etc.) to obtain the code features corresponding to the preset tag, and introduce the code features with preset fields, such as field include. If there is include in the middle, you need to further introduce the sql code entity of include, such as Figure 8 As shown;

[0078] If the code feature corresponding to the preset tag is not obtained, the preset statement in the markup language is used to supplement the code feature of the mapping layer information; this is because the code of the mapping layer mapper information is special, because the SQL code can be written in the markup language XML or in the declaration of the interface class, such as Figure 9 As shown, therefore, correspondingly, the logic of the code processing here is to parse the markup language xml first, and then parse the interface class. If the interface class code does not contain <insert> 、 <update>The preset tags are supplemented with sql code in the markup language xml to synthesize a new mapping layer mapper object.

[0079] S3: Store the extracted code files; classify and store the extracted code features in the corresponding modules in a preset file format. The preset file format can be, for example, a json file; the classification storage method needs to correspond to the extracted code feature category. For example, the class information is stored in the json file of Allclassbean, the controller information of the control layer is stored in the json file of Allcontrollerbean, the mapper information of the mapping layer is stored in the json file of Allmapperbean, and the service information of the service layer is stored in the json file of Allservicebean, so as to facilitate classification storage and subsequent calling process, such as Figure 2 As shown;

[0080] S4: extracting the preset code from the code file; in this embodiment, the preset code includes at least java code, sql code and sql code in the markup language; wherein,

[0081] When extracting Java code, you can use the abstract syntax tree method to extract Java code with the same preset method name and parameter list from the code file, such as Figure 10 As shown;

[0082] When extracting SQL code, extract the SQL code with preset tags from the code file; the preset tags include <insert> 、 <update>Etc., the code containing these preset tags is considered to be SQL code, such as Figure 11 As shown;

[0083] When extracting SQL code in markup language, first extract the data with the same ID as the interface information from the mapping layer mapper information, and then extract the data with preset tags (such as <insert> 、 <delete>etc.) and introduce code features with preset fields, such as field include. If there is include in the middle, it is necessary to further extract the SQL code associated with include, such as Figure 12 shown.

[0084] S5: Extract referenced code from the code file by finding the location where other code is referenced in the code, such as package name, class name, method name, etc., to find the referenced code in the extracted code. This specifically includes the following steps:

[0085] S5.1: Delete redundant symbols and unify the code format. For example, line breaks, \t, and extra spaces in the code will affect the subsequent regular expression extraction. First remove symbols that do not affect the code semantics but affect the regular expression extraction to further ensure the smooth and accurate extraction of the subsequent regular expression.

[0086] S5.2: Extract static reference code using regular expressions;

[0087] Regular Expression is a text pattern matching tool that searches, matches, and operates on text by defining specific patterns. For example, in this embodiment, regular expressions can be used to extract codes such as String companyName=companyService.companyName(cid); This type of code represents static reference code, such as service, component, api reference, etc. Figure 13 As shown;

[0088] S5.3: Similarly, regular expressions can be used to extract local reference codes, such as Figure 14 shown.

[0089] S6: Assemble the results and send them to the model; after all the relevant codes are extracted and before being handed over to the large language model, the codes need to be sorted and sorted. Therefore, when assembling the results, the entry method needs to be placed at the front, and the codes need to be sorted according to the call chain and the frequency of calls. This is because the entry represents the entrance to the business, so that the output information is more in line with the current business goals. According to actual needs, the service method can also be placed at the front, but this sorting method is more used for correlation analysis. Correspondingly, the prompt words handed to the large language model will also become similar to "Analyze this service, if it makes xxx modifications, will it cause which codes to need to be modified?", and the requirements of the code extraction and analysis method in this embodiment are: to be able to answer what the business logic of this interface is, so the entry method needs to be placed at the front.

[0090] S7: Call the model and complete the output. The code extraction and analysis results implemented by the method of this embodiment are as follows: Figure 15 As shown in Figure 2, using a large language model instead of manual code reading and analysis can greatly improve the efficiency of code reading and analysis.

[0091] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.< / delete> < / insert> < / update> < / insert> < / update> < / insert> < / delete> < / insert> < / delete> < / insert> < / mapper>

Claims

1. A code extraction and analysis method, characterized in that: The following steps are involved: S1: Go through the code base and copy all the code; S2: Use the corresponding method to extract code features; S3: stores the extracted code files; S4: extracting preset codes from the code file; S5: Extract reference code from code files; S6: Assemble the results and send them to the large language model; S7: Call the large language model and complete the output; In step S2, the extracted code features at least include code features in class information, control layer information, mapping layer information and service layer information; The step S2 includes the following steps: S2.1: Extract code features of class information: traverse all Java files, extract each Java file into a compilation unit, extract class information and interface information from the compilation unit to generate storage files of the corresponding module; S2.2: Extract code features of interface information: determine whether the class information is service layer information from its code features, and then obtain the code features of the corresponding interface through the interface implementation information and the import package corresponding to the interface; S2.3: Extract code features of framework information: first parse the markup language to obtain the preset space in the mapping layer information, and find the interface information corresponding to the mapping layer information through the preset space, then parse the preset tag to obtain the code feature corresponding to the preset tag, and introduce the code feature with the preset field; if the code feature corresponding to the preset tag is not obtained, use the preset statement in the markup language to supplement the code feature of the mapping layer information; In the step S4, the preset code at least includes java code, sql code and sql code in markup language.

2. A code extraction and analysis method according to claim 1, characterized in that: In the step S1, a timer is used to actively and regularly trigger a preset script to traverse the code base and copy all codes.

3. A code extraction and analysis method according to claim 1, characterized in that: In step S1, all codes are copied only from the preset branch.

4. A code extraction and analysis method according to claim 1, characterized in that: In the step S3, the extracted code features are classified and stored in the corresponding module in a preset file manner.

5. A code extraction and analysis method according to claim 1, characterized in that: In the step S4, When extracting Java code, extract the Java code with the same method name and parameter list as the preset method from the code file; When extracting SQL code, extract the SQL code with preset tags from the code file; When extracting the SQL code in the markup language, first extract the data with the same ID as the interface information from the mapping layer information, then extract the data with preset tags from the data, and introduce the code features with preset fields.

6. A code extraction and analysis method according to claim 1, characterized in that: In the step S5, The following steps are involved: S5.1: Delete redundant symbols and unify the code format; S5.2: Extract static reference code using regular expressions; S5.3: Use regular expressions to extract local reference codes.

7. A code extraction and analysis method according to claim 1, characterized in that: In step S6, when assembling the results, the entry method is placed at the front, and the codes are sorted according to the call chain and the call frequency.

Citation Information

Patent Citations

  • Large-scale code data feature extraction method and system

    CN117113347A

  • Automated translation of computer languages to extract and deploy computer systems and software

    US20230393832A1