Code interpretation assistant system
Through the code interpretation assistant system, using static code interpretation and knowledge question and answer subsystems, the difficult problems of source code knowledge mining and question and answer are solved, and the business logic can be quickly understood and flow charts can be generated, thereby improving the utilization efficiency of code assets.
Patent Information
- Application Number
- CN202510981738.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies lack effective source code-based knowledge mining and knowledge question-answering capabilities, making it difficult to fully utilize code assets to interpret business scenario processes and draw related business process diagrams.
A code interpretation assistant system is provided, including a static code interpretation subsystem and a knowledge question and answer subsystem. The code interpretation engine interprets the source code to generate a human language version of the interpretation report and business process diagram, and builds a code knowledge question and answer library for knowledge question and answer.
It enables rapid understanding of business logic and flow chart drawing based on source code, can generate a human language version of the interpretation report that matches the transaction code, and answer business-related knowledge questions through the knowledge question and answer subsystem, filling the gap in existing technology.
Smart Images

Figure CN120653298A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a code interpretation assistant system. Background Art
[0002] In the era of big models, the techniques for building and using RAG knowledge bases based on human language text are relatively common, but the techniques for building and using knowledge bases based on source code are almost nonexistent. The greatest advantage of code assets over assets written in human language (such as business requirements specifications, software requirements specifications, outline designs, and detailed designs) is that they truly reflect the most realistic implementation. If human language documents are not updated promptly, their descriptions may not align with actual implementation, while code assets always reflect the real implementation. Therefore, extracting relevant business and technical knowledge from source code assets is a very promising approach.
[0003] Existing technologies rarely address source code-based knowledge mining. Specifically, they address: 1. How to fully leverage the value of source code to provide transaction code-based business scenario process interpretation and related business process diagramming capabilities; 2. How to leverage source code knowledge assets for knowledge Q&A. These two areas are currently unavailable in the market. Summary of the Invention
[0004] To address the aforementioned issues, the present invention aims to provide a code interpretation assistant system to enable source code-based knowledge mining. The present invention's code interpretation assistant system primarily addresses two types of problems: First, when a business operator inquires about the implementation logic of a specific transaction code, the code interpretation assistant can generate a human-language interpretation report that matches the transaction code and, based on this, draw a corresponding business process diagram. Second, it creates a knowledge question-and-answer scenario based on the knowledge contained in the code. When a business operator or technician inquires about business-related knowledge, the code interpretation assistant can query the corresponding code knowledge question-and-answer knowledge base, recall matching information, and generate a corresponding knowledge-based answer.
[0005] The present invention provides a code interpretation assistant system, which includes a static code interpretation subsystem and a knowledge question answering subsystem;
[0006] The code interpretation system is used to interpret the transaction code of the core system based on the transaction code, and to interpret the transaction rules and business process diagram in the language and text version that matches the transaction code through the code interpretation engine; the knowledge question and answer subsystem is used to build a code knowledge question and answer library to conduct code knowledge question and answer in combination with the knowledge code question and answer library.
[0007] Optionally, the code interpretation subsystem implements static interpretation of code assets through a code interpretation engine, and the code interpretation engine includes a static code asset parsing module and an atomic capability building module;
[0008] The static code asset parsing module is used to obtain transaction-related static code assets based on the core system's transaction code. The static code asset parsing module includes metadata model parsing and Java file parsing. The metadata model parsing includes obtaining multiple definition contents from the metadata model based on the transaction code. The Java file parsing includes searching for transaction-related Java source files based on the definition contents.
[0009] The atomic capability analysis module is used to perform single-method business rule generation, service business rule generation, service flow chart generation, service flow chart generation, transaction flow chart generation, transaction / service related error code identification, transaction / service related business table lineage relationship identification, and transaction process orchestration identification.
[0010] Optionally, the process of static interpretation of code assets by the code interpretation engine includes: parsing the static code assets obtained based on the transaction code, obtaining the corresponding Java files by processing the metamodel; analyzing the Java files based on the abstract syntax tree AST, and establishing associations between multiple files; removing Java files that do not need to be interpreted based on the filtered file configuration, and finally interpreting the remaining Java files after filtering; wherein, the metamodel data will be stored in mongodb, and the results of the code interpretation will be placed in the code interpretation RAG knowledge base for retrieval.
[0011] Optionally, the process of interpreting a transaction code given by a user by the code interpretation engine includes: receiving a transaction code input by the user, retrieving a corresponding metadata model based on the transaction code in MongoDB, performing a RAG search in the code knowledge base, and obtaining matching information; drilling down on the calling method to obtain a list of other methods called by the service method; determining whether the transaction is multi-node, and if so, continuing to loop the process of retrieving the metadata model based on the transaction code and obtaining a list of other methods called by the service method; otherwise, assembling the output; inputting the main method and other methods obtained by drilling down into the context of the code interpretation large model, and using the code interpretation large model to form an interpretation text.
[0012] Optionally, obtaining a list of other methods for calling a service method includes: obtaining an import list and a package name from a specified Java source file, and then obtaining the current class name and the parent class name inherited by the current class name; locating and obtaining the source code of the method body of the specified method based on the abstract syntax tree and converting it into a string; preprocessing the content of the method body; preprocessing includes removing specified format content information; obtaining the input parameter information of the method body, and using regular expressions to match various method calls, including calls to static methods of tool classes, calls to static methods of the current class, chain calls, super() method calls, this() method calls, parameter object method calls, method calls to obtain a bean object using getBeans(beanName), and method calls to other objects using the dot (.) operator; uniformly processing the call relationship, specifically by checking the fully qualified class name, processing static methods, processing object methods, and processing automatically injected objects; parsing the class name and method name, and constructing a call relationship string, deduplicating and sorting, and finally returning the result.
[0013] Optionally, the knowledge question and answer subsystem includes a code knowledge question and answer knowledge base, a Chinese-English comparison table knowledge base, and a synonym comparison table knowledge base; the code knowledge question and answer base includes dividing the code into chunks and recording its meta-information, and the meta-information includes file path, package name, and class name; the English comparison table knowledge base is used to store the mapping from Chinese to English; the synonym comparison table knowledge base is used to compare data dictionaries to unify semantics.
[0014] Optionally, the workflow of the knowledge question and answer subsystem includes: receiving original query information, retrieving the synonym table in the synonym table knowledge base, optimizing the terms of the original query information, and obtaining first optimized query information; retrieving the Chinese-English comparison table knowledge base, obtaining an English description that matches the first optimized query information, and obtaining second optimized query information; extracting package paths and method call relationships related to the second question list, and retrieving and recalling target code snippets related to the first optimized query information based on the knowledge question and answer knowledge base; inputting the original query information, the first optimized query information and / or the second optimized query information, and the target code snippet into the code interpretation model, and using the code interpretation model to generate a question and answer structure.
[0015] Optionally, the code knowledge question and answer library is constructed by importing a Java project and parsing the project structure, extracting flowtran information and java class information; processing the flowtran information and java class information to obtain a csv file, and importing the csv file into the code knowledge question and answer library.
[0016] Optionally, the flowtran information processing process includes: recursively traversing a specified project directory, filtering and processing the flowtrans.xml file; deleting specified useless XML nodes, interpreting the processed XML content with LLM and obtaining an LLM interpretation result; entering a data cleaning and conversion process, extracting basic transaction information such as transaction ID, transaction name, package name, extracting transaction description information and process service information, and generating transaction usage scenarios with LLM to review and correct errors based on a business scenario corpus;
[0017] After cleaning, two CSV files are generated. The first CSV file is defined as "Project Name-Transaction Summary Information.csv", which contains full_path metadata and transaction summary information. The second CSV file is defined as "Project Name-Transaction Details.csv", which contains full_path metadata, transaction summary information, and additional detailed transaction descriptions.
[0018] Optionally, the processing of the Java class information includes: recursively traversing the specified project directory, filtering out useless Java class files, and extracting the package name, class name, class declaration type and comments of the Java file; entering the data cleaning and conversion process, allowing LLM to parse the class field information, method information, comments and code implementation and review and correct them; slicing the parsing results and storing them in multiple chunk fragments. If a class method is relatively large, it will correspond to multiple chunk fragments and their order must be maintained. Therefore, the file's full_path metadata and serial number are used as identifiers to ensure the continuity of each class method chunk; and finally generating a CSV file containing all key information of the Java class.
[0019] The present invention provides a code interpretation assistant system, which fills the gaps in the field of code interpretation of business scenario processes for transaction codes and the drawing of related business process diagrams, as well as source code knowledge assets for knowledge question and answer. It can not only help business personnel quickly understand the implementation logic of a specific transaction code and visualize it with a flowchart, but also help business personnel and technical personnel to do various types of business knowledge questions and answers based on the most authentic source code. The code interpretation assistant system of the present invention, from the perspective of its responsibilities, mainly enables two scenarios: Scenario one, based on the transaction code of the core system, interprets the human language text version of the transaction rules and business process diagrams that match this transaction code; Scenario two, based on the knowledge contained in the code to do knowledge questions and answers. Because compared to document-type knowledge, the knowledge in the code is the most accurate. The above two scenarios are scenarios enabled by the self-developed code interpretation assistant of the present invention.
[0020] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0022] Figure 1 This is a schematic diagram of the overall architecture design of static code interpretation based on core system transaction codes according to an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of a code interpretation engine according to an embodiment of the present invention statically interpreting code assets and building a meta-model database and a code interpretation knowledge base;
[0024] Figure 3 This is a schematic diagram of the code interpretation engine interpreting a transaction code given by a user, combining a meta-model database and a code interpretation knowledge base according to an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of method call relationship extraction based on regular expressions according to an embodiment of the present invention;
[0026] Figure 5 This is a schematic diagram of extracting method call relationships based on a large model according to an embodiment of the present invention;
[0027] Figure 6 This is a schematic diagram of method call relationship extraction based on Java Abstract Syntax Tree (AST) in an embodiment of the present invention;
[0028] Figure 7 1 is a schematic diagram of calling relationship extraction using a regular expression-based method according to an embodiment of the present invention;
[0029] Figure 8 The embodiment of the present invention renders the content in the plantuml syntax format into a real flowchart;
[0030] Figure 9 It is a flowchart rendered from the content in plantuml format according to an embodiment of the present invention;
[0031] Figure 10 This is the complete process of building a code knowledge question and answer knowledge base in an embodiment of the present invention;
[0032] Figure 11This is a schematic diagram of a scenario in which a code knowledge question and answer knowledge base is combined with a code knowledge question and answer knowledge base to conduct question and answer on specific business knowledge in an embodiment of the present invention;
[0033] Figures 12a to 12d This is a schematic diagram of a relatively complex business knowledge question and answer animation presentation method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to illustrate the present invention and are not intended to limit the present invention.
[0035] An embodiment of the present invention provides a code interpretation assistant system, which includes a static code interpretation subsystem and a knowledge question and answer subsystem; the code interpretation system is used to interpret the transaction code of the core system based on the transaction code, and to interpret the language and text version of the transaction rules and business process diagram that match the transaction code through a code interpretation engine; the knowledge question and answer subsystem is used to build a code knowledge question and answer library to conduct code knowledge question and answer in combination with the knowledge code question and answer library.
[0036] The code interpretation assistant system of this embodiment of the present invention provides static interpretation of specific transaction codes. Specifically, given a transaction code, it can generate a human-language interpretation report and a corresponding business process diagram. It also provides knowledge-based Q&A based on the business knowledge contained in the code. Specifically, the code interpretation assistant system of this embodiment of the present invention not only helps business personnel quickly understand the implementation logic of specific transaction codes and visualize it with flowcharts, but also enables business personnel and technical personnel to answer various business knowledge questions using authentic source code. Each subsystem is described in detail below.
[0037] A transaction code is a number used by the system to distinguish different business processes, such as transfers and balance inquiries. Each function has its own transaction code, and upon receiving the transaction code, the system knows which business logic to execute. Static code assets refer to source code and related development files stored in a code repository or local environment, without compilation or execution.
[0038] The overall architecture design of the code interpretation subsystem of the embodiment of the present invention is as follows Figure 1 As shown, the code interpretation subsystem implements static interpretation of code assets through a code interpretation engine. The code interpretation engine is an independent application service responsible for parsing all (.java / .xml) class files in the code assets and invoking a large language model for interpretation, extracting structured information and storing it in a database. The code interpretation engine includes a static code asset parsing module and an atomic capability building module.
[0039] The static code asset parsing module is used to retrieve transaction-related static code assets based on the core system's transaction code. This module includes metadata model parsing and Java file parsing. The metadata model parsing involves obtaining multiple definitions from the metadata model based on the transaction code, while the Java file parsing involves searching for transaction-related Java source files based on the definitions. The atomic capability parsing module is used to generate single-method business rules, service business rules, service flow charts, transaction flow charts, transaction / service-related error code identification, transaction / service-related business table lineage identification, and transaction process orchestration identification.
[0040] The metadata model is a general term for the following XML types: flowtrans.xml - online transaction interface definition file, apsServiceType.xml - service function definition file, apsServiceImpl.xml - service implementation definition file, tables.xml - table structure definition file, error.xml - error code definition file, d_schema.xml - data item definition file, u_schema.xml - basic type definition file, e_schema.xml - enumeration type definition file, nsql.xml - custom SQL definition file, c_schema.xml - composite type definition file, batch_tran.xml - batch transaction definition file, file_batch_tran.xml - file batch transaction definition file.
[0041] The code interpretation engine's static interpretation of code assets involves first obtaining transaction-related static code assets based on the transaction code and then performing static interpretation on them. Assets associated with a specific transaction code are defined in the metadata model. Therefore, the engine first retrieves the following definitions from the metadata model based on the transaction code: transaction flowtrans, service type apsServiceType, service implementation class apsServiceImpl, error code type ErrorType, involved database tables Tables, composite type complexType, batch transaction batch_tran, file batch transaction batch_trans, named SQL namedsql, and data dictionary data_dict.
[0042] FlowTrans is the abbreviation for the flowtrans.xml file, which defines the online transaction interface. Online transactions are the system's external interface and the smallest unit that can be called by external systems. apsServiceType is the apsServiceType.xml file, which defines the service function and is used to create a common function. apsServiceImpl is the apsServiceImpl.xml file, which defines the service implementation code. ErrorType is the abbreviation for the error.xml file, which defines the error type and corresponding error message. Tables is the abbreviation for the tables.xml file, which defines the table structure and table operation methods. complexType is the abbreviation for the c_schema.xml file, which defines a complex type and defines a collection of multiple interface field elements. batch_tran is the abbreviation for the batch transaction definition file, which is used to define business procedures for batch processing. file_batch_tran is the abbreviation for the file_batch_tran.xml file, which defines the file batch transaction definition file and is used to parse or assemble files through batch transactions. 2. namedsql is the abbreviation for the nsql.xml file, which defines the custom SQL definition file and is used to manage and write complex SQL statements. data_dict is a general term for the d_schema.xml - data item definition file + u_schema.xml - basic type definition file + e_schema.xml - enumeration type definition file, also known as the data dictionary, which is used to manage data standards and define all basic field types.
[0043] After obtaining the corresponding definition from the metadata model, we search for the corresponding Java source files, thereby obtaining multiple Java source files related to this transaction. We need to pay special attention to the following: a. The relationship between multiple service definition files; b. The main method entry point; c. Method drilldown, that is, whether the drilled-down method calls other services or methods.
[0044] After static code analysis assets, we enter the atomic capability building phase. The main tasks are as follows:
[0045] Single-method business rule generation primarily involves analyzing the corresponding business rules for a specific method. This involves two steps. The first step is to generate code implementation rules from a technical perspective, such as specific database table operations and enabling distributed transactions. Technical interpretation is necessary to help technical developers better understand the implementation logic of the corresponding method. The second step is to generate business interpretation rules in the corresponding business language based on the code implementation rules from a technical perspective.
[0046] Service business rule generation: This mainly involves generating text descriptions of specific business rules for services. Business rules must have a hierarchical description, such as chapter-based descriptions such as 1.1 and 1.1.1. Service business rules are interpreted using a description method that favors business language.
[0047] Service flow chart generation: mainly focuses on converting the business rules described in the generated text into a flow chart.
[0048] Transaction flow chart generation: mainly generates corresponding flow charts from the transaction dimension.
[0049] Transaction / service-related error code identification: mainly identifies transaction / service-related error codes.
[0050] Identification of the blood relationship between transaction / service-related business tables: This mainly identifies the relationship between transaction / service-related business tables.
[0051] Transaction process orchestration identification: mainly used to identify transaction-related process orchestration.
[0052] Once the code interpretation engine has parsed the above capabilities, it can enable a number of scenarios, such as generating transaction interpretation reports and detailed design documents.
[0053] The code interpretation engine of the embodiment of the present invention performs static interpretation on code assets and builds a meta-model database and a code interpretation knowledge base. Figure 2 The meta-model database is used to store the data content after the metadata model file is parsed by the code interpretation engine. The database type used is MongoDB.
[0054] First, associate code assets. Code assets can be offline source code zip archives, configured Git code repositories, or integrated with CI / CD pipelines. The code interpretation engine first parses the assets by processing the metadata model, extracting definitions such as FlowTrans, apsServiceType, apsServiceImpl, ErrorType, Tables, complexType, batch_tran, file_batch_tran, namedsql, and data_dict. Based on the metadata model, it then locates the corresponding Java files.
[0055] The found Java files are then analyzed based on the Abstract Syntax Tree (AST) and associations are established between multiple files. Based on the filtered file configuration, files that do not require interpretation are removed. Finally, interpretation is performed on the remaining filtered files. Metamodel data is stored in MongoDB, and the code interpretation results are stored in the Code Interpretation RAG knowledge base for retrieval.
[0056] In an optional embodiment of the present invention, the Java file abstract syntax tree parsing method can be as follows:
[0057] 1. Use the Python open source library javalang to parse the syntax tree of the Java file and decompose the Java file into class name, package name, member variables, method definition, method parameters, and class modifiers.
[0058] 2. Configure the file in advance and determine whether the current file needs to be parsed based on the file configuration content. File configuration standards: a. Java package prefixes to be parsed, multiple package configurations are separated by commas; b. Java package prefixes not to be parsed, multiple package configurations are separated by commas; c. Specify the Java class names that do not need to be parsed, using regular expressions, multiple regular expressions are separated by commas.
[0059] 3. Store the structured disassembled Java file into the MongoDB database.
[0060] Use the following code to parse the AST abstract syntax tree:
[0061]
[0062]
[0063]
[0064] File configuration in advance:
[0065]
[0066] The structured data stored in MongoDB after the abstract syntax tree is parsed is as follows:
[0067]
[0068] The embodiment of the present invention combines the meta-model database and the code interpretation knowledge base. The code interpretation engine interprets the transaction code given by the user. The overall process is as follows: Figure 3 shown.
[0069] After receiving the transaction code entered by the user, the code interpretation engine performs an asset search based on the transaction code. First, it searches MongoDB for the corresponding metadata model based on the transaction code. After obtaining the metadata model, it then performs a RAG search in the code knowledge base, combining information such as ServiceType and ServiceImpl, to obtain matching information, particularly the main method entry point (the function method corresponding to the Java file of the ServiceImpl service implementation class). After the search, it drills down into the calling method (only one level down) to obtain a list of other methods called by this service method and determines whether the transaction is multi-node (i.e., multiple serviceType service nodes can be orchestrated under a transaction, and each serviceType node will be processed in a loop). If so, the aforementioned "RAG search in the code knowledge base, combining information such as ServiceType and ServiceImpl, to obtain matching information" process continues. Otherwise, the output is assembled, placing the main method and the drilled-down methods in the context of the code interpretation model, allowing the code interpretation model to form a text interpretation based on these methods.
[0070] In an optional embodiment of the present invention, the code interpretation knowledge base construction scheme is as follows:
[0071] 1. The data dictionary metadata structured by the code interpretation engine is stored in the knowledge base in the form of a single JSON entry. The knowledge base segment identifier is: \n, the maximum segment length is 500 tokens, and the segment overlap length is 50 tokens.
[0072] 2. Parse the abstract syntax tree of the Java class file in the code interpretation engine, then call the large language model according to the set prompt, and store the output of the model in the knowledge base.
[0073] A single data dictionary is stored in the knowledge base structure:
[0074]
[0075] Optionally, the example process of the transaction code input by the user being parsed by the code interpretation engine to generate the interpreted text is as follows:
[0076] 1. The user enters the transaction code dp2130 - Current Deposit.
[0077] 2. Search MongoDB and find the corresponding metadata model:
[0078]
[0079] 3. Use ServiceType and ServiceImpl to perform RAG search in the code knowledge base. Find the TransferServiceImpl class in the knowledge base and find its main method. Transfer is the main method entry: publicTransferResulttransfer(TransferRequestrequest)
[0080] 4. Drill down to the main method transfer (one level only) to analyze which other methods it directly calls:
[0081] -`validateRequest(request)`
[0082] -`calculateFee(request)`
[0083] - `executeTransfer(request)`
[0084] - `recordTransaction(result)`
[0085] 5. Put the main method transfer and the drill-down methods validateRequest, calculateFee, executeTransfer, and recordTransaction into the context of the code interpretation model.
[0086] 6. Finally, the code interpretation model generates interpretation text:
[0087] The main method for transaction `TX123456` is `transfer`, which is responsible for processing the transfer request. First, the `validateRequest` method validates the request parameters; then the `calculateFee` method is called to calculate the fee; then the `executeTransfer` method executes the actual transfer logic; and finally, the `recordTransaction` method is called to record the transaction details. The overall process is clear and conforms to the general processing logic of transfer transactions.
[0088] In an optional embodiment of the present invention, when obtaining a list of other methods called by a service method, a regular expression-based interpretation method can be used, combined with Java language rules to analyze various call relationships and obtain a correct method call list. The entire logic is as follows Figure 4 shown.
[0089] First, the import list and package name are retrieved from the specified Java source file. The current class name and, optionally, the inherited parent class name are obtained. Then, based on the abstract syntax tree, the source code for the specified method body is located and converted to a string. The method body is preprocessed, including removing single-line comments, multi-line comments, end-of-line comments, bizlogs, and other specified formatting information. Next, the method body's input parameters are obtained and the matching of various method calls begins. Given the complexity of the Java language, the following scenarios exist: a. Calling static methods of utility classes; b. Calling static methods of the current class; c. Chaining; d. calling the super() method; e. calling the this() method; f. Calling parameter object methods; g. Calling methods on a bean object using getBeans(beanName); and h. Calling methods on other objects using the dot (.) operator.
[0090] Each case requires a regular expression match. After matching, the call relationship is processed uniformly. Specifically, the fully qualified class name is checked, followed by static methods, object methods, and finally automatically injected objects. The class and method names are then parsed, and a call relationship string is constructed. Duplicates are removed, sorted, and the result is returned.
[0091] The following describes the Java method call relationship analysis using scripts and actual cases:
[0092] S1, according to the regular expression matching code method tool class static method call, the regular expression is "([AZ][a-zA-Z0-9_]+)\.([a-zA-Z0-9_]+)\(".
[0093] S2, matches the corresponding method call in the code method according to the regular expression, the regular expression is "([a-zA-Z0-9_]+(?:\.[a-zA-Z0-9_]+)*?)\.([a-zA-Z0-9_]+)\(".
[0094] S3, matches the automatically injected objects in the code method according to the regular expression, the regular expression is "@Autowired\s+private\s+(\w+)\s+".
[0095] S4, matches method parameters, ordinary variables, and class member variables according to regular expressions. The regular expression for matching method parameters is "public\s+\w+\s*\((.*?)\)", the regular expression for matching ordinary variables is "(\w+)\s+{parts[0]}\s*[=;]", and the regular expression for matching class member variables is "private\s+(\w+)\s+{parts[0]}".
[0096] Sample input for extracting Java method call relationships:
[0097]
[0098] Sample output for extracting Java method call relationships:
[0099]
[0100]
[0101] The following is a script for extracting Java method call relationships:
[0102]
[0103]
[0104] The regular expression-based method call relation extraction in the embodiment of the present invention is 60 to 70 times faster in performance than the extraction based on a large model and the extraction based on an abstract syntax tree.
[0105] Taking the doMain() method of the DpCurAcctDept class as an example, it takes 7.5479 seconds to extract the method call relationship based on the large model. Figure 5 As shown in Figure 2. And the accuracy of this solution is very low because many method call relationships are not extracted. Using the method call relationship extraction based on Java Abstract Syntax Tree (Ast) takes 6.233 seconds, as shown in Figure 2. Figure 6 The embodiment of the present invention uses a regular expression-based method to call relationship extraction, which only takes 0.1026 seconds and is about 60 to 70 times more efficient than the previous two methods. Figure 7 As shown in the figure, method call extraction is a high-frequency operation in the code interpretation process, so improving the speed can greatly improve the overall productivity.
[0106] In an embodiment of the present invention, a PlantUML syntax format can also be generated based on the text content of the code interpretation and rendered as a flowchart. PlantUML is the syntax format of the corresponding flowchart, so the PlantUML syntax format can be generated based on the text content obtained from the code interpretation. The method is to complete it through a thought chain prompt with an example. This prompt is as follows:
[0107] You are an expert with extensive banking and finance experience and are proficient in PlantUML syntax. Your goal is to help users understand business processes and provide a comprehensive overview of the entire business process. Your task is to use the PlantUML DSL to generate a transaction flow diagram based on the transaction description provided by the user.
[0108] <Specific requirements>
[0109] 1. Use PlantUML Activity Diagram syntax to describe business processes.
[0110] 2. Generate a single, complete flowchart without splitting it.
[0111] 3. Ensure the correctness of PlantUML syntax and logic description.
[0112] 4. Strictly abide by the PlantUML syntax rules to ensure that the generated code can be correctly parsed by the PlantUML parser.
[0113] 5. Ensure that all message flows (arrows) are used correctly.
[0114] 6. Make sure there are no overhanging or unclosed structures.
[0115] 7. Simplify complex business logic into basic process steps.
[0116] 8. Use concise and clear descriptions to avoid formatting issues that may occur due to excessively long text.
[0117] 9. Limit the use of the following PlantUML elements: start, stop, if / else, while, fork, join, :action, note, partition.
[0118] 10. Use default styles and colors and do not add additional style instructions.
[0119] 11. Use notes appropriately to explain important steps or decision points, but don’t overuse them.
[0120] <Error Check List>
[0121] After generating the code, check the following points:
[0122] 1. Make sure start and stop appear only once.
[0123] 2. Check whether all if / else structures are closed correctly.
[0124] 3. Make sure all while loops have clear end conditions.
[0125] 4. Check if there are any dangling nodes or arrows.
[0126] 5. Ensure all notes are correctly attached to the corresponding activities or decision points.
[0127] <output format>
[0128] Output only the complete PlantUML code, including the @startuml and @enduml tags. Do not include any explanation or extra text. Make sure the code blocks are properly encapsulated. The format is as follows:
[0129] @startuml
[0130] xxxx
[0131] xxxx
[0132] xxxx
[0133] @enduml
[0134] When generating code, be conservative and make sure that the generated code is 100% parseable by PlantUML, even at the cost of some complexity. If you have any doubts about a syntax construct, choose a simpler, more reliable alternative.
[0135] Please generate the corresponding PlantUML activity diagram code based on the following transaction description and reference examples:
[0136] <Transaction Description>
[0137]
[0138] Example reference:
[0139] The following is a simple PlantUML activity diagram example, including the use of note and partition, for reference only:
[0140]
[0141]
[0142] Finally, the Chinese interpretation content based on the code interpretation can be correctly converted into the PlantUML syntax format, and the content in the PlantUML syntax format can be rendered into a real flowchart by calling the PlantUML server, such as Figure 8 shown.
[0143] Finally, let’s take a look at the actual effect. For example, the business rules interpreted by the code interpretation engine are as follows:
[0144] This method is mainly used to maintain application platform parameters. The business rules are as follows:
[0145] This method is mainly used to maintain application platform parameters. The business rules are as follows:
[0146] 1. Convert application service input into domain object input
[0147] 2. Call parameter maintenance processing method
[0148] 2.1 Call the "Parameter Process Maintenance" method
[0149] 2.1.1 Deserialize the input data packet and obtain the data mapping
[0150] 2.1.2 Get the class and instance objects corresponding to the table
[0151] 2.1.3 Copy the data mapping to the instance object
[0152] 2.1.4 Get the parameter processor and call the pre-processing method
[0153] 2.1.5 Execute the corresponding database operation according to the operation identifier
[0154] 2.1.5.1 If the operation indicator is Add
[0155] 2.1.5.1.1 Check if the record exists. If it does, report an error "Record already exists"
[0156] 2.1.5.1.2 Inserting New Records
[0157] 2.1.5.1.3 Calling post-processing methods
[0158] 2.1.5.1.4 Registration Parameter Audit Log
[0159] 2.1.5.2 If the operation is marked as query
[0160] 2.1.5.2.1 Generating Query SQL Statements
[0161] 2.1.5.2.2 Execute the query and return the results
[0162] 2.1.5.3 If the operation is marked as update
[0163] 2.1.5.3.1 Query and update records
[0164] 2.1.5.3.2 Calling post-processing methods
[0165] 2.1.5.3.3 Registration Parameter Audit Log
[0166] 12.1.5.4 If the operation is marked as delete
[0167] 2.1.5.4.1 Query and delete records
[0168] 2.1.5.4.2 Calling post-processing methods
[0169] 2.1.5.4.3 Registration Parameter Audit Log
[0170] 2.1.5.5 If the operation identifier is illegal, an error message "Illegal operation identifier" will be displayed.
[0171] 2.1.6 Serialize the result data packet and return it
[0172] 3. Record logs and return results
[0173] ###Detailed steps
[0174] 1. **Convert application service input to domain object input**
[0175] - Convert `cplApParaMatainIn` to `ApParaMatainProc.Output` object.
[0176] 2. **Call parameter maintenance processing method**
[0177] 2.1**Call the "Parameter Process Maintenance" method**
[0178] -Deserialize the input data packet and obtain the data map.
[0179] -Get the class and instance objects corresponding to the table.
[0180] -Copy the data map to the instance object.
[0181] -Get the parameter processor and call the pre-processing method.
[0182] -Execute the corresponding database operation according to the operation identifier.
[0183] 2.1.5**Execute the corresponding database operation according to the operation identifier**
[0184] 2.1.5.1**If the operation indicator is Add**
[0185] - Check if the record exists. If it does, report an error "Record already exists".
[0186] -Insert a new record.
[0187] -Call post-processing method.
[0188] - Register parameters to audit log.
[0189] 2.1.5.2 **If the operation is marked as query**
[0190] -Generate query SQL statements.
[0191] -Executes the query and returns the results.
[0192] 2.1.5.3 **If the operation is marked as update**
[0193] -Query records and update them.
[0194] -Call post-processing method.
[0195] - Register parameters to audit log.
[0196] 2.1.5.4**If the operation is marked as delete**
[0197] -Query records and delete them.
[0198] -Call post-processing method.
[0199] - Register parameters to audit log.
[0200] 2.1.5.5 **If the operation identifier is illegal**
[0201] - Report an error "Illegal operation identifier".
[0202] 2.1.6 **Serialize the result data packet and return **
[0203] -Serialize the result packet and return it.
[0204] 3. **Record the log and return the result**
[0205] -Record the log and return the result
[0206] The final generated content in plantuml format is rendered as follows Figure 9 shown.
[0207] As previously described, the code interpretation assistant system of the embodiment of the present invention further includes a knowledge question and answer subsystem, which is mainly used to perform knowledge questions and answers based on the knowledge contained in the code assets.
[0208] Before conducting knowledge quizzes, you need to build a code knowledge quiz knowledge base. The process of building a code knowledge quiz knowledge base is as follows: Figure 10 As shown, first import the Java project and parse the project structure to extract flowtran information and java class information.
[0209] The XML file containing the flowtran information is processed. The specific processing logic is as follows: first, recursively traverse the specified project directory, filter and process the flowtrans.xml file, and then delete designated unused XML nodes such as interface and mapping. The LLM large language model is then used to interpret the processed XML content and obtain the LLM interpretation results (the explanatory text after the LLM has processed the flowtrans.xml file). The interpretation results are manually reviewed and corrected. After the above process is completed, the data cleaning and conversion process begins. The specific steps are as follows: first, extract basic transaction information such as transaction ID, transaction name, and package name, then extract transaction description information and process service information. Finally, LLM is used to generate transaction usage scenarios. Based on the business scenario corpus, manual review and error correction are performed. After cleaning, two CSV files are generated: the first file is defined as "Project Name - Transaction Summary Information.csv", which contains full_path metadata and transaction summary information; the second file is defined as "Project Name - Transaction Details.csv", which, in addition to the full_path metadata and transaction summary information, also contains detailed transaction descriptions generated by LLM and manually reviewed.
[0210] The specific process of LLM generating transaction usage scenarios is as follows:
[0211] S1, input the transaction ID, transaction name, and process description to LLM.
[0212] S2, prompt design, inputs instructions for LLM. Based on the following transaction process description, a real business use case is generated for it, including customer goals, operation procedures, involved services, as well as possible problems and supplementary information in the transaction scenario.
[0213] S3 receives the transaction scenario results generated by the LLM. In the banking system, a customer initiates a transfer through the mobile banking app. The system first verifies the customer's identity (by calling UserValidationService). Once verified, the funds are transferred (by calling MoneyTransferService). This scenario is suitable for single, small transfers, and the system requires two-factor authentication before the transfer.
[0214] To process Java class information, the system first recursively traverses the specified project directory, filters out useless Java class files, and extracts the package name, class name, class declaration type, and comments of the Java file. The system then enters the data cleaning and conversion process, allowing LLM to parse the class's field information, method information, comments, and code implementation for manual review and correction. The parsed results are fragmented and stored in multiple chunks. If a class method is relatively large, it will correspond to multiple chunks, and their order must be strictly maintained. Therefore, the file's full_path metadata and serial number are used as identifiers to ensure the continuity of each class method chunk. Finally, a CSV file containing all key information about the Java class is generated.
[0215] Detailed description of LLM parsing java class information processing:
[0216] S1, parses the Java class field information and extracts the fields (attributes) declared in the class, including the field name, type, access modifier (`private`, `public`, etc.), whether it is static (`static`), and whether it is a constant (`final`).
[0217] S2, parse method information, including: method name, return value type, parameter list (including parameter name and type), access modifier, whether it is a static method, and method comments.
[0218] S3, parses Java class annotation information, including: class-level, field-level, and method-level annotations.
[0219] S4, sharding is performed based on the parsing results. Sharding rules: a. Each method is processed separately, and if the method is too large, it is further sharded; b. Each chunk is limited to a reasonable number of tokens, such as 2000 tokens; c. If the method is too long, it is segmented by logical blocks (such as if-else, for loop, try-catch block); d. Multiple chunks of the same method must be clearly marked with sequence numbers (such as: chunk 1 / 3, 2 / 3, 3 / 3) to ensure the correct restoration order. The following is a sharding example:
[0220]
[0221] After obtaining all the above CSV files, import the CSV files into the Code Knowledge Q&A Knowledge Base. Note that it is also possible to incrementally update the Java code here by locating the chunk fragment based on the full_path metadata.
[0222] Figure 11The following figure shows a scenario in which the knowledge question answering subsystem of an embodiment of the present invention performs question answering on specific business knowledge. Figure 11 As shown, for the code knowledge question and answer scenario to function properly, three knowledge bases are required. Knowledge base one is the code knowledge question and answer knowledge base described in the previous section. It divides all code into method-based chunks and includes metadata (such as file paths, package names, class names, etc.); knowledge base two is a Chinese-English comparison table knowledge base, which facilitates understanding the mapping from Chinese descriptions to English words, as function and variable names in the code are more likely to be composed of English words; knowledge base three is a synonym comparison table knowledge base, which is used to compare with the data dictionary to unify semantics. For example, "batch processing" is a term formally defined in the data dictionary to align data standards, but this term can also be called "running batches," "batch processing," "batch transactions," etc.
[0223] The entire process is as follows: When a query is initiated, such as "Please explain the overall process of batch running," the synonym table in the RAG knowledge base is first searched to optimize the relevant terms. After RAG is retrieved, the optimized question list 1 is obtained. This question list 1 optimizes the initial question "Please explain the overall process of batch running" to "Please explain the overall process of batch processing" described in standard terms.
[0224] Then, the Chinese-English comparison table in the Chinese-English RAG knowledge base is retrieved to obtain the matching English description. After the RAG is recalled, the optimized question list 2 is obtained. This question list 2 converts the previous step "Help me explain the overall process of batch processing" into "Please explain the biz logic for batch processing." Now, after obtaining the original question + optimized question list 1 + optimized question list 2, a common large model (such as a general-purpose LLM (large language model) for reasoning optimization) is called to extract the package path and method call relationship related to the question. Then, the code knowledge question and answer RAG knowledge base is retrieved and recalled. The original question, optimized question, and recalled code snippets are submitted to the code interpretation large model to generate the final question and answer results. The complete process of generating question and answer results by the code interpretation large model is as follows:
[0225] S1, input preparation stage, the final input to the code interpretation model includes:
[0226] a. Original question (e.g., "Explain the overall process of batch processing to me");
[0227] b. Optimized Question List 1 (more standard expression optimized for Chinese, such as: "Please explain the core process of batch processing in detail");
[0228] c. Optimized Question List 2 (translated into optimized English using a Chinese-English comparison table, such as "Please explain the biz logic for batch processing.");
[0229] d. Retrieved code snippets (main method, calling method, code logic block);
[0230] Data composition format:
[0231]
[0232] S2, calling the code to interpret the large model prompt is designed as follows:
[0233]
[0234]
[0235] S3, after receiving the input, the large model will perform the following reasoning and generation process according to the instructions:
[0236] a. Understand the question intent and comprehensively determine what the user really wants to know based on the original question and the optimized question list;
[0237] b. Understand the code context and identify the function, calling relationship, input and output of the method.
[0238] c. Structure your answer. Start with an overall summary (e.g., the entire batch processing process is divided into several steps), then describe the logic of each step in detail, interspersed with code references as appropriate (e.g., "In the `batchProcess()` method, `validateInput()` is first called to validate the parameters...");
[0239] S4, output the final question and answer, example:
[0240]
[0241] In an optional embodiment of the present invention, for more complex business knowledge questions and answers, an "animation"-based presentation method can be used. When asking about hot account mechanisms and solutions, this knowledge is more complex, so an "animation"-based presentation method is used. Animation screenshot Figures 12a to 12d The underlying animation is based on Manim, an open-source animation engine used to create explanatory animated videos. Therefore, by having the code interpret the large model, the question-and-answer results are converted into Manim's Python scripts and then rendered as animations.
[0242] The code interpretation model of the embodiment of the present invention is an artificial intelligence model for code interpretation. The functional structure and description of each layer of the code interpretation structure are as follows:
[0243] 1. Input Data
[0244] 1. Source code (Java source code file)
[0245] Refers to the standard `.java` files used within the system to implement business logic, service functions, exception handling, and data interaction. Its characteristics are: a. Functional objects, including definitions of classes (Class), interfaces (Interface), methods (Method), attributes (Field), etc.; b. Structural features, with a clear package structure (Package), comments (Comment), method call chain, and exception capture logic; c. Target content, extracting core business processes such as transaction entry (such as Controller class), service orchestration (Service layer), data operations (DAO layer), and exception handling (Try-Catch structure).
[0246] 2. XML configuration file (metadata model)
[0247] Refers to a specific type of `.xml` file* that defines metadata such as transaction processes, service functions, table structures, error codes, etc. in a project. It is not a general XML file and mainly includes the following subcategories. Each type of XML file has clear business semantics:
[0248] a. FlowTrans definition file `flowtrans.xml` defines the online transaction interface and describes the correspondence between transaction codes and processing logic.
[0249] b. The service function definition file `apsServiceType.xml` defines a reusable service function module.
[0250] c. The service implementation definition file `apsServiceImpl.xml` defines the call implementation details (class method binding) of a specific service.
[0251] d. The table structure definition file `tables.xml` defines the database table structure, fields, primary keys and other information.
[0252] e. The error code definition file `error.xml` defines the mapping between the error code returned by the system and the error information.
[0253] f. The data dictionary definition files `d_schema.xml`, `u_schema.xml`, and `e_schema.xml` define data items, basic types, and enumeration types.
[0254] g. The composite type definition file `c_schema.xml` defines a composite data type consisting of multiple data items.
[0255] h. The batch transaction definition file `batch_tran.xml` defines the business process for batch processing transactions.
[0256] i. File Batch Transaction Definition The `file_batch_tran.xml` file defines batch transaction processing based on file input and output.
[0257] j. The custom SQL definition file `nsql.xml` defines complex SQL statements and their mapping methods.
[0258] 2. Output Data
[0259] 1.PlantUML format text: standard PlantUML flowchart syntax description.
[0260] 2. draw.io structured data: data format that complies with the draw.io XML or JSON standards.
[0261] 3.manim animation script: Python code file, using the manim library for animation generation.
[0262] 3. Internal structure
[0263] 1. Preprocessing layer: Standardize semi-structured text, extract core events, actions, and object relationships, and use regex-based keyword extraction and structure decomposition. The following is a preprocessing Python script:
[0264]
[0265]
[0266]
[0267] 2. Semantic construction layer: The pre-processed structured text content is called by the LLM model to generate the interpretation result. The following is the prompt word structured by calling the LLM model:
[0268]
[0269]
[0270] 3. Format mapping layer: According to the output format syntax, it is mapped to the syntax structure required by PlantUML, draw.io, and manim. The following is the prompt design template for PlantUML syntax:
[0271]
[0272] 4. Output generation layer: finally serialize to the corresponding file format, perform syntax splicing and file encapsulation according to the target format, such as generating puml files, drawio xml, and manim scripts.
[0273] Code interpretation calling large language model prompt:
[0274]
[0275]
[0276] Code interpretation: Example of the output result of calling the large language model:
[0277]
[0278] The code interpretation assistant system in this embodiment of the present invention can rapidly build a code interpretation assistant based on a large code interpretation model, fully tapping into the potential of source code as the most authentic, accurate, and timely source of knowledge. This is particularly true in core banking scenarios, where understanding the corresponding business process details based on a transaction code and presenting them as flowcharts is a rigid requirement. The code interpretation assistant can help business personnel quickly achieve this goal.
[0279] In practical applications, both business personnel and technical personnel have a desire to understand the specific business knowledge contained in the code. In the past, they could only refer to the corresponding business and technical documents. However, due to lack of timely updates, these documents may contain inaccurate knowledge or not be the latest version, which can cause deviations in understanding and even generate technical debt. The code interpretation assistant system of the embodiment of the present invention relies on a code knowledge question and answer knowledge base, which can recall the most accurate knowledge in the source code and organize the corresponding most accurate answers through a large code interpretation model. This effectively solves the above pain points.
[0280] A key difficulty in code knowledge extraction projects lies in the extraction of method call relationships. Because the interpretation of business processes must fall on the method call level, the efficiency of method call extraction directly determines the efficiency of converting source code assets into text assets. The embodiment of the present invention deeply studies the laws of programming languages (for example, most core codes are written in Java, so this is mainly combined with the Java programming language) and creatively uses a regular expression-based method call extraction method. This method is 60-70 times faster than the method of extracting method call relationships using a large model or the method of extracting method call relationships based on the Ast abstract syntax tree.
[0281] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A code interpretation assistant system, characterized in that: The code interpretation assistant system includes a static code interpretation subsystem and a knowledge question answering subsystem; The code interpretation system is used to interpret the transaction code of the core system through the code interpretation engine to obtain the transaction rules and business process diagram in the language matching the transaction code; The knowledge question and answer subsystem is used to build a code knowledge question and answer database to perform code knowledge question and answer in combination with the knowledge code question and answer database.
2. The system according to claim 1, wherein: The code interpretation subsystem implements static interpretation of code assets through a code interpretation engine, which includes a static code asset parsing module and an atomic capability building module; The static code asset parsing module is used to obtain transaction-related static code assets based on the transaction code of the core system; The static code asset parsing module includes metadata model parsing and Java file parsing. The metadata model parsing includes obtaining multiple definition contents from the metadata model based on the transaction code; the Java file parsing includes searching for Java source files related to the transaction based on the definition contents. The atomic capability analysis module is used to perform single-method business rule generation, service business rule generation, service flow chart generation, service flow chart generation, transaction flow chart generation, transaction / service related error code identification, transaction / service related business table lineage relationship identification, and transaction process orchestration identification.
3. The system according to claim 2, characterized in that The process of static interpretation of code assets by the code interpretation engine includes: Parse the static code assets obtained based on the transaction code and obtain the corresponding Java files by processing the metamodel; Analyze the Java file based on the abstract syntax tree (AST) and establish associations between multiple files; Remove Java files that do not need to be interpreted based on the filtered file configuration, and finally interpret the remaining Java files after filtering; The metamodel data will be stored in MongoDB, and the code interpretation results will be placed in the code interpretation RAG knowledge base for retrieval.
4. The system according to claim 3, characterized in that The process of interpreting the transaction code given by the user by the code interpretation engine includes: Receive the transaction code entered by the user, search the corresponding metadata model in MongoDB based on the transaction code, perform RAG search in the code knowledge base, and obtain matching information; Drill down into the calling method to obtain a list of other methods called by the service method; Determine whether the transaction is multi-node. If it is multi-node, continue to loop the process of retrieving the metadata model based on the transaction code and obtaining the list of other methods called by the service method. Otherwise, assemble the output; The main method and other methods obtained by drilling down are input into the context of the code interpretation model, and the code interpretation model is used to form an interpretation text.
5. The system according to claim 4, characterized in that A list of other methods to get service method calls include: Get the import list and package name from the specified Java source file, and then get the current class name and the parent class name inherited by the current class name; Locate and obtain the source code of the method body of the specified method based on the abstract syntax tree and convert it into a string; Preprocess the content of the method body; preprocessing includes removing the specified format content information; Get the input parameter information of the method body and use regular expressions to match various method calls, including calls to static methods of tool classes, calls to static methods of the current class, chain calls, super() method calls, this() method calls, parameter object method calls, method calls to obtain a bean object using getBeans(beanName), and method calls on other objects using the dot (.) operator; The call relationship is processed uniformly. The specific processing methods are: checking the fully qualified class name, processing static methods, processing object methods, and processing automatically injected objects; Parse the class name and method name, construct a call relationship string, remove duplicates and sort, and finally return the result.
6. The system according to claim 1, wherein: The knowledge question and answer subsystem includes a code knowledge question and answer knowledge base, a Chinese-English comparison table knowledge base, and a synonym comparison table knowledge base; the code knowledge question and answer base divides the code into chunks and records its meta-information, and the meta-information includes file path, package name, and class name; the English comparison table knowledge base is used to store the mapping from Chinese to English; the synonym comparison table knowledge base is used to compare data dictionaries to unify semantics.
7. The system according to claim 1, wherein: The workflow of the knowledge question answering subsystem includes: Receiving original query information, searching a synonym table in a synonym table knowledge base, optimizing terms of the original query information, and obtaining first optimized query information; Searching a Chinese-English comparison table knowledge base to obtain an English description matching the first optimized query information, thereby obtaining second optimized query information; Extracting package paths and method call relationships related to the second question list, and retrieving and recalling target code snippets related to the first optimized query information based on the knowledge question and answer knowledge base; The original query information, the first optimized query information and / or the second optimized query information, and the target code snippet are input into a code interpretation model, and a question-answer structure is generated using the code interpretation model.
8. The system according to claim 6, wherein: The code knowledge question and answer database is constructed in the following manner: Import Java projects and parse the project structure to extract flowtran information and java class information; After information processing is performed on the flowtran information and the java class information, a csv file is obtained, and the csv file is imported into a code knowledge question and answer library.
9. The system according to claim 8, characterized in that The processing of the flowtran information includes: Recursively traverse the specified project directory, filter and process the flowtrans.xml file; Delete the specified useless XML nodes, use LLM to interpret the processed XML content and obtain the LLM interpretation results; Enter the data cleaning and conversion process, extract basic transaction information such as transaction ID, transaction name, package name, transaction description information and process service information, and use LLM to generate transaction usage scenarios for review and error correction based on the business scenario corpus; After cleaning, two CSV files are generated. The first CSV file is defined as "Project Name - Transaction Summary Information.csv" and contains full_path metadata and transaction summary information. The second CSV file is defined as "Project Name - Transaction Details.csv". In addition to the full_path metadata and transaction summary information, it also contains detailed transaction descriptions.
10. The system according to claim 8, wherein: The processing process of the Java class information includes: Recursively traverse the specified project directory, filter out useless Java class files, and extract the package name, class name, class declaration type, and comments of the Java file; Enter the data cleaning and conversion process, let LLM parse the class field information, method information, comments and code implementation and conduct review and correction; The parsing results are fragmented and stored in multiple chunks. If a class method is relatively large, it will correspond to multiple chunks and their order must be maintained. Therefore, the file's full_path metadata and sequence number are used as identifiers to ensure the continuity of each class method chunk. Finally, a CSV file containing all the key information of the Java class is generated.