Code inference method and device based on semantic scene, equipment, medium and product
By constructing abstract syntax trees and symbol tables, conducting deep semantic understanding and adaptive rule inference, and combining global context analysis, the problems of insufficient semantic understanding, poor flexibility, and limited inference capabilities in complex scenarios are solved in the existing technology, and better code inference and generation are achieved in complex business scenarios, improving the degree of automation and adaptability.
Patent Information
- Application Number
- CN202510113731.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-10
AI Technical Summary
Existing code inference technologies lack understanding of semantic scenarios and cannot deeply understand the business logic or functional requirements behind the code, resulting in the generated code lacking correct inference of specific semantics. They rely on static rules or data to flexibly respond to dynamically changing scenarios.
By constructing an abstract syntax tree and symbol table, deep semantic understanding and adaptive rule inference are carried out, and code inference and generation are realized in combination with global context analysis.
Better code inference and generation in complex business scenarios, improve the degree of automation and adaptability, and solve the problems of insufficient understanding of semantics, poor flexibility, and limited inference capabilities in complex scenarios in the prior art.
Smart Images

Figure CN120122950A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to a code inference method, device, equipment, medium, and product based on semantic scenarios. Background Art
[0002] During the system development and maintenance process, code inference and automatic generation technologies have become key means to improve development efficiency and reduce human errors. Traditional code inference technologies mainly include:
[0003] 1. Rule matching and templated generation: Rule matching technology generates corresponding code by combining predefined rules or patterns with the input context. For example, inferring the function body based on the function signature, or automatically generating a constructor based on the class definition. Such methods rely on rules or templates defined by developers and can handle some simple and repetitive code generation requirements.
[0004] 2. Code inference based on static analysis: Static analysis is a technology for analyzing source code without running the program. By analyzing the structure of the code, the type system, and the dependency relationships of variables, static analysis tools can make certain inferences about the code, helping developers discover potential errors or improve code quality. Such tools sometimes also provide simple code completion functions, such as automatic inference of variables or suggestions for method calls.
[0005] However, the existing code inference technologies above have the following defects.
[0006] 1. Lack of understanding of semantic scenarios: Unable to deeply understand the business logic or functional requirements behind the code, resulting in the generated code often lacking correct inferences about specific semantics. They focus more on the syntax and structure of the code rather than the actual meaning behind it.
[0007] The reason for this defect is syntax priority and the lack of a semantic model. Traditional technologies such as rule matching and static analysis mainly rely on the syntax structure rather than in-depth semantic analysis. They focus on analyzing the code form rather than the function of the code in a specific business scenario. In traditional tools, there is a lack of a connection mechanism between business semantics and code behavior, and the business context is not effectively combined with code generation, resulting in the generated code being difficult to match the real requirements of specific scenarios.
[0008] 2. Dependence on static rules or data: Templated generation and static analysis tools highly depend on predefined rules or code structures and cannot flexibly handle dynamically changing scenarios.
[0009] The reason for this defect is the limitation of predefined rules and the unpredictability of the dynamic environment. Predefined rules are fixed and static, and they can only handle patterns or syntactic features that are known in advance. If the scenario changes or new programming patterns emerge, it is difficult for the rules to adapt. During the code generation process, the program may face a complex operating environment. Static rules are difficult to capture this complexity and change, resulting in inaccurate inference results.
[0010] Therefore, it has become an urgent technical problem for those skilled in the art to provide an improved code inference technology based on semantic scenarios based on the deficiencies in the prior art. Summary of the Invention
[0011] Problems to be Solved by the Invention
[0012] The object of the present invention is to overcome the defects of the prior art and provide an improved code inference method, device, equipment and medium based on semantic scenarios. According to the improved code inference method, device, equipment and medium provided by the present invention, problems such as insufficient semantic understanding, poor flexibility, and limited inference ability in complex scenarios during the code inference and generation process of the prior art are solved, and better code inference and generation in complex business scenarios are realized, improving the degree of automation and adaptability.
[0013] Methods for Solving the Problems
[0014] The first aspect of the present invention relates to a code inference method based on semantic scenarios, including the following steps:
[0015] A construction step of constructing an abstract syntax tree and a symbol table based on the semantic scenario;
[0016] A calculation step of traversing the abstract syntax tree, calculating the parameters at the cursor position, and obtaining nodes with incomplete semantics;
[0017] An inference step of performing code inference based on the nodes with incomplete semantics and the symbol table to obtain the result of code completion;
[0018] A completion step of performing code completion according to the result of code completion.
[0019] Preferably, between the calculation step and the inference step, the ambiguity of the abstract syntax tree is also judged.
[0020] Preferably, the parameter at the cursor position is a syntax tree node.
[0021] Preferably, the parameter at the cursor position is an array of parent nodes.
[0022] Preferably, when the node with incomplete semantics is a member access node, the scope of the member access node is obtained through the symbol table.
[0023] Preferably, when the node with incomplete semantics is an identifier node, first judge according to the context of the node. If the exact content cannot be judged according to the context, then judge according to the symbol table.
[0024] The second aspect of the present invention relates to a code inference device, including:
[0025] A construction module for constructing an abstract syntax tree and a symbol table based on a semantic scenario;
[0026] A calculation module for traversing the abstract syntax tree, calculating the parameters at the cursor position, and obtaining nodes with incomplete semantics;
[0027] An inference module for performing code inference based on the nodes with incomplete semantics and the symbol table to obtain the result of code completion;
[0028] A completion module for performing code completion according to the result of code completion.
[0029] The third aspect of the present invention relates to a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the code inference method in the first aspect are implemented.
[0030] The fourth aspect of the present invention relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the code inference method in the first aspect are implemented.
[0031] The fifth aspect of the present invention relates to a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the code inference method in the first aspect are implemented.
[0032] Effects of the Invention
[0033] According to the improved code inference method provided by the present invention, in-depth semantic understanding, adaptive rule inference, and global context analysis are performed, solving the problems of insufficient semantic understanding, poor flexibility, and limited inference ability in complex scenarios in the prior art during code inference and generation. It realizes better code inference and generation in complex business scenarios, and improves the degree of automation and adaptability. Description of the Drawings
[0034] Figure 1 It is a flowchart of the code inference method according to the first embodiment of the present invention.
[0035] Figure 2 It is a structural diagram of the computer device according to the third embodiment of the present invention.
[0036] Figure 3 Schematic diagram of the code inference method in Figure 1 .
[0037] Figure 4 Schematic diagram for ambiguity determination of the code inference method in Figure 1 .
[0038] Figure 5 Schematic diagram of the inference steps of the code inference method in Figure 1 . Detailed implementation mode
[0039] Hereinafter, the code inference method involved in the present invention will be described in detail first.
[0040] Figure 1 Flowchart of the code inference method according to the first embodiment of the present invention. As Figure 1 shown, the specific process of the code inference method is as follows. First, a construction step (step S100) is performed to construct an abstract syntax tree and a symbol table based on the semantic scenario. Then, a calculation step (step S101) is performed to traverse the abstract syntax tree, calculate the parameters at the cursor position, and obtain nodes with incomplete semantics. Then, an inference step (step S102) is performed to perform code inference based on the nodes with incomplete semantics and the symbol table to obtain the result of code completion. Finally, a completion step (step S103) is performed to complete the code according to the result of code completion.
[0041] Description of step S100. The abstract syntax tree (AST) is an abstract representation of the source code structure, which shows the syntax structure of a programming language in a tree-like form, and each node on the tree represents a structure in the source code. In the method of intelligent code inference based on the semantic scenario, the cursor position information, the specific content of the text to be completed, and the accurately constructed symbol table are indispensable key input elements. Therefore, first, an elaborate and accurate abstract syntax tree and a symbol table are constructed by deeply parsing the code.
[0042] Description of step S101. Specifically, as Figure 3As shown, during the traversal of the Abstract Syntax Tree (AST), each node is visited one by one to calculate and determine the specific location of the cursor. This step is directly related to the accuracy of subsequent semantic analysis. The parameter of the cursor position is preferably a ParseTreeNode (syntax tree node) or a ParentNode[] (array of parent nodes). The syntax tree node reveals the code element pointed to by the current cursor, and its associated array of parent nodes provides progressive context information for inference. These context information together constitute the basis of semantic analysis, enabling code intelligent inference to accurately understand the code intention and context environment at the cursor position based on these detailed clues. When the text where the cursor is located is incomplete and the semantically incomplete nodes are IdentifierNode (identifier node) or MemberAccessNode (member access node), these two types of nodes are processed differently during the code completion process.
[0043] The description of step S102 is as follows. As Figure 5As shown, specifically, preferably, when the node with incomplete semantics is a MemberAccessNode, the scope of the member access node is obtained through the symbol table. The Scope of the MemberAccessNode is obtained through the constructed symbol table to achieve code completion. For example, when the input is similar to object., the cursor stops after the dot. At this time, the type of object is found through the symbol table, so as to provide intelligent prompts for member methods and properties related to this type, such as object.ToString(), object.ToArray(), etc. In this process, the data structure in the symbol table records each object, class, interface and their corresponding members, ensuring that the system can provide correct candidate completion items based on the object type. Preferably, when the node with incomplete semantics is an IdentifierNode, first judge according to the context of the node. If no exact content can be judged according to the context, then judge according to the symbol table. The possible content to be input after the cursor is inferred through the context of the code. For example, in a LINQ statement in C# code, if the cursor is in the environment of a ComputeNode (representing the compute operation scenario in C#), the possible content to be input next is inferred to be keywords or functions related to LINQ operations. For example, for var result = myList., at this time the cursor stops in front of function keywords such as.Any(),.All(), etc. By judging the context, it is possible to want to use functions related to the ComputeNode, and the system will provide intelligent completion prompts for methods such as ANY and ALL. When the context cannot draw a clear inference, the system will perform a prefix optimal match based on the symbol table. The symbol table stores all currently available identifiers, variables, methods, etc. When the cursor is at an IdentifierNode (for example, the cursor is in the middle or at the beginning of a variable name), the system will search for matching identifiers in the symbol table according to the input prefix, so as to provide completion suggestions. For example, if the user enters myOb when inputting a variable but has not completed the entire variable name, at this time the cursor is in the IdentifierNode. The system will search for all variables with myOb as the prefix in the symbol table and provide candidate items such as myObject and myObserver. Finally, the system accurately provides completion options by matching the prefix with the variable names in the symbol table.
[0044] Step S103 is described. Specifically, according to the suggested results provided by the code inference method, after carefully analyzing and understanding the code context, precise and effective code completion is performed, so as to ensure the integrity of the code, the correctness of the logic, and the implementation of the function, thereby improving the running efficiency and reliability of the code.
[0045] Preferably, between the calculation step and the inference step, the ambiguity of the abstract syntax tree is also judged. As Figure 4 shown, during the code completion process, there are problems with constructing the AST based on incorrect rules and the resulting ambiguity, especially when dealing with complex syntax structures. Taking a[0] as an example, the position of [0] may need to be inferred from the context whether it is an array or a subscript index. In the grammar definition file (grammar), if both the array ([0] itself represents an array) and the arrayIndex ([0] is the subscript of the array a) rules are matched, the parsed result may tend to the incorrect option, such as misjudging [0] as an array instead of a subscript index, which may lead to incorrect inferred scopes or completion hints. To address this ambiguity problem, a placeholder is inserted at the cursor position. By inserting a placeholder at the cursor position in the code, a more complete syntax structure can be provided to the parser, improving the context semantics of the code snippet and helping it to more accurately parse the result expected by the user. By combining the construction of the symbol table and the semantic judgment of the context, different types of nodes can be effectively processed, and accurate code completion suggestions can be dynamically generated, significantly improving the intelligence of code writing.
[0046] It can be seen that the code inference method according to the first embodiment of the present invention solves the problems of insufficient semantic understanding, poor flexibility, and limited inference ability in complex scenarios in the prior art during code inference and generation, realizes better code inference and generation in complex business scenarios, improves the degree of automation and adaptability, and performs better in code inference and generation in complex business scenarios. The prior art mainly uses syntax structure analysis for static checking and format optimization, lacking a deep understanding of the business semantics of the code. Most of the prior art can only process the local structure of the code and cannot perform inferences based on the global business context. Through AST category inference combined with semantic scenarios, this application can deeply understand the business logic behind the code. By classifying AST nodes and associating them with business semantic scenarios, a code generation scheme that meets the actual needs can be inferred. Compared with traditional syntax structure analysis, through the association between AST categories and business semantics, this application can deeply understand the business logic of the code and generate code that better meets the actual needs, rather than being limited to syntax structure analysis. In addition, the prior art usually uses fixed rules and cannot be adjusted in real time according to changes in business requirements. When facing dynamic business scenarios, the rules need to be manually updated, lacking flexibility. Through dynamic adaptive inference rules, this application can automatically adjust the inference strategy according to code context and changes in business requirements, reducing manual intervention, achieving efficient dynamic adaptation, being superior to the prior art in terms of dynamic scenario adaptability, having an adaptive ability, being able to flexibly respond to changes in complex business scenarios, having an adaptive ability, and being able to automatically adjust inference rules according to changes in business scenarios, overcoming the limitations of relying on fixed rules in traditional technologies. In addition, the prior art highly relies on manually defined rules and templates, with high update and maintenance costs and limited rule applicability. Through automated inference rule generation and optimization, this application greatly reduces the dependence on the manual rule library. The inference rule library is automatically generated and optimized according to the actual scenario to adapt to diverse business requirements. Compared with the prior art, through automated rule generation, this application significantly reduces the maintenance cost, and through automated generation and optimization of inference rules, significantly reduces the need for manual maintenance and manual rule updates, improving the flexibility and efficiency of the system. In addition, although the prior art can parse the code structure, the inferred code fragments are often inconsistent in business logic and semantics, with insufficient accuracy and relying on manual adjustment. Through deep semantic analysis and semantic verification, this application ensures that the inferred code is semantically consistent and highly accurate, reducing manual correction work, being significantly superior to traditional methods in terms of code accuracy and semantic consistency, and generating code that better meets the actual business needs.
[0047] The code inference device according to the second embodiment of the present invention includes: a construction module for constructing an abstract syntax tree and a symbol table based on a semantic scenario; a calculation module for traversing the abstract syntax tree, calculating the parameters at the cursor position, and obtaining nodes with incomplete semantics; an inference module for performing code inference based on the nodes with incomplete semantics and the symbol table to obtain the result of code completion; and a completion module for performing code completion according to the result of code completion. The code inference device corresponds to the code inference method of the first embodiment. Therefore, various deformation methods in the first embodiment are also applicable to the second embodiment and will not be elaborated here.
[0048] As described above, the code inference device according to the second embodiment of the present invention solves the problems in the prior art such as insufficient semantic understanding, poor flexibility, and limited inference ability in complex scenarios during the process of code inference and generation, realizes better code inference and generation in complex business scenarios, and improves the degree of automation and adaptability.
[0049] The third embodiment of the present invention provides a computer device, and the internal structure diagram of the computer device can be as Figure 2 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to connect to external devices for data interaction with external devices. When the computer program is executed by the processor, it implements the code inference method according to the first embodiment of the present invention.
[0050] Those skilled in the art can understand that Figure 2 the structure shown in
[0051] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0052] The fourth embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the processor executes the computer program, it implements the code inference method according to the first embodiment of the present invention.
[0053] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0054] Industrial applicability
[0055] According to the code inference method, device, computer device, storage medium, and computer program product involved in the present invention, the problems in the prior art such as insufficient semantic understanding, poor flexibility, and limited inference ability in complex scenarios during the code inference and generation process are solved. It realizes better code inference and generation in complex business scenarios, and improves the degree of automation and adaptability.
[0056] Although the present invention has been illustrated and described by referring to some preferred embodiments of the present invention, those of ordinary skill in the art should understand that various changes can be made to it in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A code inference method based on semantic scenarios, characterized in that: The following steps are involved: The construction step builds the abstract syntax tree and symbol table based on the semantic scenario; Calculation step, traverse the abstract syntax tree, calculate the parameters of the cursor position, and obtain the semantically incomplete nodes; Inference step: code inference is performed based on semantically incomplete nodes and symbol tables to obtain code completion results; Completion step: Complete the code according to the result of code completion.
2. The code inference method according to claim 1, characterized in that: Between the calculation step and the inference step, the ambiguity of the abstract syntax tree is also judged.
3. The code inference method according to claim 1, characterized in that: The argument at the cursor position is a syntax tree node.
4. The code inference method according to claim 1, characterized in that: The parameter for the cursor position is an array of parent nodes.
5. The code inference method according to claim 1, characterized in that: When the semantically incomplete node is a member access node, the scope of the member access node is obtained through the symbol table.
6. The code inference method according to claim 1, characterized in that: When the semantically incomplete node is an identifier node, it is first judged based on the context of the node. If the exact content cannot be judged based on the context, it is then judged based on the symbol table.
7. A code inference device based on semantic scenarios, characterized in that: include: Building module, used to build abstract syntax tree and symbol table based on semantic scenarios; A calculation module is used to traverse the abstract syntax tree, calculate the parameters of the cursor position, and obtain semantically incomplete nodes; The inference module is used to perform code inference based on semantically incomplete nodes and symbol tables to obtain code completion results; The completion module is used to complete the code according to the results of code completion.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the code inference method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the code inference method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the code inference method according to any one of claims 1 to 6 are implemented.