An industrial code intelligent perception method based on declaration implementation separation and path stack

CN122594233BActive Publication Date: 2026-09-29GUODIAN NANJING AUTOMATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611079967.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-29
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

基于坐标的补全要求服务端始终维护一棵与该文件完全同步的巨型AST;这不仅占用巨大内存,且任何局部修改都可能触发AST的大面积更新;补全时,从坐标点出发的回溯操作需要在庞大的AST中进行深度搜索,延迟随之增大

Benefits of technology

[0047]1、本发明在项目初始化或更新时,仅提取工业工程文件中的声明代码(包括POU变量声明区、全局变量声明区、用户自定义数据类型定义区),跳过所有POU的逻辑实现代码体,构建全量符号索引,只有对POU中代码进行语义高亮时,才解析POU的逻辑实现代码体,进而将索引构建时间与代码体规模解耦,实现秒级全量符号索引构建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594233B_ABST
    Figure CN122594233B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of code intelligent perception, and discloses an industrial code intelligent perception method based on declaration implementation separation and path stack, which comprises the following steps: parsing declaration codes of engineering files when a project is loaded, constructing and maintaining a full-amount symbol index table based on the declaration codes, and querying a highlighted program organization unit logical code body by using the full-amount symbol index table; when a user edits the program organization unit logical code body, extracting an access path stack according to a cursor interception and operator matching rule, combining the access path stack with the full-amount symbol index table, and executing industrial code intelligent perception processing. When a project is initialized or updated, only the declaration codes in the industrial engineering files are extracted, all logical implementation code bodies are skipped, a full-amount symbol index is constructed, the index construction time is decoupled from the code body scale, and second-level full-amount symbol index construction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent code perception technology, and more specifically, to an industrial intelligent code perception method based on declaration implementation separation and path stack. Background Technology

[0002] Currently, the common technical solution in the field of intelligent code recognition is the Language Server Protocol, which works as follows in mainstream languages ​​such as C / C++, Java, and Python:

[0003] Project Index Building: After the server starts, it scans all source files in the project. For each source file, it runs a complete lexical and syntactic analysis to generate an abstract syntax tree (AST) and extracts symbol definitions (classes, functions, variables, etc.) from it. The ASTs and symbol tables of all files are aggregated into a global data pool.

[0004] Semantic highlighting: After the project index is built, the server distinguishes semantic token types such as local variables, global variables, function calls, and type names based on the symbol table, and encodes them as token type labels to return to the editor, achieving the highlighting effect. Semantic highlighting requires a complete AST and symbol table.

[0005] Code completion: After the project index is built, when the developer enters obj.member and triggers completion, the editor (client) sends the Uniform Resource Identifier (URI) of the current file and the text coordinates (line number, column number) of the cursor to the server. After receiving the request, the server locates the AST of the corresponding file, starts from the node at the cursor coordinates, traverses the syntax tree upwards, parses the type of obj, searches for the candidate member in the member list of the type, and returns the completion result.

[0006] However, current industrial control programming software that follows the IEC 61131-3 standard generally stores all POUs, global variables, data types, tasks, and other information in a single project file. Consequently, the LSP solution faces the following problems when processing industrial code:

[0007] Slow project index construction: LSP's file-level full parsing processing object is each source code file, and each source file is an independent processing unit. However, industrial engineering files combine all POUs into one large file, and the logical code body accounts for a huge proportion. If applied directly, the server will be forced to parse the entire logical code body, and this code body is completely redundant information for the construction of the global symbol table. This unnecessary parsing causes the initialization and index update time to increase catastrophically.

[0008] High latency during nested member code completion: When engineers frequently perform deep member accesses such as ABC in their code, existing completion solutions rely on the client sending text cursor coordinates, requiring the server to load the entire project file and backtrack the syntax tree to deduce the type; industrial project files can be considered as extremely long texts. Coordinate-based completion requires the server to maintain a giant AST that is completely synchronized with the file at all times; this not only consumes a huge amount of memory, but any local modification may also trigger a large-scale update of the AST; during completion, the backtracking operation starting from the coordinate point needs to perform a deep search in the huge AST, which increases the latency accordingly.

[0009] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0010] To address the problems in related technologies, this invention proposes an industrial code intelligent perception method based on declaration implementation separation and path stack, in order to overcome the aforementioned technical problems existing in the existing related technologies.

[0011] Therefore, the specific technical solution adopted by the present invention is as follows:

[0012] In a first aspect, the present invention provides an industrial code intelligent perception method based on declaration-based separation and path stack, the method comprising:

[0013] When the project loads, the declaration code of the project file is parsed, a full symbol index table is built and maintained based on the declaration code, and the logic code body of the program organization unit is queried and highlighted using the full symbol index table;

[0014] When the user edits the program to organize the unit logic code body, the access path stack is extracted according to the cursor truncation and operator matching rules, and combined with the full symbol index table to perform industrial code intelligent perception processing.

[0015] Preferably, the declaration code of the project file is parsed during project loading, a full symbol index table is constructed and maintained based on the declaration code, and the logical code body of the highlighted program organization unit is queried using the full symbol index table, including:

[0016] When a user opens a project file, the declaration code of all sub-projects in the project file is extracted and sent to the server. After receiving the declaration code, the server performs lightweight parsing processing to build a parse tree.

[0017] The visitor pattern is used to traverse the parse tree to extract all symbol definitions in the declaration code, a full symbol index table is built, and the full symbol index table is optimized based on user editing behavior and symbol index update rules.

[0018] When a user opens a program unit, the client extracts the corresponding logic code and sends it to the server.

[0019] The server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, and outputs the highlighted logical code body of the program organization unit.

[0020] Preferably, when a user opens a project file, the declaration code of all sub-projects in the project file is extracted and sent to the server. After receiving the declaration code, the server performs lightweight parsing processing to construct a parse tree, including:

[0021] When a user opens a project file, the client extracts the program organization units, variable declarations, global variable declarations, and user-defined data types contained in each sub-project within the project file as declaration code, and sends the declaration code to the server.

[0022] The server defines lexical and grammatical rules based on the industrial code syntax rules, and constructs a lexical analyzer and a grammatical analyzer in conjunction with an analyzer generation tool;

[0023] The declaration code is input into the lexical analyzer to perform character-by-character scanning of the string. According to the lexical rules, the declaration code is divided into a stream of the smallest semantic units, and irrelevant content contained in the declaration code is discarded.

[0024] The minimum semantic unit stream is input into the parser, which matches the sequence of minimum semantic unit streams step by step according to the grammar rules, verifies whether the minimum semantic unit streams conform to the defined language structure, and constructs a parse tree based on the verification results.

[0025] Preferably, the visitor pattern is used to traverse the parse tree to extract all symbol definitions in the declaration code, establish a full symbol index table, and optimize the full symbol index table based on user editing behavior and symbol index update rules, including:

[0026] The visitor pattern is adopted. The visitor interface generated by the analyzer generation tool traverses the parse tree and extracts symbol definition information during the traversal, and fills the predefined symbol index table one by one.

[0027] After the visitor traversal is completed and the predefined symbol index table is filled, a full symbol index table containing all program organization unit names, type names, variable names, and their interrelationships is obtained;

[0028] The client listens for user editing behavior and optimizes the full symbol index table by using timers and full and incremental symbol index update operations when the target editing behavior is detected.

[0029] The target editing behaviors include: renaming, deleting, and modifying program organization unit names and variables within program organization units; renaming, deleting, and modifying global variables; and renaming, deleting, and modifying user-defined data types.

[0030] Preferably, the server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, outputting the highlighted logical code body of the program organization unit, including:

[0031] The server divides the logical code of the program organization unit into the smallest independently identifiable semantic unit according to the lexical rules of the structured text language. The smallest semantic unit contains the text content and the position information of the text content in the logical code.

[0032] The server labels each smallest semantic unit with lexical information including keywords, operators, literals and delimiters, and queries the semantic information corresponding to each smallest semantic unit in the full symbol index table to assign semantic information to the smallest semantic unit.

[0033] Following the execution rule of semantics first and lexical second, based on the lexical and syntactic information of each smallest semantic unit, the client assigns a color to each smallest semantic unit to produce a highlighting effect, and outputs the highlighted program organization unit logic code body.

[0034] Preferably, when the user edits the program organization unit logic code body, the access path stack is extracted according to the cursor truncation and operator matching rules, and combined with the full symbol index table to perform industrial code intelligent perception processing, including:

[0035] When the user's edit program organization unit logic code body is detected, the text and position of the current cursor line are obtained, and the identifier prefix that the user is typing is extracted from the cursor position to the left as the name to be completed;

[0036] Determine if there is an access operator to the left of the name to be completed. If there is no access operator or the termination condition is met, then process it as a normal variable completion and set the path stack to empty.

[0037] If a member access operator exists, the expression is parsed in reverse from the left side of the member access operator to the left. During the reverse parsing process, array subscripts, function calls and spaces are automatically skipped. Multi-level access operators are extracted and added to the path stack in sequence until the termination condition is met and the addition stops.

[0038] Based on the path stack and access operators, an access path stack for code completion is obtained. The access path stack is then used as a key and combined with the full symbol index table to perform intelligent code completion processing.

[0039] Preferably, the process of obtaining an access path stack for code completion based on the path stack and access operators, and using the access path stack as a key in conjunction with the full symbol index table to perform intelligent code completion processing includes:

[0040] Based on the path stack addition result and the name to be completed, generate context information containing the complete path and completion information as an access path stack that meets the expected requirements.

[0041] The client sends the access path stack to the server. The server uses the access path stack and the full symbol index table as a basis to perform top-down type inference and completion processing to filter candidate members.

[0042] Based on the name to be completed and the scoring rules, the matching score of each candidate member name is calculated, and the candidate members are sorted according to the matching score. The selected candidate members are used as the perception result to perform industrial code completion processing.

[0043] Secondly, the present invention also provides an industrial code intelligent perception system based on declaration implementation separation and path stack, the system comprising:

[0044] The declaration implementation separation module is used to parse the declaration code of the project file when the project is loaded, build and maintain a full symbol index table based on the declaration code, and use the full symbol index table to query and highlight the logical code body of the program organization unit;

[0045] The path stack generation module is used to extract the access path stack according to the cursor truncation and operator matching rules when the user edits the program's organizational unit logic code body, and combine it with the full symbol index table to perform industrial code intelligent perception processing.

[0046] The beneficial effects of this invention are as follows:

[0047] 1. During project initialization or update, this invention extracts only the declaration code (including POU variable declaration area, global variable declaration area, and user-defined data type definition area) from the industrial engineering file, skips all POU logic implementation code bodies, and constructs a full symbol index. The POU logic implementation code body is only parsed when semantic highlighting is performed on the code in the POU, thereby decoupling the index construction time from the code body size and achieving second-level full symbol index construction.

[0048] 2. The present invention is based on the code completion mechanism of the access path stack, which enables the client to reverse parse the structured access path stack (ordered member access chain and prefix to be completed) from the cursor position when the user inputs, and send the path stack to the server. The server directly uses the path stack to perform top-down type deduction and member lookup in the pre-built full symbol index table, without loading any source code files or backtracking syntax tree, completely abandoning the coordinate backtracking mode of the existing LSP.

[0049] 3. When extracting the path stack, the client proposed in this invention can automatically skip interfering symbols such as array subscripts, function call parentheses, and spaces, correctly handle nested expressions and various member access operators, and linearize the entire access expression before the cursor into a path stack. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart of an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention;

[0052] Figure 2 This is a flowchart of the construction, updating and semantic highlighting of the full symbol index in an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention.

[0053] Figure 3 This is a flowchart of the access path stack extraction process in an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention.

[0054] Figure 4 This is a flowchart of code completion in an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention;

[0055] Figure 5 This is a code effect diagram of an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention;

[0056] Figure 6 This is a schematic diagram of background query and evaluation scores in an industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention. Detailed Implementation

[0057] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0058] According to an embodiment of the present invention, an industrial code intelligent perception method based on declaration implementation separation and path stack is provided.

[0059] It should be noted that in this embodiment, the declaration code refers to the code snippets in the IEC 61131-3 standard engineering file used to define data structures, global variables, and POU interfaces, including the VAR declaration area, global variable area, TYPE type definition, etc.

[0060] The implementation code refers to the logic code body (ST, LD, FBD, etc., the specific implementation code) in POU.

[0061] A path stack is a data structure that stores an ordered sequence of nested member access expressions, such as [variable name, member 1, member 2...].

[0062] The coordinate backtracking mode is the code completion method used in existing LSP solutions, where the client sends the cursor coordinates, and the server loads the file and traverses the syntax tree to deduce the type.

[0063] IEC 61131-3 is an industrial automation programming language standard developed by the International Electrotechnical Commission, which defines programming languages ​​such as Structured Text (ST), Function Block Diagram (FBD), and Ladder Diagram (LD).

[0064] A POU (Program Organization Unit) is the basic building block of a PLC program, which includes a program, function blocks, and functions.

[0065] Project files are structured files in industrial control software used to centrally manage all POUs, global variables, data types, task configurations, and other content within a project.

[0066] A token is the smallest semantic unit obtained after parsing and splitting the source code.

[0067] The technical background of this embodiment is an industrial control integrated development environment that conforms to the IEC 61131-3 standard. Unlike mainstream programming languages ​​such as C++ and Java, which store each class or function as an independent source file, industrial control projects typically adopt a highly centralized project file organization method:

[0068] A PLC project is managed by a single project file, which centrally stores all the information required for the project, including at least: User-defined data types (UDTs): structures, enumerations, arrays, etc.; a global variable table (GLOBAL VAR); multiple program organization units (POUs), including programs, function blocks, and functions. Each POU is physically separated into a variable declaration area (VAR, VAR_INPUT, VAR_OUTPUT, VAR_TEMP, etc.) and a logic code body (implementation in languages ​​such as ST, LD, and FBD); and task configuration: defining the execution order and cycle of the POUs.

[0069] In addition, industrial programming uses domain-specific programming languages. Compared with mainstream programming languages, industrial code has a simpler syntax, concise expression and statement structures (such as clear rules for assignment, condition, and loop statements), and no complex template programming, lambda expressions, or other advanced features. This centralized storage, file storage structure that separates declaration and implementation, and relatively simple syntax provide an optimization basis for this embodiment.

[0070] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, the industrial code intelligent perception method based on declaration implementation separation and path stack according to an embodiment of the present invention includes:

[0071] Step S1: When the project is loaded, parse the declaration code of the project file, build and maintain a full symbol index table based on the declaration code, and use the full symbol index table to query and highlight the logical code body of the program organization unit.

[0072] In one embodiment, parsing the declaration code of the project file when the project is loaded, constructing and maintaining a full symbol index table based on the declaration code, and querying the highlighted logical code body of the program organization unit using the full symbol index table includes: extracting the declaration code of all sub-projects in the project file when the user opens the project file, and sending the declaration code to the server. After receiving the declaration code, the server performs lightweight parsing processing to construct a parse tree; traversing the parse tree using the visitor pattern to extract all symbol definitions in the declaration code, establishing a full symbol index table, and optimizing the full symbol index table based on user editing behavior and symbol index update rules; extracting the corresponding logical code using the client when the user opens the program organization unit and sending it to the server; the server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, outputting the highlighted logical code body of the program organization unit.

[0073] In one embodiment, when a user opens a project file, the declaration code of all sub-projects in the project file is extracted and sent to the server. After receiving the declaration code, the server performs lightweight parsing processing to build a parse tree, including: when the user opens the project file, the client extracts the program organization units, variable declarations, global variable declarations, and user-defined data types contained in each sub-project of the project file as declaration code, and sends the declaration code to the server; the server defines lexical rules and syntax rules according to the industry code syntax rules, and constructs a lexical analyzer and a syntax analyzer in conjunction with the analyzer generation tool; the declaration code is input into the lexical analyzer to perform character-by-character string scanning processing, and the declaration code is segmented into a stream of minimal semantic units according to the lexical rules, and irrelevant content contained in the declaration code is discarded; the stream of minimal semantic units is input into the syntax analyzer, and the sequence of minimal semantic unit streams is matched step by step according to the syntax rules to verify whether the stream of minimal semantic units conforms to the defined language structure, and a parse tree is built based on the verification results.

[0074] In one embodiment, the process of using the visitor pattern to traverse the parse tree, extract all symbol definitions from the declaration code, build a full symbol index table, and optimize the full symbol index table based on user editing behavior and symbol index update rules includes: using the visitor pattern, the visitor interface generated by the analyzer generation tool traverses the parse tree, extracts symbol definition information during the traversal, and fills the predefined symbol index table line by line; after the visitor traversal is completed and the predefined symbol index table is filled, a full symbol index table containing all program organization unit names, type names, variable names, and their interrelationships is obtained; the client listens for user editing behavior, and when a target editing behavior is detected, optimizes the full symbol index table using a timer and full symbol index update operations and incremental symbol index update operations; wherein the target editing behavior includes: renaming, deleting, or modifying program organization unit names and variables within program organization units; renaming, deleting, or modifying global variables; and renaming, deleting, or modifying user-defined data types.

[0075] In one embodiment, the server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, outputting the highlighted logical code body of the program organization unit. This includes: the server segments the logical code of the program organization unit into independently identifiable smallest semantic units according to the lexical rules of structured text languages, with each smallest semantic unit containing text content and its position information in the logical code; the server annotates each smallest semantic unit with lexical information including keywords, operators, literals, and delimiters, and queries the corresponding semantic information of each smallest semantic unit in the full symbol index table to assign semantic information to the smallest semantic unit; following the execution rule of semantics first and lexical second, based on the lexical and syntactic information of each smallest semantic unit, the client assigns a color to each smallest semantic unit to produce a highlighting effect, and outputs the highlighted logical code body of the program organization unit.

[0076] It needs to be explained that the purpose of step S1 is to achieve second-level full symbol index construction and real-time semantic highlighting based on the declaration-implementation separation on-demand parsing method. Specifically, a unified full symbol index needs to be built during project initialization. Semantic highlighting involves segmenting the logical code body in the POU and implementing semantic-level highlighting for each segmented unit. The core of the declaration-implementation separation on-demand parsing method is to strictly separate the declaration and implementation parts of the industrial code. The declaration part includes POU variables, global variables, and data types, used to build the symbol index; the implementation part includes the logical code body of all POUs, used for highlighting. The implementation process is as follows:

[0077] Building a full symbol index: When a user opens the project, the server needs to build a full symbol index table. Subsequent functions such as semantic highlighting and code completion need to retrieve detailed symbol information from the index table. Symbols include custom data types, global variables, POUs, and variables defined in POUs. Leveraging the feature of separating declarations and implementations within project files, only the declaration code from all sub-projects is extracted. The extracted content includes:

[0078] Each POU and its variable declarations (such as FUNCTION_BLOCK, PROGRAM, and the variable declaration area of ​​FUNCTION: VAR, VAR_INPUT, VAR_OUTPUT, etc.); global variable declarations (VAR_GLOBAL); user-defined data types (structures, enumerations, arrays, etc. in TYPE…END_TYPE); and the logical code body of all POUs (ST, FBD, and other implementation code) are not extracted.

[0079] The client sends all declaration code to the server. After receiving the declaration code, the server initiates a lightweight parsing process for each piece of declaration code: lexical analysis and syntax analysis are only performed on the declaration code. The implementation process is as follows:

[0080] Define lexical and grammatical files: Write industrial code grammar rules that conform to the IEC 61131-3 standard, define lexical rules (to identify keywords, identifiers, operators, literals, etc.) and grammar rules (to define the structure of declared code, such as POU variable declarations, global variable declarations, and custom data type definitions, etc.). ANTLR grammar rules begin with lexical rules, and each lexical rule defines the name of a word and a regular expression to match that word.

[0081] Generate lexical analyzer and parser: Use the ANTLR (parser generation tool) to generate the corresponding lexical analyzer and parser source code based on the .g4 grammar file.

[0082] Lexical analysis: The lexical analyzer scans the input declaration code string character by character and divides it into a token stream (keywords, identifiers, operators, literals, delimiters, etc.) according to lexical rules, which is the smallest semantic unit stream mentioned above, while discarding whitespace characters and comments and other irrelevant content.

[0083] Performing syntax analysis: The parser receives the token stream (the stream of the smallest semantic units) generated by the lexical analyzer, matches the token sequence step by step according to the syntax rules, verifies whether it conforms to the defined language structure, and constructs the corresponding parse tree, also known as the abstract syntax tree (AST).

[0084] Traversing the parse tree using the visitor pattern: The visitor pattern is adopted. By implementing the visitor interface generated by ANTLR, a corresponding visit method is written for each grammar rule node. During the traversal, symbol definition information is extracted and the symbol index table is filled one by one.

[0085] Constructing a full symbol index table: After the visitor traversal is complete, a full symbol index table containing all POU names, type names, variable names, and their relationships is obtained. This index table is the project-level unique symbol source, used for subsequent semantic highlighting and code completion. An abstract syntax tree is constructed, traversed using the visitor pattern, extracting all symbol definitions from the declaration code to build the full symbol index table. This index table is the project-level unique symbol source, containing all POU names, type names, variable names, and their relationships, but does not contain any token or syntax tree information within the code body. This process does not parse any POU logic code body, therefore the parsing speed is extremely fast and the memory usage is minimal.

[0086] Update symbol index: The project-level symbol index will be updated when the user performs the following editing actions:

[0087] Rename, delete, and modify POU names and POU internal variables;

[0088] Rename, delete, or modify global variables;

[0089] Rename, delete, and modify user-defined data types.

[0090] The client continuously monitors the user's editing behavior and records the changed position when the above operation is detected. To avoid frequent index updates, a debouncing algorithm is used, which does not immediately execute the update operation. Debouncing is an algorithm strategy to limit the high-frequency execution of functions. The core idea is: when an event is triggered consecutively, only the last (or the first) triggered operation is executed, and all intermediate triggers are ignored.

[0091] In this embodiment, when a user edits the declaration code (such as renaming POU, adding a global variable, etc.), it may trigger continuous keyboard input, pasting, deleting and other operations. If each operation immediately triggers the symbol index update, it will result in: frequent parsing and index rebuilding, causing a surge in CPU usage; and editor lag, affecting the user experience.

[0092] Therefore, a timer is added in this embodiment. The timer executes every 2 seconds. If the user frequently performs editing operations within 2 seconds, the index update will not be performed. Instead, an index update will be performed uniformly only after the user stops editing for a period of time. A batch update will be performed uniformly after 2 seconds. There are two ways to update the symbol index:

[0093] Full update of symbol index: Full update as follows Figure 2 As shown, the update process reuses the step of building a full symbol index, that is, extracting all the declaration code in the project and rebuilding the symbol index. This implementation method is relatively simple, and the symbol index update speed can basically meet the needs of industrial programming.

[0094] Incremental update of symbol index: Incremental update only performs lexical analysis, syntax analysis, builds syntax tree, traverses syntax tree, extracts symbols, and performs local additions and deletions in the original full symbol index table. This method is more complex to implement, but it can improve the speed of symbol index update.

[0095] The semantic highlighting function requires the use of the project implementation, specifically the logic code body of the POU. When a user opens a POU or edits its logic code, the client extracts all the logic code of that POU and sends it to the server. The server first segments the logic code of the POU into independently identifiable tokens. Token segmentation follows the lexical rules of the Structured Text (ST) language in the IEC 61131-3 standard, cutting a continuous stream of characters into a series of tokens, i.e., the smallest semantic unit mentioned above. This is achieved by writing a lexical file according to the IEC 61131-3 programming rules, using ANTLR to parse the lexical file and generate a lexical parser, and then using the lexical parser... The parser divides the logic code of the POU into token units. Each token contains text content and its position information in the source code. For example, the code snippet `a:=1;` will be divided into four tokens: `{a,:=,1,;}`. During the first parsing process, the server annotates the lexical information of each token, including keywords, operators, literals, delimiters, and other categories. During the second parsing process, the server combines the global symbol index table to query the semantic information corresponding to each token in the index table, such as global variables, local variables, function block instances, function call names, enumeration members, etc., thereby assigning complete semantic information to each token. Semantic information is used first for each token; if no semantic information is available, lexical information is used. The client assigns a color to each token based on its lexical and syntactic information, ultimately achieving a highlighted display effect.

[0096] Step S2: When the user is editing the program to organize the unit logic code body, the access path stack is extracted according to the cursor truncation and operator matching rules, and combined with the full symbol index table to perform industrial code intelligent perception processing.

[0097] In one embodiment, when a user is editing the logic code body of a program organization unit, the access path stack is extracted according to the cursor truncation and operator matching rules, and combined with the full symbol index table to perform intelligent industrial code perception processing. This includes: when the user is detected editing the logic code body of the program organization unit, obtaining the text and position of the line where the user's current cursor is located, and extracting the prefix of the identifier that the user is currently inputting from the cursor position to the left as the name to be completed; determining whether there is an access operator to the left of the name to be completed. If there is no access operator, or the termination condition is met, then ordinary variable completion processing is performed, and the path stack is directly set to empty; if there is a member access operator, then the expression is parsed backward from the left of the member access operator, and array subscripts, function calls, and spaces are automatically skipped during the backward parsing process. Multi-level access operators are extracted and added to the path stack in sequence until the termination condition is met; based on the path stack and access operators, an access path stack for code completion is obtained, and the access path stack is used as a key to perform intelligent industrial code perception completion processing in combination with the full symbol index table.

[0098] In one embodiment, obtaining an access path stack for code completion based on the path stack and access operators, and using the access path stack as a key, combined with a full symbol index table, to perform intelligent code completion processing includes: generating context information containing the complete path and completion information based on the path stack addition result and the name to be completed, as an access path stack that meets the expected requirements; the client sends the access path stack to the server, and the server performs top-down type deduction and completion processing based on the access path stack and the full symbol index table to filter candidate members; calculating the matching score of each candidate member name according to the name to be completed and the scoring calculation rules, sorting the candidate members based on the matching scores, and selecting the candidate members as the perception result to perform industrial code completion processing.

[0099] It should be explained that the main implementation of step S2 is to achieve low-latency, coordinate-independent code completion based on the access path stack. Since the syntax of industrial code is simpler than that of mainstream programming languages ​​(such as C++ and Java), the following example illustrates the code completion scenario involved in this embodiment:

[0100] Example 1: obj.x | Path stack: accessPath={obj}, suffix=x;

[0101] Example 2: ptr->val | Path stack: accessPath={ptr}, suffix=val;

[0102] Example 3: data[0].name|path stack: accessPath={data}, suffix=name;

[0103] Example 4: func(obj.ab|) path stack: accessPath={obj,a}, suffix=b;

[0104] Example 5: myVar | Path stack: accessPath={}, suffix=myVar;

[0105] Example 6: obj.| or ob->| path stack: accessPath={obj}, suffix=;

[0106] Example 7: data[ax|].name path stack: accessPath={a}, suffix=x;

[0107] Example 8: ptr->val.c->ma | Path stack: accessPath={ptr,val,c,m},suffix=a;

[0108] Example 9: a->b.ad.a | Path stack: accessPath={d}, suffix=a.

[0109] Where | represents the cursor position, accessPath is an ordered list of strings stored in the order of access, and suffix is ​​the prefix of the identifier to be completed.

[0110] like Figure 3 As shown, when a user enters code in the POU editor, the access path stack extraction process for code completion is automatically triggered. The specific code completion path stack extraction process is as follows:

[0111] The process retrieves the text and position of the current cursor line, extracts the prefix of the identifier the user is currently typing from the cursor position to the left as the name to be completed, and determines whether the left side of the prefix is ​​a "." or "->" member access operator. If no such operator exists, or if a space, line beginning, or other termination condition is encountered during parsing, it is processed as a normal variable completion, and the path stack is immediately emptied. If a member access operator exists, the expression is parsed in reverse order from the left side of the operator to the left, automatically skipping array indices [], function calls (), and spaces, extracting identifiers from multi-level member access chains and adding them to the path stack until a termination condition is encountered. Finally, the parsed access path is combined with the prefix to be completed to generate a context object containing the complete path and completion information. This process can correctly handle various complex scenarios such as obj.x|, func(obj.ab|), and data[ax|].name, recognizing and parsing the entire access expression before the cursor into a linear access path stack, obtaining the expected access path and prefix result.

[0112] After the client extracts the access path stack used for code completion, it sends this information to the server. The server performs a top-down type deduction and completion based on the path stack and the full symbol index table. In the fuzzy matching stage, a matching score is calculated for each candidate member name based on the prefix to be completed by the user (i.e., the suffix in the path stack). The lower the score, the higher the priority. The specific rules are as follows:

[0113] Rule 1: First letter matching, that is, the prefix and the first character of the candidate word (ignoring capitalization) must be the same, otherwise it will be eliminated directly;

[0114] Rule 2: Base score, which is the initial score obtained by subtracting the prefix length from the candidate word length;

[0115] Rule 3: CamelCase bonus, which is -5 points for matching uppercase letters;

[0116] Rule 4: Underscore bonus, meaning if the matched character is an underscore, -5 points;

[0117] Rule 5: Skip penalty, i.e., the current match is not consecutive with the previous match in the candidate words, +5 points;

[0118] Rule 6: Prefix matching bonus, i.e., in the final statistics, if all matching characters match consecutively from the beginning of the candidate word, -30 points are awarded;

[0119] Rule 7: Perfect match bonus, that is, in the final statistics, if the prefix length is equal to the candidate word length (and the prefix match is satisfied), -50 points, for a total of -80 points;

[0120] Rule 8: Score normalization, that is, the final score plus a fixed bias of 200 and the maximum value of 1 are taken;

[0121] Rule 9: The lower the final score, the higher the candidate word ranks. Candidates that do not match receive a negative score (or are directly excluded).

[0122] By following the rules and steps described above, candidate members can be ranked as follows: exact match > prefix match > consecutive match > skip match. Members named with camelCase and underscores receive an additional ranking advantage. The entire process does not require access to any source code files or syntax trees. The process is as follows: Figure 4 As shown, the entire process uses the path stack as a key to perform continuous searches in the pre-built full symbol index table, which is highly efficient and unaffected by the size of the POU code.

[0123] According to another embodiment of the present invention, an industrial code intelligent perception system based on declaration implementation separation and path stack is also provided, the system comprising:

[0124] The declaration implementation separation module is used to parse the declaration code of the project file when the project is loaded, build and maintain a full symbol index table based on the declaration code, and use the full symbol index table to query and highlight the logical code body of the program organization unit;

[0125] The path stack generation module is used to extract the access path stack according to the cursor truncation and operator matching rules when the user edits the program's organizational unit logic code body, and combine it with the full symbol index table to perform industrial code intelligent perception processing.

[0126] To facilitate understanding of the above technical solutions of the present invention, the working principle or operation method of the present invention in actual process will be described in detail below.

[0127] One of the objectives of this embodiment is to achieve second-level full symbol index construction and real-time semantic highlighting. When the project is loaded, only the declaration part of the project file is parsed, completely skipping the logic code body that is not currently in use. A unified full symbol index is constructed and maintained. The parsing of the POU code body is delayed until each POU is opened or modified on demand. On the local syntax tree constructed on demand, combined with the global symbol index, semantic-level highlighting is achieved through two token parsings (lexical parsing and semantic parsing), thereby decoupling global resource consumption from code body size.

[0128] The second objective of this embodiment is to achieve low-latency, coordinate-independent code completion. The path parsing of nested expressions is moved to the client side, and the server directly uses the structured path sequence to look up the table in the symbol index, eliminating the need for source file loading and syntax tree backtracking. A comparison of Java, C++ code projects, and industrial programming language projects is shown in Table 1.

[0129] Table 1 Comparison of Java, C++ code projects, and industrial programming language projects

[0130]

[0131] For example, the project file project.asac, after removing device configuration and other content unrelated to code understanding, mainly contains the following:

[0132] "TYPE MyStruct :

[0133] STRUCT

[0134] name : STRING;

[0135] age : USINT;

[0136] END_STRUCT;

[0137] END_TYPE

[0138] TYPE ColorEnum : (

[0140] RED,

[0141] YELLOW );

[0143] END_TYPE

[0144] TYPE MyArray :

[0145] ARRAY [0..1] OF BOOL;

[0146] END_TYPE

[0147] PROGRAM PrgMain

[0148] VAR stu : MyStruct; END_VAR

[0149] VAR color1 : ColorEnum; END_VAR

[0150] VAR bVar1 : UINT; END_VAR

[0151] bVar1 := 1;

[0152] FOR bVar1 := 1 TO 100 BY 1 DO

[0153] stu.age:=bVar1;

[0154] END_FOR;

[0155] FOR bVar1 := 1 TO 100 BY 1 DO

[0156] stu.age:=bVar1;

[0157] END_FOR;

[0158] FOR bVar1 := 1 TO 100 BY 1 DO

[0159] stu.age:=bVar1;

[0160] END_FOR;

[0161] END_PROGRAM

[0162] FUNCTION_BLOCK MyFb

[0163] VAR bVar1 : DINT; END_VAR

[0164] FOR bVar1 := 1 TO 100 BY 1 DO

[0165] bVar1 := 2;

[0166] END_FOR;

[0167] END_FUNCTION_BLOCK

[0168] CONFIGURATION config

[0169] VAR_GLOBAL

[0170] bGlobalVar1 : BOOL;

[0171] END_VAR

[0172] VAR_GLOBAL

[0173] bGlobalVar2 : BOOL;

[0174] END_VAR

[0175] VAR_GLOBAL

[0176] bGlobalVar3 : BOOL;

[0177] END_VAR

[0178] END_CONFIGURATION.

[0179] The bolded parts are declarations, and the rest is the implementation. The example shows relatively little implementation code; in real projects, the implementation code is the core functionality and will far exceed the declaration code. `TYPE…END_TYPE` represents custom data types. In `PROGRAM…END_PROGRAM` and `FUNCTION_BLOCK…END_FUNCTION_BLOCK`, `VAR…END_VAR` represents variables defined in the POU, and `VAR_GLOBAL…END_VAR` represents global variables. These are all declaration code. The remaining parts in `PROGRAM…END_PROGRAM` and `FUNCTION_BLOCK…END_FUNCTION_BLOCK` are the implementation code.

[0180] The project organization structure shows that the industrial code project is managed by a single project file, with declarations and implementations stored separately, exhibiting clear boundaries. When building the symbol index, only the declaration code is analyzed. After construction, the symbol index includes the structure type MyStruct, the enumeration type ColorEnum, the array type MyArray, global variables bGlobalVar1, bGlobalVar2, bGlobalVar3, POU PrgMain and its internal variables, and POU MyFb and its internal variables. During semantic highlighting, if POU PrgMain is enabled, only the implementation code of POU PrgMain needs semantic highlighting; there is no need to parse the implementation code of other POUs. The effect is as follows: Figure 5 As shown.

[0181] When completing the code, if the user inputs "stu" on the front end, the system will parse the path stack: accessPath={" "}, suffix="stu". After the back end queries and evaluates the score, the result is as follows: Figure 6 As shown.

[0182] The performance evaluation data for this embodiment is shown in Table 2:

[0183] Table 2. Performance Evaluation Data Table

[0184] This embodiment significantly improves the speed of project initialization and index update. Existing LSP solutions require full parsing of the entire project file, while the large proportion of logic code does not contribute to the global symbol index but consumes a lot of parsing time. This embodiment only parses the declaration code during index construction, completely skipping all logic implementation code. Since the amount of declaration code is small, the parsing overhead is reduced. Therefore, the index construction time is decoupled from the total amount of project code. Compared with existing solutions, this embodiment can shorten the project initialization time from tens of seconds to seconds. Moreover, the index update is only triggered when the user modifies the declaration (such as adding global variables or adding data types), and there is no need to update the index when modifying the logic implementation code that accounts for the largest proportion of the project, further reducing unnecessary computational overhead. This embodiment achieves a significant reduction in code completion response latency, and its performance is unaffected by code size. Because the existing solution's coordinate backtracking mode requires the server to traverse the syntax tree from the cursor position upwards for each completion, the deeper the nesting level and the larger the file, the higher the latency. This embodiment moves path resolution to the client, and the server only needs to perform a pure table lookup operation in the pre-built full symbol index table, without performing any source file loading or syntax tree backtracking. This lightweight design not only reduces the latency of a single completion, but also makes the server performance almost unaffected by the size of the POU code, and the completion can respond instantly, ensuring a continuous smooth experience under fast coding.

[0185] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An industrial code intelligent perception method based on declaration-based implementation separation and path stack, characterized in that, The method includes: When a user opens a project file, the program organization units contained in each sub-project within the project file, the variable declarations within the program organization units, the global variable declarations, and the user-defined data types are extracted as declaration code. A full symbol index table is built and maintained based on the declaration code, and the logical code body of the highlighted program organization unit is queried using the full symbol index table; When the user edits the program to organize the unit logic code body, the access path stack is extracted according to the cursor truncation and operator matching rules, and combined with the full symbol index table to perform industrial code intelligent perception processing.

2. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 1, characterized in that, The process of constructing and maintaining a full symbol index table based on the declaration code, and using the full symbol index table to query the logical code body of the highlighted program organization unit includes: The declaration code is sent to the server. After receiving the declaration code, the server performs lightweight parsing processing to build a parse tree. The visitor pattern is used to traverse the parse tree to extract all symbol definitions in the declaration code, a full symbol index table is built, and the full symbol index table is optimized based on user editing behavior and symbol index update rules; When a user opens a program unit, the client extracts the corresponding logic code and sends it to the server. The server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, and outputs the highlighted logical code body of the program organization unit.

3. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 2, characterized in that, After receiving the declaration code, the server performs lightweight parsing processing to construct the parse tree, including: The server defines lexical and grammatical rules based on the industrial code syntax rules, and constructs a lexical analyzer and a grammatical analyzer in conjunction with an analyzer generation tool; The declaration code is input into the lexical analyzer to perform character-by-character string scanning. According to the lexical rules, the declaration code is divided into a stream of the smallest semantic units, and irrelevant content contained in the declaration code is discarded. The minimum semantic unit stream is input into the parser, which matches the sequence of minimum semantic unit streams step by step according to the grammar rules, verifies whether the minimum semantic unit streams conform to the defined language structure, and constructs a parse tree based on the verification results.

4. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 3, characterized in that, The process of using the visitor pattern to traverse the parse tree to extract all symbol definitions in the declaration code, establishing a full symbol index table, and optimizing the full symbol index table based on user editing behavior and symbol index update rules includes: The visitor pattern is adopted. The visitor interface generated by the analyzer generation tool traverses the parse tree and extracts symbol definition information during the traversal, and fills the predefined symbol index table one by one. After the visitor traversal is completed and the predefined symbol index table is filled, a full symbol index table containing all program organization unit names, type names, variable names, and their interrelationships is obtained; The client listens for user editing behavior and optimizes the full symbol index table by using timers and full and incremental symbol index update operations when the target editing behavior is detected. The target editing behaviors include: renaming, deleting, and modifying program organization unit names and variables within program organization units; renaming, deleting, and modifying global variables; and renaming, deleting, and modifying user-defined data types.

5. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 4, characterized in that, The timer is used to perform state detection at target intervals. If a target editing behavior is detected within each target interval, a full update symbol index operation or an incremental update symbol index operation is performed at the end of the current target interval. If no target editing behavior is detected within the current target interval, the full update symbol index operation or the incremental update symbol index operation is not performed.

6. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 5, characterized in that, The server segments the logical code of the program organization unit to generate the smallest semantic unit, and uses the full symbol index table to query the semantic information of the smallest semantic unit, outputting the highlighted logical code body of the program organization unit, including: The server divides the logical code of the program organization unit into the smallest independently identifiable semantic unit according to the lexical rules of the structured text language. The smallest semantic unit contains the text content and the position information of the text content in the logical code. The server labels each smallest semantic unit with lexical information including keywords, operators, literals and delimiters, and queries the semantic information corresponding to each smallest semantic unit in the full symbol index table to assign semantic information to the smallest semantic unit. Following the execution rule of semantics first and lexical information second, based on the lexical and syntactic information of each smallest semantic unit, the client assigns a color to each smallest semantic unit to produce a highlighting effect, and outputs the highlighted program organization unit logic code body.

7. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 1, characterized in that, The process of extracting the access path stack according to the cursor truncation and operator matching rules when the user edits the program organization unit logic code body, and combining it with the full symbol index table to perform industrial code intelligent perception processing includes: When the user's edit program organization unit logic code body is detected, the text and position of the current cursor line are obtained, and the identifier prefix that the user is typing is extracted from the cursor position to the left as the name to be completed; Determine if there is an access operator to the left of the name to be completed. If there is no access operator or the termination condition is met, then process it as a normal variable completion and set the path stack to empty. If a member access operator exists, the expression is parsed in reverse from the left side of the member access operator to the left. During the reverse parsing process, array subscripts, function calls and spaces are automatically skipped. Multi-level access operators are extracted and added to the path stack in sequence until the termination condition is met. Based on the path stack and access operators, an access path stack for code completion is obtained. The access path stack is then used as a key and combined with the full symbol index table to perform intelligent code completion processing.

8. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 7, characterized in that, The process of obtaining an access path stack for code completion based on the path stack and access operators, and using the access path stack as a key in conjunction with the full symbol index table to perform industrial code intelligent perception and completion processing includes: Based on the path stack addition result and the name to be completed, generate context information containing the complete path and completion information as an access path stack that meets the expected requirements. The client sends the access path stack to the server. The server uses the access path stack and the full symbol index table as a basis to perform top-down type inference and completion processing to filter candidate members. Based on the name to be completed and the scoring rules, the matching score of each candidate member name is calculated, and the candidate members are sorted according to the matching score. The selected candidate members are used as the perception result to perform industrial code completion processing.

9. The industrial code intelligent perception method based on declaration implementation separation and path stack according to claim 8, characterized in that, The scoring rules include initial letter matching, base score, camel hump bonus, underline bonus, jump penalty, perfect match bonus, and score normalization.

10. An industrial code intelligent perception system based on declaration-based implementation separation and path stack, used to implement the industrial code intelligent perception method based on declaration-based implementation separation and path stack as described in any one of claims 1-9, characterized in that, The system includes: The declaration implementation separation module is used to parse the declaration code of the project file when the project is loaded, build and maintain a full symbol index table based on the declaration code, and use the full symbol index table to query and highlight the logical code body of the program organization unit; The path stack generation module is used to extract the access path stack according to the cursor truncation and operator matching rules when the user edits the program's organizational unit logic code body, and combine it with the full symbol index table to perform industrial code intelligent perception processing.

Citation Information

Patent Citations

  • AST callback mechanism-based code sensitive information removal method, device and medium

    CN121786881A

  • Project configuration file analysis method and device and electronic equipment

    CN121879858A