Extensible structured text generation method and device in PLC field, equipment and medium
By building case library and instruction library, using language models and ANTLR parsers to generate and optimize structured text code in the PLC field, the problems of low ST code generation quality and steep learning curve in the existing technology are solved, and efficient and reliable code generation and adaptability are achieved.
Patent Information
- Application Number
- CN202510065128.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology has difficulty generating high-quality structured text (ST) code, especially in the PLC field, and the existing documentation fails to provide a comprehensive explanation of the full functionality of ST, resulting in developers facing steep learning curves and high possibility of errors.
By building a multi-task case library and a instruction library of commonly used programming instructions, the language model uses the semantic similarity of task description to retrieve similar task cases and instructions, generate structured text draft code, and perform syntax and semantic analysis through the ANTLR parser to correct error information to improve code quality.
It significantly reduces the time and work intensity of manually writing ST code, improves the quality and reliability of code generation, reduces the frequency of compilation errors, and adapts to ST variants designed by different manufacturers.
Smart Images

Figure CN120085869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an extensible structured text generation method, device, equipment and medium in the field of PLC. Background Art
[0002] Programmable Logic Controllers (PLCs) are the cornerstone of Industrial Control Systems (ICSs) in the broader industrial automation field and play a key role in the operation and management of critical infrastructure in different industries such as energy, manufacturing, and transportation.
[0003] Structured Text (ST) is a high-level language that follows the IEC 61131-3 standard and is crucial for PLCs as it can concisely express logic and seamlessly integrate with other languages within the same standard. The syntax and structure of ST are similar to high-level languages such as C or Fortran, which can simplify the learning curve for novice developers in industrial automation with experience in these programming languages. Additionally, ST can integrate more complex algorithms and data processing tasks in PLCs.
[0004] Despite the many benefits that ST offers, developers face difficulties due to its complex syntax and strict timing constraints. Compounding these challenges is the fact that existing documentation fails to provide a comprehensive explanation or definition of the full functionality of ST. Official manuals often introduce language features through only a few examples, which is usually insufficient for readers to thoroughly understand the language. Moreover, differences in language implementation among different vendors to accommodate their specific devices exacerbate the steepness of the learning curve and increase the likelihood of errors in important industrial automation projects.
[0005] In related technologies for ST code generation, the ST version is adjusted by modifying the syntax or enhancing the library to ensure better compatibility with their devices. These custom designs pose challenges to widely recognized public compilers such as MATIEC and formal verification tools such as nuXmv, thereby limiting the utility and generality of existing research and making the generated ST code difficult to be widely applied in different devices and scenarios. Additionally, current research does not evaluate the generated code based on actual industrial datasets or extensive benchmarks, lacking an evaluation of real industrial datasets or larger benchmarks. Summary of the Invention
[0006] The present invention provides an extensible structured text generation method, device, equipment and medium in the field of PLC, which solves the problem of how to improve the quality of automatically generated ST code.
[0007] To achieve the above object, the present application adopts the following technical solutions: In a first aspect, there is provided an extensible structured text generation method in the field of PLC, including: Construct a case library including multiple task cases and an instruction library including multiple common programming instructions; Based on the semantic similarity of the task description, retrieve a set number of similar task cases from the case library through a language model and sort them according to the similarity score; For each of the sorted task cases, retrieve the common programming instructions used by it from the instruction library, and generate an instruction information set C according to the corresponding instruction information D as the knowledge in a specific domain; where, Construct a context prompt and input it into the language model to generate a structured text draft code; the context prompt includes a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and the current problem requirement description; Use the structured text draft code as the input of the ANTLR parser for syntax analysis to obtain syntax error information and construct an abstract syntax tree; traverse the abstract syntax tree through the ANTLR Walker for semantic analysis to obtain semantic error information; If there is the syntax error information and / or semantic error information, provide it to the language model to generate a corresponding code repair solution and improve the code corresponding to the error information.
[0008] In the first possible implementation manner of the first aspect, each of the task cases includes a requirement description and a code snippet corresponding to the requirement description, and the requirement description is represented in vector form; divide the case library into a specific vendor library, an open-source function library, and a competition data set according to the source of the task case.
[0009] Based on the first possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, extract the corresponding requirement description from the code snippet through the language model combined with heuristic rules, or obtain the requirement description through reverse engineering of the code snippet.
[0010] In the third possible implementation manner of the first aspect, the instruction library further includes the instruction information corresponding to each of the common programming instructions, and the instruction information includes an instruction name, a function description and the corresponding application scenario, instruction parameters, and sample code; where, the instruction parameters include input and output parameters, parameter names, types, and explanations.
[0011] Based on the third possible implementation manner of the first aspect, in the fourth possible implementation manner of the first aspect, extract the instruction information from the vendor technical documents through the language model; where, the sources of the vendor technical documents include system manuals, online help websites, and instruction case code libraries.
[0012] In the fifth possible implementation of the first aspect, the syntax error information includes an error ID, a location, a context window, specific error information, and unmet syntax rules; the semantic error information includes error information about function calls, function definitions, parameter matching, variable initialization, and statement expressions.
[0013] In the sixth possible implementation of the first aspect, generating the corresponding code repair plan and improving the code corresponding to the error information specifically includes: Guiding the language model to analyze the cause of each error and the association between multiple errors according to the syntax error information and / or semantic error information; Adopting enhanced chain-of-reasoning prompts to list specific repair suggestions; generating the corresponding repair plan according to the repair suggestions; Using regular expressions to extract the corresponding code fragments and repair content from the repair plan, gradually replacing the original code, and generating a revised code version; Performing the syntax analysis and the semantic analysis on the revised code version and iteratively improving it.
[0014] In a second aspect, there is provided an extensible structured text generation device in the field of PLCs, including: An infrastructure construction module for constructing a case library including multiple task cases and an instruction library including multiple common programming instructions; A retrieval-enhanced generation module for retrieving a set number of similar task cases from the case library based on the semantic similarity of the task description through a language model and sorting them according to the similarity score; For each of the sorted task cases, retrieving the common programming instructions used by it from the instruction library, and generating an instruction information set C according to the corresponding instruction information D as domain-specific knowledge; where A code generation module for constructing context prompts and inputting them into a language model to generate a draft structured text code; the context prompts include a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and the current problem requirement description; A code inspection module for taking the draft structured text code as input to an ANTLR parser for syntax analysis to obtain syntax error information and constructing an abstract syntax tree; traversing the abstract syntax tree through an ANTLR Walker for semantic analysis to obtain semantic error information; A code improvement module, for a module, which, if there is the syntax error information and / or semantic error information, provides it to a language model, generates a corresponding code patching solution and improves the code corresponding to the error information.
[0015] In a third aspect, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the steps of the extensible structured text generation method in the PLC field as described in the first aspect are implemented.
[0016] In a fourth aspect, there is provided a readable storage medium, on which a program or instruction is stored, where when the program or instruction is executed by a processor, the steps of the extensible structured text generation method in the PLC field as described in the first aspect are implemented.
[0017] The extensible structured text generation method in the PLC field of the present invention has the following effects and advantages: This application provides an innovative method for automatically generating vendor-specific ST code from natural language requirements. By leveraging the powerful context understanding and transfer capabilities of large language models, automatically generating a code draft through automated retrieval of similar cases and instructions, combined with syntax and semantic checks using a customized ANTLR parser, and combined with an iterative error repair mechanism, an efficient knowledge retrieval and code generation process is achieved, significantly reducing the time and labor intensity of manually writing ST code, effectively reducing compilation errors in the code, and ensuring the quality and reliability of the generated code. It shows excellent domain knowledge adaptability for ST variants designed by different manufacturers.
[0018] The device, electronic device, and readable storage medium corresponding to the extensible structured text generation method in the PLC field of the present invention can achieve the same technical effects. To avoid repetition, they are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flowchart of an extensible structured text generation method in the PLC field provided by an embodiment of this application; Figure 2 It is a schematic flowchart of another extensible structured text generation method in the PLC field provided by an embodiment of this application; Figure 3 It is a schematic diagram of the content of a case library provided by an embodiment of this application; Figure 4 It is a schematic structural diagram of an extensible structured text generation device in the PLC field provided by an embodiment of this application; Figure 5Schematic diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the technical solutions in the embodiments of the present application are clearly described. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.
[0022] In the present application, the description of the method flow in the specification and the steps in the flowchart in the accompanying drawings of the present invention do not necessarily have to be strictly executed according to the step numbers. The method steps can change the execution order. Moreover, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0023] The following will be a detailed description of a method, device, equipment, and medium for generating extensible structured text in the PLC field provided by the embodiments of the present application in combination with the accompanying drawings and preferred embodiments.
[0024] First, the application scenario of the method for generating extensible structured text in the PLC field of the embodiments of the present application will be described in detail.
[0025] The embodiments of the present application provide a multi-agent framework, AutoPLC, for generating extensible structured text in the PLC field, to achieve the automatic generation of ST code and improve the development efficiency of engineers; it includes an infrastructure InfraPLC and a generation framework GenPLC.
[0026] The infrastructure InfraPLC includes: A case library containing numerous task cases, providing references for coding tasks; An instruction library containing common programming instructions, providing references for function calls in coding; A retrieval-augmented generation module, whose function is to retrieve highly similar cases and the instructions used in the cases from the established knowledge base and provide this knowledge to the core large language model (LLM).
[0027] These components are semi-automatically built based on different data sources. GenPLC utilizes the language resources provided by InfraPLC and the generation capabilities of the LLM to achieve end-to-end PLC programming.
[0028] The generation framework GenPLC includes: A code generation framework that includes a retrieval LLM for ranking the relevance of similar cases, an encoding LLM for generating ST code, and a verification LLM for implementing feedback-based optimization of the code with external tools; A customized open-source parser based on EBNF grammar that designs and implements a flexible and cost-effective ST variant code checker, supporting the identification of syntax and semantic errors in ST code.
[0029] Based on this, AutoPLC can automate the process from requirement understanding to code generation and ensure the quality of the generated code without manual intervention; it reduces the situation of code hallucinations, lowers the frequency of compilation errors, and improves the speed and flexibility of code generation by combining RAG technology and an external inspection tool mechanism; AutoPLC is easy to adapt to different ST variants, thus significantly improving the PLC programming efficiency.
[0030] Please refer to Figure 1-2 , the embodiments of the present application provide a scalable structured text generation method in the field of PLC, as Figure 1-2 shown, the recommended method of the embodiments of the present application includes: Step S1, construct a case library including multiple task cases and an instruction library including multiple common programming instructions.
[0031] In some possible implementation manners, each task case includes a requirement description and a code snippet corresponding to the requirement description, and the requirement description is represented in vector form; the case library is divided into a specific vendor library, an open-source function library, and a competition data set according to the source of the task case.
[0032] To facilitate efficient and effective code generation, the case library contains requirements and corresponding code implementations, which is a comprehensive knowledge base that promotes knowledge transfer and ensures code quality assurance. Specifically: the case library contains 914 code snippets collected from the following three strictly reviewed sources, which accurately reflect real-world domain requirements: vendor-specific libraries, open-source function libraries, and competition data sets.
[0033] OSCAT Library (718 cases): An open-source library widely used in industrial automation, providing modules for mathematical operations, logical control, and data processing. An active developer community and long-term maintenance ensure its reliability. Well-maintained CODESYS versions are available for selection.
[0034] Siemens LGF Library (151 cases): A general-purpose library renowned for its device control functions, adhering to strict industrial standards and optimized for the stable operation of Siemens automation systems.
[0035] Siemens Competition Dataset (45 cases): Real-world cases from Siemens, focusing on general function programming and device control tasks.
[0036] Furthermore, by combining a language model with heuristic rules, the corresponding requirement description is extracted from the code snippet, or the requirement description is obtained through reverse engineering of the code snippet.
[0037] In specific implementation, due to the complex descriptions and parameter specifications in the PDF documents, collecting the requirements of OSCAT and LGF poses challenges. To address this issue, we combine a large language model with heuristic rules to extract the requirement description for each code snippet in the collected libraries. For the code that already has a requirement document, we first split the document into chunks, each chunk describing a code snippet, and use GPT-4o to extract key information (such as function descriptions and input / output parameter tables) to convert this data into JSON format.
[0038] For the code without a requirement document, we use GPT-4o to directly reverse engineer the requirement description from the code. Figure 3 A data example of the case library in the embodiments of this application is shown. Generally, OSCAT is considered the simplest, LGF has medium difficulty, and according to the metrics of average implementation code length and average number of input / output parameters, Competition is considered the most challenging.
[0039] To facilitate faster retrieval, we convert the requirement description into vectors and perform fast matching based on the semantic similarity of the task description to identify potentially useful cases.
[0040] In some possible implementation manners, the instruction library includes a plurality of common programming instructions and instruction information corresponding to each common programming instruction. The instruction information includes an instruction name, a function description and a corresponding application scenario, instruction parameters, and example code; wherein, the instruction parameters include input / output parameters, parameter names, types, and explanations.
[0041] The correct use of specific instructions is crucial for generating accurate ST code. In practical applications, PLC experts often refer to official documents and learn relevant codes to ensure the correct use of instructions. To achieve the automated integration of instructions in the code pipeline, the embodiments of this application collect these instructions and store them in machine-readable formatted information.
[0042] Furthermore, instruction information is extracted from vendor technical documents through a language model; among them, the sources of vendor technical documents include system manuals, online help websites, and instruction case code libraries.
[0043] In specific implementation, instruction information is extracted from vendor technical documents, including system manuals, online help websites, and instruction case code libraries, etc. To enhance the readability of the description, the embodiments of this application further combine GPT-4o to summarize the functions of instructions and provide application scenarios. Each instruction in the instruction library of the embodiments of this application includes the following fields:
[0044] Name: The name of the instruction.
[0045] Description: The function description and application scenario of each instruction summarized by GPT-4o.
[0046] Parameters: The input and output parameters of each function, including parameter names, types, and explanations, etc.
[0047] Example Code: Typical usage examples of each function, and each code snippet is accompanied by annotations that clearly explain its function and usage.
[0048] Step S2, based on the semantic similarity of the task description, retrieve a set number of similar task cases from the case library through a language model and sort them according to the similarity score.
[0049] In the specific implementation process, considering that the LLM exhibits strong in-context learning ability and can transfer the knowledge in the given examples to the current task, the embodiments of this application design a RAG-style retriever to locate valuable information from the knowledge base we have established and use them to enhance the prompts of the LLM. To achieve efficient retrieval, the embodiments of this application first use vector retrieval to obtain the top few candidate cases with the highest matching degree from the case library. The information for the query is constructed through the requirements of the question (including task description and parameter information), and then the top k similar cases are returned according to the similarity score.
[0050] Specifically, to further select relevant cases from the candidate set, this application implements re - ranking the candidate case list by the LLM according to the function of each code case and its contribution to the current coding task. This process adopts a chain - of - reasoning design: first, it instructs the LLM to summarize the function of each case according to the task requirements and analyze its contribution to the current task, and then sorts the cases to ensure that the most relevant and beneficial cases are prioritized.
[0051] Step S3, for each sorted task case, retrieve the commonly used programming instructions it uses from the instruction library, and generate an instruction information set C according to the corresponding instruction information D as domain - specific knowledge; where 。
[0052] Step S4, construct context prompts and input them into the language model to generate a draft structured text code; the context prompts include a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and the description of the current problem requirements.
[0053] In the code generation stage, this application embodiment constructs a rich - information context prompt to guide the language model to generate code that meets the requirements. As Figure 2 shown, this prompt includes two parts: a knowledge part and a task part.
[0054] The knowledge part provides domain - related coding knowledge, including programming guidance and instruction details. The programming guidance is customized for the ST programming language and includes coding standards and formats to be followed; the instruction details are composed of the extracted instruction set.
[0055] The task part includes the tuples of the retrieved case requirements and code implementations, as well as the description of the current problem requirements. Incorporating the implementation details of similar cases into the prompt can activate the problem - related knowledge in the language model and enable it to better handle the code details of the ST programming language. For example, due to the cyclic execution characteristic of PLC programs, a common difficulty lies in how to correctly use local variables to maintain the state and process them in each new cycle. Such coding practices and specific PLC programming patterns are usually not explicitly presented in the problem requirements, making it difficult for the model to naturally follow the corresponding patterns without specialized training. Therefore, providing cases related to the requirements can help the model follow specific coding styles and patterns that conform to the actual situation of the project, thereby generating code that is easier to maintain and understand.
[0056] After constructing the context prompts, input them into the language model, and extract the generated draft structured text code for further optimization and improvement in subsequent stages.
[0057] Step S5: Use the structured text draft code as the input of the ANTLR parser for syntax analysis to obtain syntax error information and construct an abstract syntax tree; traverse the abstract syntax tree through the ANTLR Walker for semantic analysis to obtain semantic error information.
[0058] Furthermore, the syntax error information includes error ID, location, context window, specific error information, and non - satisfied syntax rules; the semantic error information includes error information about function calls, function definitions, parameter matching, variable initialization, and expressions of statements.
[0059] In the specific implementation process, although there are open - source ST verification tools such as MATIEC and IECChecker, it is difficult for them to meet the complex requirements of various variants in the modern industrial environment. On the other hand, although commercial programming platforms provide checking functions for specific ST implementations, their closed - source nature makes it difficult to integrate them into the automated code pipeline. Therefore, the embodiments of this application have developed a new flexible and cost - effective code checker for ST variants used in the modern industrial environment. Specifically, the embodiments of this application have designed a parsing solution based on the EBNF grammar definition and used ANTLR as the parsing engine to build the code checker. The checker of the embodiments of this application takes ST code as the input, sequentially executes multiple checking processes and returns checking feedback messages, which are divided into two parts: syntax analysis and semantic analysis.
[0060] Syntax analysis: Customize an ANTLR parser based on the open - source ST grammar to identify syntax errors in the code and provide a comprehensive error report. For the given ST code, the customized parser will analyze the code and construct an abstract syntax tree. When an error is found during the analysis process, the parser will provide detailed syntax error information, including the error ID, the location where the error occurs, the context window at the interruption position, as well as the specific error information and non - satisfied syntax rules. Then, the parser bypasses the error token and continues the parsing process. After the analysis is completed, detailed syntax error information and the abstract syntax tree will be returned.
[0061] Semantic analysis: For semantic analysis, customize an ANTLR Walker to perform a depth - first traversal of the abstract syntax tree generated during the syntax analysis process to check for semantic errors. The embodiments of this application have implemented semantic error checking for three main data types, including functions, variables, and statements. For functions, focus on checking function calls to confirm that the called function has been defined or an instruction library is provided, and the provided parameters match the parameter specifications declared by the function. For variables, mainly check whether they have been correctly initialized before use. For statements, perform basic verification to ensure that the expressions used in them are legally formed.
[0062] For example, in SCL, applying Typeof to non-IF and CASE statements is illegal. Once the semantic analysis is completed, detailed semantic error messages will be returned. Given that ST variants generally follow the IEC-61131 standard and share similar abstract syntax tree structures, the code inspection tool of the embodiments of the present application can easily adapt to different ST variants. By updating the EBNF grammar and making appropriate modifications to our customized AST Walker, promising verification functions tailored to specific ST variants can be achieved.
[0063] Step S6, if there are syntax error messages and / or semantic error messages, provide them to the large model to generate corresponding code patching solutions and improve the code corresponding to the error messages.
[0064] At this stage, we use a custom code checker to analyze the previously generated draft ST code. If an error is detected, the checker will provide detailed error messages and initiate an iterative improvement process. During this process, AutoPLC will guide the language model to analyze the cause of the error and the association between multiple errors, and generate corresponding patching solutions based on the error messages and ST code.
[0065] Specifically, the process of generating corresponding code patching solutions and replacing the code corresponding to the error messages includes the following steps: Step S601, guide the language model to analyze the cause of each error and the association between multiple errors according to the syntax error messages and / or semantic error messages; Step S602, adopt enhanced chain reasoning prompts to list specific repair suggestions; generate corresponding patching solutions according to the repair suggestions; Step S603, use regular expressions to extract the corresponding code fragments and patching content from the patching solution, and gradually replace the original code to generate a revised code version.
[0066] In some possible implementation manners, the method further includes: Step S604, perform iterative improvement on the syntax analysis and semantic analysis of the revised code version.
[0067] By repairing the prompt to guide the language model to comprehensively understand the code context with its own knowledge to determine the root cause of the error, rather than relying solely on the error feedback at the breakpoint. To improve the repair efficiency and reduce the risk of getting into an infinite loop, the embodiments of this application provide all error information at once, rather than performing independent repairs one by one. Using enhanced chain-of-reasoning prompts, list specific repair suggestions first, and then generate corresponding code patching solutions to improve the correctness of the patching. Further utilize regular expressions to extract relevant code snippets and patching content from the LLM's response, and gradually perform replacement operations on the code to generate a revised code version. Subsequently, these revised codes will be checked in subsequent self-improvement iterations. The iteration process will terminate when no errors are detected or the preset maximum number of iterations is reached.
[0068] Based on the above technical solutions, this application has the following effects and advantages: The AutoPLC of this application significantly reduces the time and workload of manually writing ST code through automated ST code generation. It provides an innovative method to automatically generate vendor-specific ST code from natural language requirements. In addition, AutoPLC has shown significant advantages over other current ST code generation technologies in comprehensive tests on three test datasets. Different from the existing methods that rely on manual interaction and only generate native ST code, AutoPLC has a lightweight design, does not require any form of manual participation in the intermediate generation steps, and has a low coupling degree between modules. Only a few modules need to be replaced to generate code for ST variants. The applicant further verified the effectiveness of AutoPLC in code generation in specific real-world scenarios through extensive experiments, and confirmed the positive impact of each module on code generation through ablation experiments.
[0069] Through the RAG technology, AutoPLC can effectively provide tutoring support for similar cases for the current generation task, enabling the large model to have more domain knowledge and reducing the hallucination situation in the generated code. After generation, external syntax and semantic checking tools can repeatedly perform syntax and semantic checks on the generated code, which reduces the frequency of compilation errors in the code. Compared with traditional formal code generation methods, AutoPLC has natural flexibility and speed advantages for the task of generating code from natural language requirements. Compared with the method of generating code with the help of large models, it has an innovative retrieval-enhanced generation and external checking tool mechanism, which ensures the quality of the generated code.
[0070] These innovative points of the AutoPLC in the embodiments of this application give it significant advantages in automatically generating ST code, improving the speed of writing ST code and effectively assisting engineers in coding requirements.
[0071] The method AutoPLC in the embodiments of this application realizes fully automatic ST code generation for the first time, and innovatively combines RAG and a replaceable syntax and semantics checking module to assist code generation, showing excellent domain knowledge adaptability for ST variants designed by different manufacturers.
[0072] See Figure 4 , corresponding to the embodiments of the extensible structured text generation method in the above PLC field, the embodiments of this application provide an extensible structured text generation device in the PLC field, and the device includes: An infrastructure construction module 1001, configured to construct a case library including a plurality of task cases and an instruction library including a plurality of common programming instructions; A retrieval enhanced generation module 1002, configured to retrieve a set number of similar task cases from the case library based on the semantic similarity of the task description through a language model and sort them according to the similarity score; For each of the sorted task cases, retrieve the common programming instructions used by it from the instruction library, and generate an instruction information set C according to the corresponding instruction information D as domain-specific knowledge; where A code generation module 1003, configured to construct context prompts and input them into a language model to generate a structured text draft code; the context prompts include a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and the current problem requirement description; A code checking module 1004, configured to use the structured text draft code as an input to an ANTLR parser for syntax analysis to obtain syntax error information and construct an abstract syntax tree; traverse the abstract syntax tree through an ANTLR Walker for semantic analysis to obtain semantic error information; A code improvement module 1005, configured to, if there is the syntax error information and / or semantic error information, provide it to the language model to generate a corresponding code repair solution and improve the code corresponding to the error information.
[0073] The above extensible structured text generation device in the PLC field implements the steps and various processes of the above embodiments of the extensible structured text generation method in the PLC field, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0074] See Figure 5, corresponding to the embodiment of the extensible structured text generation method in the above PLC field, the embodiment of the present application provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps and various processes of the embodiment of the extensible structured text generation method in the above PLC field, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0075] The memory 1009 can be used to store software programs and various data. The memory 1009 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 can include volatile memory or non-volatile memory, or the memory 1009 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0076] The processor 1010 can include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010.
[0077] Corresponding to the embodiment of the extensible structured text generation method in the above PLC field, an embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps and various processes of the embodiment of the extensible structured text generation method in the above PLC field are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.
[0078] Wherein, the processor is the processor in the electronic device described in the embodiment of the present application above. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.
[0079] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0081] It can be understood that the embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. As is known to those skilled in the art, without departing from the spirit and scope of the present invention, various changes or equivalent substitutions can be made to these features and embodiments. In addition, those of ordinary skill in the art can modify these features and embodiments in light of the inspiration or teachings of the present application to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.
Claims
1. A method for generating scalable structured text in the field of PLC, characterized in that: include: Build a case library including multiple task cases and an instruction library including multiple commonly used programming instructions; Based on the semantic similarity of the task descriptions, a set number of similar task cases are retrieved from the case library through a language model and sorted according to similarity scores; For each of the sorted task cases, the commonly used programming instructions are retrieved from the instruction library, and an instruction information set C is generated according to the corresponding instruction information D as knowledge in a specific field; wherein, Constructing contextual prompts and inputting language models to generate structured text draft codes; the contextual prompts include a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and current problem requirement description; The structured text draft code is used as an ANTLR parser input for syntax analysis, syntax error information is obtained and an abstract syntax tree is constructed; the abstract syntax tree is traversed by ANTLR Walker for semantic analysis, and semantic error information is obtained; If the grammatical error information and / or semantic error information exists, it is provided to the language model, a corresponding code patching solution is generated, and the code corresponding to the error information is improved.
2. The method for generating scalable structured text in the PLC field according to claim 1, characterized in that: Each of the task cases includes a requirement description and a code snippet corresponding to the requirement description, and the requirement description is expressed in a vector form; the case library is divided into a specific supplier library, an open source function library, and a competition data set according to the source of the task case.
3. The method for generating scalable structured text in the PLC field according to claim 2, characterized in that: The corresponding requirement description is extracted according to the code snippet by combining a language model with heuristic rules, or the requirement description is obtained by reverse engineering the code snippet.
4. The method for generating scalable structured text in the PLC field according to claim 1, characterized in that: The instruction library also includes instruction information corresponding to each of the commonly used programming instructions, and the instruction information includes instruction name, function description and corresponding application scenario, instruction parameters and sample code; wherein the instruction parameters include input and output parameters, parameter name, type and explanation.
5. The method for generating scalable structured text in the PLC field according to claim 4, characterized in that: The instruction information is extracted from the supplier's technical documents through a language model; wherein the sources of the supplier's technical documents include system manuals, online help websites, and instruction case code libraries.
6. The method for generating scalable structured text in the PLC field according to claim 1, characterized in that: The grammatical error information includes error ID, location, context window, specific error information and unsatisfied grammatical rules; the semantic error information includes error information of function call, function definition, parameter matching, variable initialization and statement expression.
7. The method for generating scalable structured text in the PLC field according to claim 1, characterized in that: The generating of the corresponding code patching solution and improving the code corresponding to the error information specifically includes: Guide the language model to analyze the cause of each error and the relationship between multiple errors according to the grammatical error information and / or semantic error information; Using enhanced chain reasoning prompts, listing specific repair suggestions; generating corresponding repair plans based on the repair suggestions; Extracting corresponding code snippets and patch contents from the patch solution using regular expressions, gradually replacing the original code, and generating a revised code version; The revised code version is subjected to the syntax analysis and the semantic analysis and iterative improvement.
8. An extensible structured text generation device in the field of PLC, characterized in that: include: An infrastructure building module, used for building a case library including a plurality of task cases and an instruction library including a plurality of commonly used programming instructions; A retrieval enhancement generation module, for retrieving a set number of similar task cases from the case library through a language model based on the semantic similarity of the task description and sorting them according to similarity scores; For each of the sorted task cases, the commonly used programming instructions are retrieved from the instruction library, and an instruction information set C is generated according to the corresponding instruction information D as knowledge in a specific field; wherein, A code generation module, used to construct contextual prompts and input a language model to generate structured text draft code; the contextual prompts include a knowledge part and a task part, the knowledge part includes programming guidance and the instruction information set, and the task part includes the sorted task cases and current problem requirement description; A code checking module is used to perform syntax analysis on the structured text draft code as an input of an ANTLR parser, obtain syntax error information and construct an abstract syntax tree; perform semantic analysis by traversing the abstract syntax tree through an ANTLR Walker, and obtain semantic error information; The code improvement module is used for providing the syntax error information and / or semantic error information to the language model if there is any, generating a corresponding code patching solution and improving the code corresponding to the error information.
9. An electronic device, characterized in that: The electronic device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for generating structured text extensible in the PLC field as claimed in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or an instruction, and when the program or the instruction is executed by the processor, the steps of the method for generating an extensible structured text in the PLC field as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
PLC software programming aided design method
CN105159656A
Edge controller integrating deep learning and PLC language and code generation method
CN117667045A
Code generation optimization method and device for large language model, equipment and medium
CN117724695A
Code generation method and device, computer equipment and storage medium
CN118778939A
Artificial intelligence auxiliary programming construction method
CN119179467A
Cited By
Large language model JSON output repairing method and system based on ANTLR grammar analysis
CN120509399A
Method and system for repairing JSON output of large language model based on ANTLR grammar parsing
CN120509399B
Structured text code generation and review method
CN120743231A
A method for generating and reviewing structured text code
CN120743231B
Code generation method and device, equipment and storage medium
CN121477769A