Large language model and semantic feedback iterative optimization process control program generation method

By employing an iterative optimization method based on compiler and semantic expert feedback, the problems of data scarcity and syntactic semantics in generating process control programs from large language models are solved, enabling efficient and reliable generation of process control programs suitable for the field of industrial automation.

CN121069878APending Publication Date: 2025-12-05GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511222187.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing Large Language Models (LLMs) face problems such as data scarcity, syntax errors, and semantic bias when generating process control programs, making it difficult for the generated programs to meet the high reliability and safety requirements of industrial automation, and lacking an effective automated feedback mechanism.

Method used

By combining compiler feedback and semantic expert feedback in an iterative optimization method, a preference dataset and a multi-level expert evaluation system are constructed to optimize the generation of the process control program in real time. The compiler checks the grammatical correctness and the semantic evaluation assesses the logical rationality. Multiple rounds of iterative training are conducted to improve the generation quality of the model.

Benefits of technology

It significantly improves the generation quality and efficiency of process control programs, ensures that programs meet syntactic and semantic requirements in industrial applications, reduces reliance on large-scale datasets, and improves compilation pass rate and semantic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069878A_ABST
    Figure CN121069878A_ABST
Patent Text Reader

Abstract

The invention provides a large language model and semantic feedback iterative optimization process control program generation method, and belongs to the technical field of automatic control. And generating a preference data set through a compiler and semantic expert feedback, and performing model optimization by using the data set. The training bottleneck of a traditional data-driven model is avoided, and the generation capability of the model can be continuously optimized through dynamic adjustment and feedback under the condition of data scarcity; a dynamic iterative optimization mechanism is constructed by introducing the feedback of a compiler and a semantic expert, and the generated process control program can be checked and corrected in real time. The compiler feeds back to ensure the grammar correctness of the generated program, and a semantic expert evaluates from the aspects of logic and task intention. The problems of grammar errors and semantic mismatching in a traditional method are avoided. According to the method, the compiling passing rate and semantic accuracy of the process control program are effectively improved, and the quality and adaptability of the generated process control program are remarkably optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention provides a method for generating process control programs using large language models and semantic feedback iterative optimization, belonging to the field of automation control technology. Background Technology

[0002] With the rapid development of industrial automation technology, programmable logic controllers (PLCs), as core control devices, are widely used in intelligent manufacturing, energy management, and machinery control. The process control program defined by the IEC 61131-3 standard is one of the mainstream languages ​​for PLC programming. Its rigorous syntax and powerful functionality can meet the implementation needs of complex industrial logic. In recent years, code generation technology based on Large Language Models (LLMs) has rapidly emerged. For example, AI programming assistants generate code through natural language interaction, significantly improving development efficiency. However, while code generation technology for general-purpose programming languages ​​(such as Python) is relatively mature, generating process control programs for industrial automation still faces unique challenges. Industrial scenarios have extremely high requirements for the reliability and safety of process control programs. The generated process control programs must not only conform to strict syntactic specifications but also accurately map the semantic requirements of the actual control logic. Although LLMs perform excellently in general code generation, their application in the field of process control programs is limited by the scarcity of publicly available training data, the complexity of language features, and the stringent code quality requirements of industrial scenarios. Therefore, targeted optimization methods are urgently needed to adapt to industry needs.

[0003] Currently, the main bottleneck in LLM (Limited Linear Modulation) generation of process control programs lies in the limitations of data and evaluation mechanisms. First, publicly available datasets of process control program samples in the industrial sector are extremely scarce, and process control programs involving enterprise-specific logic are often not publicly available, leading to a severe data shortage problem for LLM training. Second, existing LLM generation methods are prone to syntax errors (such as type mismatches and control structure errors) and semantic deviations (such as logic not matching requirements). Traditional methods rely on manual annotation or fine-tuning using static datasets, which is costly and difficult to cover diverse industrial control scenarios. Furthermore, existing technologies lack efficient automated feedback mechanisms: compiler syntax checks only provide binary results and cannot guide model optimization for semantic accuracy; while manual verification can evaluate semantics, it is difficult to meet the real-time and large-scale deployment requirements of industrial scenarios. These problems often require multiple iterations of correction to the generated process control programs, significantly reducing development efficiency. Therefore, there is an urgent need for an iterative optimization framework that does not rely on massive amounts of labeled data and can integrate automated syntax verification and semantic evaluation to improve the reliability and practicality of LLM-generated process control programs, thereby promoting the evolution of industrial automation programming towards intelligence and low-code implementation. Summary of the Invention

[0004] This invention provides a method for generating a process control program for large language models and semantic feedback iterative optimization, which solves the following technical problems:

[0005] Data scarcity is a significant challenge in the field of process control program generation. Due to the scarcity of high-quality instruction-response datasets, traditional supervised learning-based training methods struggle to achieve high code generation quality, and the models exhibit weak generalization ability. This problem is particularly pronounced in complex industrial automation applications, where existing public datasets typically fail to cover all syntactic and semantic requirements. Therefore, effectively training efficient models despite the lack of a large number of high-quality samples has become a major technical challenge in process control program generation. This invention innovatively utilizes a preference dataset for iterative optimization by combining compiler feedback and semantic expert feedback, without relying on large-scale datasets, thus overcoming the training bottleneck caused by data scarcity.

[0006] Existing process control program generation models, especially those based on large-scale language models (LLMs), often struggle to avoid generating programs with syntactic errors or semantic inconsistencies. This is because many existing methods rely solely on the reasoning capabilities of pre-trained models, lacking effective error detection and correction mechanisms. Consequently, the generated programs may fail to compile or fail to accurately implement specific control logic. To overcome this problem, this invention proposes a dual evaluation mechanism combining compiler and semantic expert feedback. This mechanism integrates syntax checking and semantic evaluation to optimize the generated programs in real time during training. Through continuous feedback and iterative optimization, the model's compilation success rate and the program's semantic accuracy are effectively improved, ensuring that the generated process control programs meet practical application requirements.

[0007] The specific technical solution of this invention:

[0008] A method for generating a process control program for large language models and semantic feedback iterative optimization includes the following steps:

[0009] (1) The input is the pre-trained language model θ0 and the supervised fine-tuning dataset D. sft ,definition For V * →[0,1] represents the sequence of tokens on vocabulary V mapped to a probability distribution. D sft ={(d1,dr1),...,(d m ,dr m )} represents a dataset of instruction-response pairs consisting of several natural language intent descriptions and their corresponding process control procedures. Where d i User intent description; dr i Describe the corresponding process control procedures for this task.

[0010] (2) Constructing domain-specific instruction-response mappings: For the process control program generation task, the mapping function from natural language intent to process control program is established as: f:D intent →D ST The input is a user intent description d∈D. intent The output is the corresponding process control program dr∈D ST .

[0011] (3) Define the optimization objective of supervised fine-tuning of the SFT: Perform supervised fine-tuning of the initial parameter θ0, with the objective of maximizing the conditional log probability of instruction-response pairs in the training set.

[0012]

[0013] in, The objective function for supervised fine-tuning (SFT) is θ; the initial parameters of the model are π. θ (dr|d) represents the probability that the model generates a process control program dr given instruction d. The optimization objective is to improve the model's ability to understand natural language instructions and accurately generate process control programs that conform to grammatical rules.

[0014] (4) Supervise the implementation process of fine-tuning;

[0015] Data preprocessing stage: Standardize the original instruction-response samples; then use a syntax parser to verify the grammatical correctness of the process control program; construct training and validation sets to ensure balanced data distribution.

[0016] Model training phase: Input sample pairs (d, dr) ∈ D sft Calculate the cross-entropy loss:

[0017] loss = -logπ θ (dr|d) (2)

[0018] The gradient descent method is used to iteratively optimize the parameter θ0.

[0019] (5) Obtaining an initial model for domain adaptation: Through supervised fine-tuning, an initial model θ with basic generalization ability is obtained for the process control program generation task. sft θ sft =SFT(θ0,D sft );

[0020] (6) Constructing an instruction-preference triplet dataset for DPO training: Introducing a direct preference optimization mechanism for DPO, based on model θ sft Construct a preference-labeled dataset Where: d iA user intent description consistent with the SFT phase; Outputs of process control procedures that have been evaluated as being of high quality; Output of process control procedures that have been assessed as being of low quality.

[0021] Preference relationship is defined as: Preference tags are derived from expert evaluations and are determined based on multi-dimensional indicators such as program quality, functionality, and syntax correctness.

[0022] (7) Constructing an expert evaluation system integrating LLM and compiler feedback: To effectively evaluate the quality of candidate process control programs, a multi-level, multi-source expert feedback system is constructed. total It includes the following two components:

[0023] A. Compiler Feedback Module E compiler :

[0024] Syntax check: Syntax_Check(dr)∈{0,1}

[0025] Compilation success rate: Compile_Success(dr)∈{0,1}

[0026] Static analysis information: Static_Analysis(dr)

[0027] B. LLM Expert Feedback Module E LLM :

[0028] Program quality score: QualityScore(dr)∈[0,1]

[0029] Functionality: Function Match(d,dr)∈[0,1]

[0030] Best practice conformance: BestPractice(dr)∈[0,1]

[0031] The comprehensive evaluation function is defined as follows:

[0032] E total (d,dr)=αE compiler (dr)+βE LLM (d,dr) (3)

[0033] Where α+β=1 is used to adjust the weight ratio of the two types of feedback in the evaluation.

[0034] (8) Implement the online iterative DPO optimization process: DPO training is performed using an online iterative method. The t-th iteration includes the following three stages:

[0035] (a) Sample generation stage: Input the current model parameters θ t With instruction set D instructions For each instruction d∈D instructions Generate multiple candidate process control programs: A sampling strategy is introduced to enhance the diversity and coverage of candidate outputs.

[0036] (b) Preference labeling stage: Comprehensive expert evaluation of all candidate outputs.

[0037] Construct preference pairs from the rating results: (dr win ,dr lose whereE total (d,dr win E total (d,dr lose The preferences dataset for the current round is then compiled and generated.

[0038] (c) Model update phase: updating the dataset according to preferences Calculate the DPO loss; update the parameters to obtain the model state for the next round:

[0039]

[0040] (9) Define the optimization objective function for DPO;

[0041] To achieve preference-supervised model optimization, a direct preference optimization loss function is introduced:

[0042]

[0043] in, σ is the mathematical expectation symbol, representing the expected value over all samples in the dataset; σ is the sigmoid function, mapping probability differences to the (0,1) interval; β is a temperature parameter used to adjust the intensity of the contrast in preference distributions; π θ This is the currently optimized model; π ref For reference model; π θ (dr win |d) Generate the optimal process control program dr for the current model under the given instruction d. win The probability of π; ref (dr win |d) Generates the optimal process control program dr based on the reference model under given instruction d. win The probability of π; θ (dr lose |d) Generate the undesirable selection process control program dr for the current model under the given instruction d. lose The probability of π;ref (dr lose |d is a reference model that generates a suboptimal process control program dr under given instruction d. lose The probability of.

[0044] (10) Set stopping conditions for iterative optimization: To ensure the convergence of the online DPO optimization process, define the following three types of stopping criteria:

[0045] (a) Gain threshold:

[0046] ΔPerformance (t) =Performance (t) -Performance (t-1) (6)

[0047] When ΔPerformance (t) Iteration stops when the threshold is less than ε, where ε is a preset performance convergence threshold.

[0048] (b) Define the preference consistency ratio as When this indicator tends to stabilize, it is considered that the optimization is approaching convergence;

[0049] (c) Maximum number of rounds limit: When the preset iteration limit t is reached. max Training will automatically terminate when the time is right.

[0050] (11) Establish a distribution-aligned online preference sampling mechanism: each round of samples is generated by model θ t Distribution by current strategy Generate, forming a positive feedback iterative loop:

[0051] (12) Iterative process of preference sample collection:

[0052] Instruction set: Let X = {x1, x2, ..., x} M} represents a set of task instructions;

[0053] Sampling strategy for each round: In the i-th round, a subset X is extracted from X. i ={x i,1 ,…,x i,N},X i ~π i (X);

[0054] Where π i (X) represents the instruction sampling distribution of the i-th round, and N is the sampling size of each round.

[0055] (13) Implement multi-response sampling and diversity enhancement strategies;

[0056] For each instruction x i,n ∈Xi From the current model θ i Multiple responses are generated, and the response set is:

[0057]

[0058] Among them, Y i y is the set of responses generated in the i-th round of sampling; i,j This represents the j-th response generated in the i-th round; T represents the number of responses generated in each round of sampling. Let x be the model strategy / probability distribution for the i-th round; i,n This is the input for the nth instruction in the i-th round.

[0059] Diversity Enhancement: Temperature adjustment and Top-k strategies are introduced during the sampling process to enhance the semantic and structural coverage of candidate outputs.

[0060] (14) Construct a two-layer expert evaluation system for sample filtering and scoring;

[0061] The first-level expert evaluation system, compiler feedback: for each candidate response y i,j Perform a compilation check; if the candidate response compiles successfully, Compile. Check (y i,j =Success, and include it in the evaluable sample set.

[0062] The second-level expert evaluation system, LLM semantic scoring, evaluates the compiled sample set. The semantic reasonableness is evaluated using an LLM model, with evaluation dimensions including functional correctness, logical integrity, and security compliance. The scoring function is defined as follows:

[0063] Semantic Label (y i,j ,x i,n )∈{Positive,Negative} (8)

[0064] (15) Implement a hierarchical preference labeling mechanism;

[0065] For the candidate response set Y i ={y i,1 ,…,y i,T Perform two-layer expert annotation to generate preference pairs for training.

[0066] Specifically, it includes:

[0067] Standard path: If a response exists that compiles successfully and has correct semantics, mark it as a positive sample Y. pos The rest are negative samples Y neg Generate preference pairs (y w ,yl ), where y w ∈Y pos y l ∈Y neg .

[0068] Special path: If no positive samples are available, use the Large Language Model (LLM) score to generate relative preference pairs (y). w ,y l The rating is derived from the previous step, where the rating y w Higher than y l .

[0069] (16) Generate structured preference training data, and organize the annotation results of each round into a preference dataset:

[0070]

[0071] Among them, D i The structured preference training dataset generated in the i-th round; y w,n For a better response; y l,n A poor response is indicated by ">"; > indicates a preference relationship; N is the total number of instruction-response pairs.

[0072] (17) Perform preference consistency and quality verification;

[0073] Expert consensus verification:

[0074]

[0075] Wherein, Consistency Rate is the consistency rate, used to evaluate preference consistency and quality verification; |Y i | represents the total number of responses; Compile Check (y) is a function that performs a compilation check on the response y; Success indicates a successful compilation check; Semantic Label(y) is a function that performs a semantic label on the response y; Positive semantic label indicates a positive / good result.

[0076] Preference transitivity check: ensuring chained preference relationships Established.

[0077] Fuzzy sample filtering: Remove samples with preference score differences less than a threshold δ min The sample pairs are selected, and the pairs with significant preferences are retained for training.

[0078] (18) Dynamic sampling strategy optimization instruction selection strategy π i Design of (X):

[0079] Uniform sampling: π i(X) = Uniform(X), applicable to the initial round.

[0080] Difficult example sampling: π i (X)∝Error R ate(x) focuses on instructions that perform poorly in the model.

[0081] Diversity sampling: π i (X)∝Diversity S core(x) ensures instruction type coverage.

[0082] Adaptive sampling parameter adjustment:

[0083] The response count per instruction is dynamically adjusted based on the model's performance.

[0084] The quality of N is adjusted based on preference, which is the number of instructions per round.

[0085] (19) Construct a high-quality preference dataset and merge the datasets generated in all iteration rounds to form the final training set:

[0086]

[0087] Among them, D final To produce a final high-quality preference dataset; The union of all rounds of the dataset; (x, y w ,y l ) represents a sample from the final dataset, containing instructions, preferred responses, and corresponding outputs. w >y l This indicates a preference relationship that is superior to the previous one.

[0088] (20) Perform preference-driven model update training: Collect preference dataset θ in the i-th iteration. i Then, the DPO training process is executed to update the model parameters. Input:

[0089] Current model θ i ;

[0090] Preference data D i ={(x,y w ,y l );

[0091] Reference Model π ref =θ sft .

[0092] Execution preference optimization loss function Gradient updates yield new model parameters θ. i+1 .

[0093] (21) Apply the DPO loss function for parameter optimization;

[0094] Using the collected preference dataset D i The model is trained using the DPO loss function, with the training objective being to minimize the following loss function:

[0095]

[0096] in Generate the optimal response y for the current model under a given instruction x. w The probability of π; ref (y w |x) represents the optimal response y generated by the reference model under given instruction x. w The probability of; Generate a poorer response y for the current model under a given instruction x. w The probability of π; ref (y l |x) represents the poor response y generated by the reference model under given instruction x. w The probability of.

[0097] Parameter update process:

[0098] Calculate the gradient of the loss function:

[0099] Perform gradient descent update:

[0100] Where α is the learning rate parameter.

[0101] (22) Complete the model update for a single iteration and use the DPO update function to adjust the parameters:

[0102] θ i+1 =DPO Update (θ i D i (12)

[0103] This process includes batch processing preference dataset D i Forward propagation calculates the preference probability, backpropagation calculates the parameter gradient, and the optimizer updates the model parameters.

[0104] (23) For the updated model θ i+1 Performance verification: Test the quality of the process control program's generation on the validation set, calculate the compilation success rate and semantic accuracy, and compare it with the previous model θ. i To conduct a performance comparison.

[0105] (24) Determine whether to continue to the next round:

[0106] If \(i + 1 < I\) and there is still room for performance improvement, set the next iteration: \(i\leftarrow i + 1,\theta\) i \(\leftarrow\theta\) i+1 , return to step (12) of the preference collection process and continue the iteration. If the maximum number of iterations \(I\) is reached or the performance converges, terminate the iterative training process. Output the final optimized model.

[0107] (25) Obtain the final DPO optimized model: After \(I\) rounds of iterative training, the optimized model parameters \(\theta\) are finally obtained: final \(=\theta\) I .

[0108] (26) Comprehensive performance verification of the final model \(\theta\): final Evaluate the overall performance on an independent test set.

[0109] Conduct a comparative analysis with the baseline model.

[0110] Verify the performance of the model in the actual process control program generation task.

[0111]

[0112] The technical effects of the present invention are as follows:

[0113] 1) The implementation of the technical solution of the present invention effectively solves the problems of data scarcity and model training bottlenecks. Through the iterative optimization method combining the compiler and semantic expert feedback, the quality and efficiency of process control program generation are significantly improved. In traditional PLC programming, generating a process control program that meets requirements usually depends on a large-scale, high-quality instruction-response pair dataset. However, the relevant datasets in this field are scarce, resulting in limited accuracy and generalization ability of existing methods in process control program generation. Through the iterative optimization mechanism of the present invention, by using the preference dataset and combining real-time feedback, the model can be effectively optimized with fewer training samples, thus avoiding the dependence on a large amount of data. The feedback mechanism of the present invention can still generate high-quality process control programs in the case of data scarcity.

[0114] ​2) This invention also addresses common syntax and semantic errors in PLC programming. Traditional process control program generation models often lack real-time error detection and correction mechanisms, leading to potential syntax errors or logic that does not meet actual task requirements. This invention employs a dual mechanism of compiler feedback and semantic expert feedback to not only perform syntax checks on the process control program to ensure compilation success but also evaluate its semantics to ensure compliance with predetermined control logic. This feedback mechanism allows the model to correct errors promptly, thereby improving the quality and reliability of the process control program. After multiple rounds of iterative training, the final generated process control program exhibits greater operability and stability in industrial applications, providing a reliable guarantee for the efficient development of automated control systems. Attached Figure Description

[0115] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0116] The process of this invention is as follows Figure 1 As shown, the specific technical solution of the present invention will be described in accordance with this process and in conjunction with specific embodiments.

[0117] Example 1: Automatic generation and online iterative optimization of process control program for constant temperature heating tank (temperature closed-loop control)

[0118] Target scenario: The liquid in the process tank is maintained at 60℃±1℃. When the temperature exceeds 85℃, the interlock is triggered and the machine is immediately shut down under emergency stop conditions.

[0119] 1) Multi-response sampling and generation

[0120] Following steps (12) and (13), the system performs multi-response sampling (T=8) on the command, setting the temperature sampling parameters τ=0.8 and Top-k=5 to cover different control strategies. A candidate set {y1,…,y8} is obtained, which includes schemes based on PI control, schemes with dead-zone control, and schemes with different safety interlock combinations.

[0121] 2) Two-layer expert evaluation and preference construction

[0122] Compiler feedback (first level): Check the syntax and variable types of the 8 candidate programs, eliminate erroneous candidates, and retain the compilable subset.

[0123] LLM semantic scoring (second layer): Scoring is performed based on dimensions such as "whether over-temperature interlocking, emergency stop priority, integral anti-saturation, and dead-zone control are implemented". Positive and negative sample pairs (y) are constructed based on the scoring results. + ,y - ), then proceed to DPO optimization.

[0124] 3) DPO parameter update and convergence criteria

[0125] The probability distributions of the reference model and the current model are compared to calculate the DPO loss and update the parameters. The convergence criteria are: a single-round gain threshold ε = 0.5%, a preference consistency ratio consistently above 95%, and a maximum of 6 iteration rounds. In actual operation, the model met the convergence conditions in the 4th round.

[0126] 4) Functional characteristics of the final generated program

[0127] After iterative optimization, the generation program achieved the following:

[0128] The heating will automatically be cut off and the integral will be reset when the temperature exceeds 85℃.

[0129] When an emergency stop is triggered, the output is interrupted first.

[0130] Set a dead zone of ±0.5℃ around the target temperature of 60℃ to avoid frequent start-stop cycles;

[0131] The integral term has anti-saturation constraints to ensure control stability.

[0132] Example 2: Program generation and optimization for a two-stage conveyor line (sequential start-up, interlocking, and fault reset)

[0133] Target scenario: Motors M1 and M2 are connected in series for transmission. M1 starts first, and M2 starts after a 2-second delay. If a Jam alarm or emergency stop occurs, the machine will stop immediately and can only be restarted after manual reset.

[0134] 1) First round of generation and special path triggering

[0135] In the first round of generation (T=10, τ=0.9, Top-k=6), some candidates were able to compile successfully, but all of them lacked the "fault retention and manual reset" logic. Since there were no candidates that simultaneously met the requirements of "compilation + semantic correctness", the system entered the special path of step (15) and sorted the candidates by semantic scoring.

[0136] 2) Difficult example sampling and stepwise convergence

[0137] To strengthen the "safety interlock and manual reset" logic, the system introduced hard example sampling in subsequent iterations and adjusted the number of samples from 10 to 6. From the second to third rounds onwards, candidates that balanced compilation and semantic correctness emerged, and the process returned to the standard path to continue DPO updates.

[0138] 3) Functional characteristics of the final generated program

[0139] After the 3rd to 5th iterations, the program achieved:

[0140] Sequential startup: After M1 starts, M2 will run after a 2-second delay;

[0141] Emergency stop and alarm interlock: When any Jam or emergency stop is triggered, the entire line will stop immediately;

[0142] Fault retention and reset: The alarm remains latched after being triggered and can only be cleared and restarted manually with a reset.

[0143] Priority control: Stop and emergency stop commands have the highest priority to ensure safety.

[0144] 4) DPO Training and Convergence Determination

[0145] During the DPO training process, the construction of candidate pairs gradually transitioned from "only compiling correctly" to "compiling correctly + semantically correct", and finally reached the convergence condition in the 5th round (gain is below the threshold and preference consistency tends to stabilize).

[0146] This invention proposes a preference learning method that does not rely on large-scale, high-quality datasets. It generates a preference dataset through feedback from compilers and semantic experts, and then uses this dataset for model optimization. This method avoids the training bottleneck of traditional data-driven models, enabling continuous optimization of the model's generation capabilities through dynamic adjustment and feedback even in data-scarce environments. This technological innovation can reduce the workload of data collection and manual annotation in industrial automation programming, lowering costs while improving model performance and accuracy for specific tasks.

[0147] This invention constructs a dynamic iterative optimization mechanism by introducing feedback from compilers and semantic experts, enabling real-time checking and correction of the generated process control programs. Compiler feedback ensures the syntactic correctness of the generated programs, while semantic experts evaluate them from the perspectives of logic and task intent. This combination allows the model to not only generate programs that meet syntactic requirements but also ensure their practical application effectiveness in industrial automation, avoiding syntax errors and semantic mismatches found in traditional methods. This mechanism effectively improves the compilation pass rate and semantic accuracy of process control programs, significantly optimizing the quality and adaptability of the generated programs.

[0148] In this embodiment, other solutions may also be adopted:

[0149] 1) Model optimization method based on reinforcement learning (RL) and reward mechanism. In this alternative, reinforcement learning can be used instead of traditional supervised learning and DPO feedback mechanism. In this approach, the model gradually improves its generated process control program through interaction with the environment. The reinforcement learning algorithm provides rewards and penalties based on the compilation results and semantic evaluation of the generated process control program. The model adjusts its generation strategy based on feedback, thereby gradually optimizing the quality of the generated process control program over multiple rounds. Although this method does not directly utilize the preference dataset for iterative optimization, it can ensure that the generated process control program gradually improves its accuracy and ability to meet industrial requirements by dynamically adjusting the model's output. Although this method relies on long-term interaction with the environment and may require a long training period, it can automatically adapt to complex task requirements and directly optimize the generated process control program through the reward mechanism. The advantage of adopting this alternative is that it avoids over-reliance on human labels and can continuously learn and adjust in practical applications.

[0150] 2) Template- and rule-driven code generation method. This approach guides code generation through predefined rule sets and code templates. The model selects an appropriate template based on the input instructions and generates and fills in the code according to the rules. This method does not rely on a data-driven training process but generates compliant process control programs in a programmatic manner based on manually set rules and templates. The rules and templates can cover most common application scenarios in PLC programming, ensuring that the generated programs meet basic syntax and logic requirements. Although this approach has weaker flexibility and adaptability, it effectively reduces the complexity and resource consumption of model training, especially when dealing with structured and repetitive programming tasks, enabling the rapid generation of efficient and controllable process control programs. The advantage of using this approach is that it does not require large-scale datasets and complex training processes, making it suitable for some simple and standardized industrial automation tasks. However, this method may not be able to handle complex and dynamic programming needs. Therefore, for tasks requiring a high level of customization, the iterative optimization method of this invention still has greater advantages.

Claims

1. A process control program generation method of large language model and semantic feedback iterative optimization, characterized in that, Comprising the following steps: (1) input is a pre-trained language model θ0and a supervised fine-tuning dataset D sft , define V * →[0, 1] is a token sequence mapping on vocabulary V to probability distribution; D sft = {(d1, dr1),..., (d m , dr m )} is an instruction-response pair dataset consisting of a number of natural language intent descriptions and their corresponding process control programs; wherein d i is a user intent description; dr i is the process control program corresponding to the task description; (2) Constructing field-specific instruction-response mapping relationship: For process control program generation task, the mapping function from natural language intention to process control program is: f: D intent → D ST ; the input is user intention description d∈D intent , and the output is the corresponding process control program dr∈D ST ; (3) Define the optimization objective of supervised fine-tuning SFT: Supervised fine-tuning on initial parameters θ0, the goal is to maximize the conditional log probability of instruction-response pairs in the training set: wherein, is an optimization objective function for supervising fine-tuning of the SFT; θ is an initial parameter of the model; π θ (dr|d) represents the probability of the model generating the process control program dr under the condition of a given instruction d; the optimization objective is to improve the understanding ability of the model for natural language instructions and accurately generate process control programs conforming to grammar specifications; (4) Supervised fine-tuning implementation process: Data preprocessing phase: Standardize the original instruction-response samples; Then use the syntax parser to verify the grammatical correctness of the process control program; Construct training set and validation set division to ensure balanced data distribution; Model training phase: input sample pair (d, dr) e D sft , calculate cross-entropy loss: loss = -logπ θ (dr|d) (2) Using gradient descent method to iteratively optimize the parameters θ0; (5) Obtain the initial model adapted to the field: through the supervised fine-tuning process, obtain the initial model θ with basic generalization ability on the process control program generation task sft , θ sft = SFT(θ0, D sft ); (6) Construct instruction-preference triple data set for DPO training; (7) Constructing an expert evaluation system that integrates LLM with compiler feedback: Constructing a multi-level, multi-source expert feedback system E total , which includes the following two components: A, compiler feedback module E compiler : B, LLM expert feedback module E LLM : The comprehensive evaluation function is defined as: E total (d,dr) = aE compiler (dr) + bE LLM (d,dr) (3) Where α+β=1, used to adjust the weight proportion of the two types of feedback in the evaluation; (8) Implement the optimization process of online iterative DPO: Use online iterative method for DPO training; (9) Define the optimization objective function of DPO; (10) Set the stopping condition of iterative optimization; (11) Establish a distribution-aligned online preference sampling mechanism: each round of samples is generated by model θ t According to the current policy distribution Generate, form a positive feedback iterative cycle: (12) Iterative process of preference sample collection: Instruction set: Let X = {x1, x2, …, x M} be the task instruction set; Each round sampling strategy: At the i-th round, draw a subset from X: X i = {x i,1 ,…,x i,N}, X i ~ p i (X); where π i (X) denotes the i-th round instruction sampling distribution, N is the sampling size of each round; (13) Implement multi-response sampling and diversity enhancement strategy; For each instruction x i,n ∈ X i Generate a plurality of responses from the current model θ i The response set is: where Y i is the set of generated responses for the i-th round of sampling; y i,j is the j-th generated response for the i-th round of sampling; T is the number of generated responses per round of sampling; is the model policy / probability distribution for the i-th round; x i,n is the n-th instruction input for the i-th round; Diversity enhancement: Introduce temperature adjustment and Top-k strategy in the sampling process to enhance the coverage of candidate output in the semantic and structural layers; (14) Construct a double-layer expert evaluation system to filter and score the samples; (15) Implement a hierarchical preference labeling mechanism; The candidate response set Y i = {y i,1 ,…,y i,T} is double-layer expert annotated to generate training preference pairs; (16) Generate structured preference training data, and organize the labeling results of each round into a preference data set: where D i is the structured preference training dataset generated for the i-th round; y w,n is the better response; y l,n is the worse response; > denotes the preference relation; N is the total number of instruction-response pairs; (17) Perform preference consistency and quality verification; Expert consistency verification: wherein Consistency Rate is a consistency rate for evaluating preference consistency and quality verification; |Y i | is the total number of response sets; Compile Check (y) is a function of compiling checking for response y; Success is a state of successful compiling checking; Semantic Label(y) is a function of semantic labeling for response y; Positive semantic labeling is a positive / good result; Preference transitivity check: Ensures that a chain of preference relations holds; Fuzzy sample filtering: reject samples with preference score difference less than threshold δ min , and keep significant preference pairs for training; (18) Dynamic sampling policy optimization instruction selection policy p i Design of (X): Uniform sampling: π i (X) = Uniform(X), for the initial round; Hard example mining: π i (X)∝Error R ate(x), focusing on instructions where the model performs poorly; Diversity sampling: π i (X)∝Diversity S core(x), ensuring instruction type coverage; Adaptive sampling parameter adjustment: Adjust T, the number of responses per instruction, dynamically according to model performance; Adjust N, the number of instructions per round, based on preference quality; (19) Construct a high-quality preference data set, combine all the data generated in each iteration round to form the final training set: where D final is the final high-quality preference dataset; is the union of all round datasets; (x,y w ,y l ) are samples in the final dataset, containing instructions, preferred responses, and corresponding outputs; y w >y l denotes the preference relation of better than. (20) Perform preference-driven model update training: collect preference dataset Θ in the i-th iteration i After, perform DPO training process to update model parameters; input: Current model θ i ; Preference data D i = {(x, y w , y l ) ; Reference model p ref = 0 sft ; Perform gradient updates of the preference optimization loss function to get a new round of model parameters θ i+1 ; (21) Apply DPO loss function for parameter optimization; Utilizing the collected preference dataset D i The model is trained using the DPO loss function, with the training objective being to minimize the following loss function: in σ is the mathematical expectation symbol, representing the expected value of all samples in the dataset; σ is the sigmoid function, which maps the probability differences to the (0,1) interval; β is a temperature parameter used to adjust the intensity of the contrast in the preference distribution. Generate the optimal response y for the current model under a given instruction x. w The probability of π; ref (y w |x) represents the optimal response y generated by the reference model under given instruction x. w The probability of; Generate a poorer response y for the current model under a given instruction x. l The probability of π; ref (y l |x) represents the poor response y generated by the reference model under given instruction x. l The probability of; Parameter update process: Compute the gradient of the loss function: Performing gradient descent update: Where α is the learning rate parameter; (22) Complete single-round iteration model update, adjust parameters using DPO update function: θ i+1 = DPO Update (θ i , D i ) (12) The process comprises batch processing a preference dataset D i , forward propagation to compute preference probabilities, backward propagation to compute parameter gradients, and an optimizer to update model parameters; (23) On the updated model θ i+1 Performance verification: Test the process control program generation quality on the validation set, calculate the compilation success rate and semantic correctness rate, and compare with the previous round model θ i , performance comparison; (24) Determine whether to continue to the next round: If i + 1 < I and there is still room for performance improvement, set the next round of iteration: i i ← θ i+1 , return to step (12) preference collection process, continue to execute iteration; if the maximum number of iterations is reached or the performance converges, terminate the iteration training process; output the final optimized model; (25) Obtain the final DPO optimization model: After I rounds of iterative training, the final optimized model parameters: θ final = θ I ; (26) Final model θ final Performance verification: Evaluate overall performance on independent test set; Compare with baseline model for analysis; Verify the performance of the model in the actual process control program generation task.

2. The process control program generation method of claim 1, wherein: The specific method of step (6) is: Introducing a direct preference optimization (DPO) mechanism, based on a model θ sft , constructing a preference-labeled dataset where: d i is a description of the user intent consistent with the SFT phase; is a process control program output that is assessed to be of high quality; is a process control program output that is assessed to be of low quality; Preference relations are defined as: Preference labels are derived from expert evaluation, based on multi-dimensional indicators of program quality, function implementation, and grammatical correctness.

3. The process control program generation method of claim 1, wherein: In step (7): A, compiler feedback module E compiler : Syntax check: Syntax_Check(dr)∈{0,1}; Compilation success rate: Compile_Success(dr)∈{0,1}; Static analysis information: Static_Analysis(dr); B, LLM expert feedback module E LLM : Program quality score: Quality Score(dr)∈[0,1]; Function implementation degree: Function Match(d,dr)∈[0,1].

4. The process control program generation method of claim 1, wherein: The specific method of step (8) is: Use online iterative method for DPO training; The tth iteration includes the following three stages: (a) Sample generation phase: input current model parameters t with instruction set D instructions ; for each instruction d ∈ D instructions , generate multiple candidate process control programs: Introduce sampling strategy to enhance the diversity and coverage of candidate outputs; (b) Preference annotation phase: comprehensive expert evaluation of all candidate outputs Construct preference pairs from the score results: (dr win , dr lose ) where E total (d, dr win ) > E total (d, dr lose ); aggregate to generate the preference dataset for the current round (c) Model update phase: update the model according to the preference on the dataset Compute DPO loss; perform parameter update, resulting in the next model state as:

5. The process control program generation method of claim 1, wherein: The specific method of step (9) is: To achieve preference-based model optimization, introduce direct preference optimization loss function: Where, π θ (dr win |d) Generate the optimal process control program dr for the current model under the given instruction d. win The probability of π; ref (dr win |d) Generates the optimal process control program dr based on the reference model under given instruction d. win The probability of π; θ (dr lose |d) Generate the undesirable selection process control program dr for the current model under the given instruction d. lose The probability of π; ref (dr lose |d is a reference model that generates a suboptimal process control program dr under given instruction d. lose The probability of.

6. The process control program generation method of claim 1, wherein: The stopping condition of step (10) is: Ensure the convergence of online DPO optimization process, define the following three types of stopping criteria: (a) Gain threshold: ΔPerformance (t) = Performance (t) - Performance (t-1) (6) When ΔPerformance (t) Iteration stops when the threshold is less than ε, where ε is a preset performance convergence threshold. (b) defining a preference consistency ratio as When the index tends to stabilize, it is considered that the optimization tends to converge; (c) maximum number of wheel iterations limit: training is automatically terminated when a preset iteration upper limit t is reached. max training is automatically terminated.

7. The process control program generation method of claim 1, wherein: Step (14) specifically includes; First tier expert evaluation system, compiler feedback: for each candidate response y i,j Perform compilation checks; If the candidate response compiles successfully, Compile Check (y i,j ) = Success, include it in the assessable sample set Second layer expert evaluation system, LLM semantic score: on the sample set passed by compilation Evaluate its semantic reasonableness using LLM model, evaluation dimensions include functional correctness, logical integrity and security compliance, and the score function is defined as: Semantic Label(y i,j ,x i,n )∈{Positive,Negative} (8).

8. The process control program generation method of claim 1, wherein: Step (15) specifically includes: Standard path: if there exists a response that is both compiled and semantically correct, marked as positive sample Y pos The rest are negative samples Y neg ; generate preference pairs (y w , y l ) where y w ∈ Y pos y l ∈ Y neg ; Special path: if no positive samples, enable large language model LLM score to generate relative preference pair (y w ,y l ), score from previous step where score y w is higher than y l .

Citation Information

Cited By

  • Iterative optimization system and method for expert agent

    CN121998051A