Optimal set point tracking control method of injection molding system based on output feedback

By constructing an augmented system and an infinite-domain cost function based on output feedback, designing an output feedback Q function, and iteratively solving for the optimal control strategy, the problem of improper selection of discount factor in injection molding system is solved. This achieves optimal setpoint tracking control without state measurement, ensuring the stability and accuracy of the system.

CN121290727AActive Publication Date: 2026-01-09NORTHEASTERN UNIV CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511844696.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-01-09
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

Existing setpoint tracking control methods for injection molding systems rely on improper selection of discount factors, leading to stability risks and making it difficult to achieve optimal control when the state is not fully known.

Method used

By adopting an output feedback-based approach, an augmented system and an infinite-domain cost function are constructed. The augmented state vector is equivalently represented as a non-minimum state vector using a state parameterization method. An output feedback Q function is designed, and the optimal control strategy is iteratively solved using an output feedback Q learning algorithm, thus avoiding dependence on discount factors.

Benefits of technology

It achieves optimal setpoint tracking control of the injection molding system without requiring all system state measurement information, ensuring closed-loop stability and tracking accuracy, and effectively suppressing unknown disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121290727A_ABST
    Figure CN121290727A_ABST
Patent Text Reader

Abstract

The invention provides an injection molding system optimal set point tracking control method based on output feedback, and relates to the technical field of automatic control of an injection molding system.The method comprises the steps that a discrete time state-space equation of an injection molding system is established, and the equation at least comprises a state vector and an output vector; subtracting the output vector from a corresponding set point in the reference trajectory to obtain a tracking error, and constructing an augmentation system according to the state vector and the tracking error; constructing an infinite domain cost function based on an augmentation system; obtaining a non-minimum state vector based on the augmented system by using a state parameter method, and defining an output feedback Q function based on the vector; input and output data of an injection molding system are collected to design a dead-beat controller, the controller serves as an initial stable control strategy to execute an output feedback strategy Q learning algorithm, an output feedback Q function is converged through iterative solution, an optimal output feedback control strategy for minimizing an infinite domain cost function is obtained, and the control precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic control of injection molding systems, and particularly relates to an optimal set point tracking control method for an injection molding system based on output feedback. BACKGROUND

[0002] The set point tracking control problem is a typical control task in practical control systems. Taking the injection molding process as an example, the process mainly includes three stages of filling, pressure maintaining (or compression / holding) and cooling. Among them, in the pressure maintaining stage, the nozzle pressure value is a key process variable, and accurate set point tracking of the key process variable is the key to ensuring the quality of the final product. The traditional set point tracking control method obtains the optimal control strategy (or controller) by minimizing a cost function composed of tracking error and control input. However, when the reference trajectory cannot asymptotically converge to zero, the integral of the cost function in the infinite time domain will tend to infinity, resulting in invalid problem definition. To solve this problem, a common practice in the prior art is to introduce a discount factor into the cost function to ensure its boundedness. However, the stability of this method depends heavily on the selection of the discount factor. If the value is not properly selected, the stability of the closed-loop system may be affected. In particular, most existing methods do not analyze the impact of the discount factor on stability, but only conservatively select a larger value, which limits its reliability in practical industrial applications.

[0003] On the other hand, the injection molding process has strong nonlinearity, time-varying nature and complex internal mechanism, and factors such as material unevenness and valve nonlinearity make it difficult to establish an accurate mathematical model. This makes it difficult to apply many model-based optimal control methods (such as Linear Quadratic Regulator, LQR) in practice. In recent years, data-driven control methods that do not rely on models have received widespread attention. Such methods can be divided into two categories: indirect and direct. Indirect methods (such as system identification based on least squares method, neural network, etc.) first determine the system model using measurement data, and then design the controller based on the model. However, errors and uncertainties in the model determination step will be passed to the controller, affecting the final performance. Direct methods bypass the modeling step and directly use data to design the controller. Representative methods include reinforcement learning and data-driven behavior system theory.

[0004] Q-learning, as a mainstream model-free reinforcement learning algorithm, is widely used in optimal control problems. For the infinite domain tracking problem, some scholars have proposed a Q-learning algorithm that can solve the augmented algebraic Riccati equation online. For nonlinear systems, some scholars have used neural networks to approximate the Q function. However, these methods all use cost functions with discount factors, and therefore face the risk of stability caused by improper selection of the discount factor.

[0005] To avoid the problem of improper discount factor selection, some scholars have introduced relative control variables and incremental control variables into the cost function and designed a corresponding Q-learning algorithm, but this method relies on the strong assumption that all states of the system are measurable. In actual injection molding systems, due to harsh working conditions such as high temperature, high pressure, and closedness, the cost of accurately measuring all states is extremely high or difficult to achieve, resulting in limited applicability of the above-mentioned method based on full-state feedback. Of course, the existing data-driven control method based on behavior system theory can also avoid the problem of improper discount factor selection, but this method also needs to know all states of the system.

[0006] Therefore, for a multivariable process such as an injection molding system whose states are not completely known, how to design an optimal set point tracking strategy without relying on the discount factor and without determining the system model is still an important technical problem to be solved in the field. SUMMARY

[0007] The present application provides an optimal set point tracking control method for an injection molding system based on output feedback, aiming to achieve optimal set point tracking without relying on the discount factor and the accurate system model. The technical solution is as follows: The present application provides an optimal set point tracking control method for an injection molding system based on output feedback, which comprises: establishing a discrete-time state space equation of the injection molding system, wherein the discrete-time state space equation comprises a state vector, an input vector, and an output vector; subtracting the output vector from the corresponding set point in the reference trajectory to obtain a tracking error, and constructing an augmented system according to the state vector and the tracking error; based on the augmented system, constructing an infinite-domain cost function without discount factor; using a state parameterization method, equivalently representing the augmented state vector of the augmented system as a non-minimal state vector composed of historical input data and output data; based on the non-minimal state vector, defining an output feedback Q function; collecting input data and output data of the injection molding system to design a zero-error controller, and using the zero-error controller as an initial stable control strategy to execute an output feedback off-policy Q-learning algorithm, and converging the output feedback Q function through iteration to obtain an optimal output feedback control strategy that minimizes the infinite-domain cost function.

[0008] In one possible implementation, constructing an augmented system according to the state vector and the tracking error comprises: calculating the difference value of the state vector at adjacent time points; An augmented state vector is defined based on the state vector difference and the tracking error; An augmented system is obtained based on the augmented state vector.

[0009] In a possible implementation, the augmented state vector is: ; wherein, represents the state vector difference, and , represents a state vector at the time point t, represents a state vector at the time point t; represents the tracking error, and , , respectively represent an output vector and a set point at the time point t, , , respectively are a dimension of the state vector and a dimension of the output vector.

[0010] In a possible implementation, the augmented system is: ; wherein, , , represents the augmented state vector, represents an increment of the input vector, 、 、 respectively represent a system matrix, an input matrix and an output matrix of the injection molding system, represents an identity matrix with a dimension of , , , respectively represent a dimension of the vector or matrix; wherein, an input of the augmented system is an increment of the input vector, a state of the augmented system is an augmented state vector composed of a state vector difference and a tracking error , and when the augmented state vector is adjusted to zero, the injection molding system can reach a steady state and optimal set point tracking.

[0011] In a possible implementation, the augmented state vector of the augmented system is equivalently represented as a non-minimal state vector composed of historical input data and output data by using the state parameterization method, including: Based on the state parameterization method, a state parameterization of the augmented state vector is obtained, and the state parameterization of the augmented state vector contains the non-minimal state vector.

[0012] In a possible implementation, the state parameterization of the augmented state vector is: ; Wherein, represents the non-minimal state vector, represents a non-minimal state an increment along a time direction, represents an observable index of the injection molding system, is a row full-rank transformation matrix dependent on model parameters, .

[0013] In a possible implementation, based on the non-minimal state vector, an output feedback Q function is defined, including: determining a value function according to an optimal state feedback control strategy of the augmented system; constructing a state feedback Q function according to the value function; constructing an output feedback Q function based on the state parameterization of the augmented state vector and the state feedback Q function.

[0014] In a possible implementation, the output feedback Q function is constructed based on the state parameterization of the augmented state vector and the state feedback Q function, including: substituting the state parameterization of the augmented state vector into the state feedback Q function to obtain: ; Wherein, is an auxiliary vector of the state feedback Q function, is an auxiliary vector for constructing the output feedback Q function, , represents a unit matrix with a dimension of ; based on the auxiliary vector, the output feedback Q function is: ; Wherein, , , respectively represent parameter matrices of the output feedback Q function and the state feedback Q function, , , respectively represent the block matrices of the output feedback Q function parameter matrix, , , , respectively represent the block matrices of the state feedback Q function parameter matrix, wherein, is the solution of the augmented system algebraic Riccati equation, , , represents a preset weight matrix.

[0015] In a possible implementation, the method further comprises: After obtaining the output feedback Q function, an optimal output feedback control strategy is obtained by minimizing the output feedback Q function.

[0016] In a possible implementation, the output feedback off-policy Q-learning algorithm is executed with the no-lag controller as an initial stable control strategy, and the output feedback Q function is converged by iterative solving to obtain an optimal output feedback control strategy that minimizes the infinite-domain cost function, comprising: processing the input data and the output data of the injection molding system to obtain a nonsingular matrix Z and a matrix W ; multiplying the state feedback Q function by on the left and multiplying by on the right to obtain a to-be-solved equation containing a parameter matrix of the output feedback Q function: ; wherein, , , represents the gain of the output feedback Q function; if the to-be-solved equation needs to be iterated times, the to-be-solved equation is: ; and in the first iteration of solving the to-be-solved equation, let , and select an initial stable control strategy that makes the injection molding system asymptotically stable from an initial state; iteratively solving the to-be-solved equation to constantly update the old strategy to a new control strategy , ; when is true, is a preset constant, the program is ended, and the optimal output feedback control strategy of the injection molding system is output.

[0017] The technical scheme provided by the embodiments of the present application can obtain the following technical effects: (1) On the one hand, the augmented state vector of the present application contains a state vector difference and a tracking error Two components, each of which corresponds to an independent control target: when the component , it indicates that the state vector does not change at the front and rear moments, that is, the injection molding system enters a steady state; when the component , it indicates that the output vector is consistent with the corresponding set point in the reference trajectory, that is, the pressure at the nozzle of the injection molding system reaches the optimal set point tracking. Further, an augmented system is generated based on the augmented state vector, which defines the input of the augmented system as the increment of the input vector and defines the state of the augmented system as the augmented state vector , so that the control target of the injection molding system is no longer to track the reference trajectory, but to adjust the augmented state vector to zero, thereby successfully converting the complex set point tracking problem into a typical adjustment problem. At the same time, since the injection molding system is in a steady state when the augmented state vector is adjusted to zero, the integral term in the cost function eventually decays to zero, thereby ensuring the convergence of the infinite integral and completely getting rid of the dependence on the discount factor.

[0018] (2) On the other hand, based on the state parameterization method, the augmented state vector of the augmented system is first equivalently represented as a non-minimal state vector composed of historical input data and output data, and then an output feedback Q function is constructed based on the non-minimal state vector and an optimal output feedback control strategy is designed, so that the present application can achieve optimal set point tracking without relying on the measurement of all state information of the injection molding system, that is, without knowing all states of the injection molding system.

[0019] (3) In addition, the present application uses the input data and output data generated during the operation of the injection molding system to design a non-ideal controller based on the input data and output data, and uses the controller as the initial stable control strategy of the output feedback off-policy Q learning algorithm, and solves the optimal output feedback control strategy by executing the output feedback off-policy Q learning algorithm. Therefore, the entire solving process does not need to be modeled in advance, but can still guarantee that the optimal output feedback control strategy is solved, and the performance of the full-state feedback control method based on the accurate model is achieved in terms of closed-loop stability, tracking accuracy and anti-interference ability. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical scheme of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. In the drawings: Figure 1 is a flow chart of an optimal set point tracking control method of an output feedback based injection molding system according to an embodiment of the present application; Figure 2 is a schematic diagram of an optimal set point tracking control principle according to an embodiment of the present application; Figure 3 is a flow chart of an output feedback off-policy Q-learning algorithm according to an embodiment of the present application; Figure 4 is an output curve trajectory of an injection molding system in a non-disturbance case according to an embodiment of the present application; Figure 5 is an output curve trajectory of an injection molding system in an unknown constant disturbance case according to an embodiment of the present application; Figure 6 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] Exemplary embodiments of the present application will be described more fully hereinafter with reference to the accompanying drawings; however, they are not intended to limit the present application to particular embodiments. Rather, the present application is intended to cover all modifications, equivalents and alternatives falling within the scope of the present application. Like numbers refer to like elements throughout the description of the figures.

[0022] It is to be understood that the terms "first", "second", and the like, used herein do not necessarily connote any order, quantity, or importance, but are used to distinguish one element from another, and do not necessarily indicate a required or particular sequence. It is to be understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0023] The present application provides an optimal set point tracking control method of an output feedback based injection molding system, as shown in FIG. 1, the method mainly includes the following steps S101-S106. Figure 1

[0024] Step S101, a discrete-time state space equation of the injection molding system is established, the discrete-time state space equation includes a state vector, an input vector and an output vector.

[0025] ​The injection molding system generally refers to a pressure maintaining process control system with hydraulic control valve opening degree as control input and nozzle pressure as controlled output. In the embodiment, the system specifically refers to the dynamic process from the control input to the process output. The control of the system faces two big problems: the internal state cannot be directly measured and the complex mechanism leads to difficulty in establishing an accurate mathematical model. To overcome the above problems, the embodiment discards the traditional accurate modeling idea and instead uses a discrete-time state-space equation as a theoretical framework for describing the dynamic characteristics of the system, thereby providing data support for subsequent data-driven control methods that do not rely on model parameters at all.

[0026] Specifically, the calculation formula of the discrete-time state-space equation is: (1) wherein subscript represents time, is a set of natural numbers, , , are state vector, input vector and output vector of the injection molding system respectively, is the initial state of the injection molding system, represents a real number space of dimension, , , are the dimensions of the state vector, input vector and output vector respectively, 、 、 is a coefficient matrix with corresponding dimensions, specifically, 、 、 respectively represent the system matrix, input matrix and output matrix of the injection molding system.

[0027] The state vector described above refers to a set of variables that describe the internal dynamics of the injection molding system, which are usually not directly measurable, such as melt pressure, temperature, etc. The input vector refers to the opening degree of the hydraulic control valve. The output vector refers to the physical quantity that can be directly measured by the sensor and reflects the key performance of the injection molding system, which in the embodiment specifically refers to the pressure at the nozzle. The state vector, input vector and output vector together form a dynamic causal chain of the injection molding system: by changing the input vector (opening degree of the hydraulic control valve) to drive the injection molding system, the state vector (such as melt pressure, temperature) of the system changes, and this internal change is ultimately perceived as a change in the output vector (pressure at the nozzle).

[0028] Step S102: Subtract the output vector from the corresponding set point in the reference trajectory to obtain the tracking error, and construct an augmented system based on the state vector and the tracking error.

[0029] First, subtract the output vector from the corresponding setpoint in the reference trajectory to obtain the tracking error: (2) in, , They represent in The output vector and setpoints at each moment, with multiple setpoints arranged in chronological order, form a reference trajectory. For example... Figure 2 As shown, to achieve optimal setpoint tracking in this embodiment, it is necessary to ensure that the output vector is consistent with the corresponding setpoint in the reference trajectory. That is, the pressure at the nozzle of the injection molding system is the same as the pre-set pressure value. After a period of time, the curve formed by the pressure at the nozzle is the same as the reference trajectory and covers the reference trajectory, thereby achieving optimal setpoint tracking of the pressure at the nozzle.

[0030] Then, calculate the state vector difference between adjacent time steps: (3) in, express The state vector at time t, express The state vector at time step, therefore It can also be called a state increment.

[0031] Based on the obtained state vector difference and tracking error Define the augmented state vector: (4) It should be noted that all superscript "T" appearing in this application for any parameter indicates the transpose symbol.

[0032] The augmented state vector described above is derived from the state vector difference. and tracking error A new vector is formed by these components, defined such that each of the two components corresponds to an independent control objective: when the components... When the state vector no longer changes at the preceding and following moments, it indicates that the injection molding system has entered a steady state; when the components When the output vector is adjusted to zero, it indicates that the output vector is consistent with the corresponding setpoint in the reference trajectory, meaning that the pressure at the nozzle of the injection molding system has reached the optimal setpoint tracking. Therefore, when the augmented state vector is adjusted to zero, it indicates that the injection molding system has reached a steady state and simultaneously achieved optimal setpoint tracking.

[0033] Based on the augmented state vector, the augmented system is obtained: (5) wherein, , , denotes a unit matrix with dimension .

[0034] From the above formula (5), the augmented system defines its input as the increment of the input vector , and defines its state as the augmented state vector composed of the state vector difference and the tracking error , so that the control objective of the injection molding system is no longer to track the reference trajectory, but to adjust the augmented state vector to zero, because when the augmented state vector is adjusted to zero, not only can the injection molding system reach a steady state, but also the optimal set point tracking is achieved, so the embodiment can successfully convert the complex set point tracking problem into a typical regulation problem.

[0035] Step S103, based on the augmented system, construct an infinite-domain cost function without discount factor.

[0036] Based on the above augmented system, the infinite-domain optimization problem can also be fundamentally solved, because the regulation target (the augmented state vector adjusted to zero, i.e. ) of the embodiment naturally reaches a steady state, resulting in the integral term in the cost function eventually decaying to zero, thereby ensuring the convergence of infinite integration and completely getting rid of the dependence on the discount factor.

[0037] In the embodiment, the calculation formula of the infinite-domain cost function is: (6) wherein, , , is a weight matrix set in advance.

[0038] The infinite-domain cost function comprehensively considers all tracking performance and control cost from the current time to the future through the integral form, and its minimum value corresponds to the theoretically globally optimal control performance. The function directly determines the optimization direction of the output feedback Q function used subsequently, guiding the Q function to find the control strategy that can minimize the long-term cumulative cost, thereby ensuring that the closed-loop system operates stably while achieving accurate tracking of the set point.

[0039] Step S104, using the state parameterization method, the augmented state vector of the augmented system is equivalent to a non-minimal state vector composed of historical input data and output data.

[0040] From the state parameterization method, the system state can be equivalent to the historical input data and output data, and the preset parameter matrix, that is: (7) Wherein, is a defined non-minimal state, represents a vector stacked by the input of the previous time, can be obtained by calculating the observability index, is a transformation matrix calculated according to the input and output data of the system, is a row full rank transformation matrix depending on the model parameters, and the specific parameters satisfy: (8) Wherein, is an invertible transformation matrix, contains linearly independent rows of , contains the remaining rows, and and are constructed in the same way as and , can be obtained by calculating the inverse matrix of , .

[0041] Based on the above state parameterization method, the state parameterization of the augmented state vector of the embodiment can be expressed as: (9) Wherein, is the non-minimal state vector of the augmented system, represents the increment of the non-minimal state in the time direction, represents the observability index of the injection molding system.

[0042] Since the state vector of the injection molding system cannot be directly measured, and the non-minimal state vector is a vector stacked by the input and output data of the injection molding system in the past period of time, the non-minimal state vector is used, and the equivalent relationship between the non-minimal state vector and the theoretical augmented state vector is established by the state parameterization method, thereby laying a foundation for subsequent optimal output feedback control strategy design based on input data and output data.

[0043] Step S105, defining the output feedback Q function based on the non-minimal state vector.

[0044] Firstly, the state feedback Q function is constructed.

[0045] According to the standard linear quadratic regulation theory, the optimal state feedback control strategy of the augmented system is: (10) For any feedback gain , the value function at the given time is quadratic, specifically: (11) where is the value function parameter of the feedback gain .

[0046] According to the Bellman optimality principle, the above value function (11) can be written in the following recursive form: (12) Combining formulas (11) and (12), the state feedback Q function is constructed: (13) where the vector is the auxiliary vector of the state feedback Q function, and the matrix is the parameter matrix of the state feedback Q function and , , , , respectively represent the block matrices of the state feedback Q function parameter matrix, is the solution of the Riccati equation of the augmented system.

[0047] Specifically, the optimal state feedback control strategy (10) of the augmented system can be obtained by minimizing the state feedback Q function (13), that is: (14) Due to the equivalence of formula (12) and formula (13), the state feedback Q function can be further written in the following recursive form: (15) Further, formula (15) can be written in the following simplified form: (16) Substituting the control strategy Equation (16) can also be equivalently described as a matrix equation as follows: (17) where, , .

[0048] From equation (14), the optimal state feedback gain can be obtained by solving the matrix in equation (17).

[0049] The state feedback Q-function described above is a function that measures the long-term total cost of the injection molding system after taking a certain action in a certain state, for example, when preparing to execute a certain control action (including increasing, decreasing, or keeping unchanged the input vector) based on the current augmented state vector of the injection molding system, the function will correspondingly give a score, which represents the total sum of all future costs accumulated after starting from the current augmented state vector and executing the optimal control action. That is, the state feedback Q-function not only considers the immediate cost of the current control action, but more importantly, it inherently predicts and includes all future impacts of the control action. The optimal state feedback control strategy is obtained by minimizing the state feedback Q-function, that is, the optimal state feedback control strategy represents what should be the optimal control action in the process of selecting the control action at any time.

[0050] Secondly, the output feedback Q-function is constructed. Specifically, the auxiliary vector for constructing the output feedback Q-function is first parameterized by the state of the augmented state vector and the state feedback Q-function, that is, by substituting the state parameterization of the augmented state vector into the calculation formula (13) of the state feedback Q-function, the following is obtained: (18) where, is the auxiliary vector for constructing the output feedback Q-function, , represents a unit matrix with a dimension of .

[0051] Based on the constructed auxiliary vector, the output feedback Q-function is obtained as follows: (19) where, , represents the parameter matrix of the output feedback Q-function, , , represent the block matrices of the output feedback Q-function parameter matrix, respectively.

[0052] The output feedback Q function described above is a Q function whose input is entirely composed of measurable data, which can evaluate and find the optimal control action (also referred to as the optimal control strategy) only with input and output data without relying on the measurement of any state vector, so after obtaining the output feedback Q function in this embodiment, the optimal output feedback control strategy is obtained by minimizing the output feedback Q function.

[0053] Specifically, the process of obtaining the optimal output feedback control strategy is as follows: First, the state parameter of the augmented state vector is substituted into the optimal state feedback control strategy to obtain the initial optimal output feedback control strategy as follows: (20) wherein , .

[0054] Based on the initial optimal output feedback control strategy, the final optimal output feedback control strategy is obtained by minimizing the output feedback Q function as follows: (21) wherein the final optimal output feedback control strategy of the injection molding system in the form of proportional integral is obtained by summing the left and right sides of the initial optimal output feedback control strategy from 0 to (22) Step S106, collect the input data and output data of the injection molding system, design a deadbeat controller, and execute the output feedback off-policy Q learning algorithm with the deadbeat controller as the initial stable control strategy, and converge the output feedback Q function by iterative solution to obtain the optimal output feedback control strategy that minimizes the infinite domain cost function.

[0055] First, a deadbeat controller is designed based on offline data as an initial stable control strategy. For the input signal of the injection molding system with an application order of , collect the input vector, output vector and tracking error generated during the operation of the injection molding system to obtain a data sequence with a length of N , defined as the following matrix: (23) Let , construct the following non-singular matrix and matrix : , (24) Multiply the left side of formula (17) by and multiply the right side by ​This yields a parameter matrix that includes the output feedback Q function. The equation to be solved: (25) in, , , This represents the gain of the output feedback Q function.

[0056] Based on the data matrix in the above formula (23) , and The deadbeat controller is designed through the following steps a1 to a5: Step a1: Through the data matrix , and Define the matrix of the virtual system and ; Step a2: Let for Moore-Penrose pseudo-reverse, for The basis of the null space is defined as follows: the matrix of the virtual system is... , ; Step a3: Definition For a set of linearly independent matrices The full column rank matrix formed by the columns of is such that ; Step a4: Use the transformation matrix Virtual system Transform into the canonical form of a multiple-input multiple-output controller. ; Step a5: For the system Design a time-lapse controller , making It is zero-power.

[0057] Based on the obtained deadbeat controller By calculating the matrix and in the matrix Add zero rows to construct a new matrix. The position where the zero line is added is the same as The middle is formed The columns that are removed are in the same position, making Then, through calculation of the deadbeat controller... Gain , recorded as In order to gain this The deadbeat controller This serves as the initial stable control strategy for the subsequent output feedback policy Q-learning algorithm.

[0058] The output feedback Q-learning algorithm is a reinforcement learning method that iteratively solves for the optimal output feedback control strategy using the system's operational data without requiring prior modeling of the injection molding system. The core principle of this method is that through continuous iteration of policy evaluation and improvement, the parameters of the output feedback Q-function converge, thereby obtaining an optimal output feedback control strategy that minimizes the infinite-domain cost function.

[0059] like Figure 3 As shown, the process of executing the output feedback Q-learning algorithm in this embodiment is as follows: steps b1 to b6: Step b1: Collect the input and output data of the injection molding system, and process the collected data according to formulas (23) and (24) to obtain a non-singular matrix. sum matrix .

[0060] Step b2: Initialize the injection molding system, i.e., solve the equation to be solved in the iterative solution formula (25). (Assuming iteration) Next, then When ), first order At the same time, an initial stability control strategy is selected to make the injection molding system asymptotically stable from the initial state.

[0061] Step b3: Iterative solution Because of the parameter matrix in the equation to be solved Derived from the output feedback Q function, therefore, for the first... The control strategy used in the next iteration yields a corresponding parameter matrix. , the parameter matrix Substituting the values ​​into the output feedback Q function, we obtain a score for the output feedback Q function, which is used to evaluate the first... In the next iteration, is there room for improvement in the control strategy used, so as to continuously optimize the selected control strategy?

[0062] Step b4: In the iterative solution of the equation to be solved in step b3, the Q-function outputs a score in each iteration. Based on this score, the selected control strategy is further optimized, i.e., the old strategy is used. It can drive the output feedback Q function to output a score, and this score can then be used to generate a new control strategy. .

[0063] Step b5: In the process of constantly optimizing the selected control strategy (iterative output feedback Q function), if is established, and is a very small constant, the program ends. That is, the score output by the output feedback Q function is constantly selected until the selected control strategy is basically unchanged, indicating that the selected control strategy is the optimal control strategy (also referred to as the optimal output feedback control strategy).

[0064] Step b6: Output the optimal output feedback control strategy of the injection molding system: (26) In this embodiment, the infinite-domain cost function is a standard for measuring the long-term performance of the injection molding system, but it is difficult to directly optimize this function, and the output feedback Q function is actually another form of this cost function, similar to a "real-time scoring system". Therefore, the output feedback off-policy Q learning algorithm in this embodiment drives the output feedback Q function to converge, and its essence is to make the Q function gradually approach and eventually satisfy its corresponding Bellman optimality equation. In mathematics, the solution of this equation and the global minimum solution of the infinite-domain cost function are mutually necessary and sufficient conditions. Therefore, when the algorithm converges, the optimal output feedback control strategy derived from the output feedback Q function is the global optimal solution under the definition of the infinite-domain cost function, thereby ensuring that the injection molding system meets both stability and optimal set point tracking.

[0065] As can be seen, in this embodiment, an initial error-free controller that guarantees the stability of the injection molding system is first designed using offline data in a completely data-driven manner, and then based on this, the optimal output feedback control strategy consistent with the global optimal solution of the infinite-domain cost function is iteratively solved by executing the output feedback off-policy Q learning algorithm, without any prior knowledge about the system model throughout the process.

[0066] In order to facilitate the explanation of the effectiveness of the output feedback-based optimal set point tracking control method of the injection molding system of the present application, a specific injection molding system is taken as an example in the following: First, based on open-loop testing and analysis, the change of the pressure at the nozzle with the opening of the hydraulic control valve is determined as the following discrete-time state space equation: (27) where the input vector is the opening of the hydraulic control valve, the output vector is the pressure at the nozzle, the initial state vector is set to , , , , the weight matrix is , . The set point is , , is an integer set. The random input signal satisfying the persistent excitation condition is applied to the above formula (27) and data is collected, the collected offline input and output data are preprocessed, and the incremental input data and the incremental output data with a sequence length of are obtained and , and the tracking error is . Let , construct the nonsingular matrix and the matrix , and execute the output feedback off-policy Q-learning algorithm. During the execution process, after 10 iterations of solving, the gain of the obtained zero-error controller is . The numerical result is exactly the same as the optimal solution obtained by using the MATLAB idare function based on the accurate model, which verifies the effectiveness of the algorithm.

[0067] For the injection molding system in the above example, when the system does not exist disturbance, the output curve is as shown in Figure 4 , when the system is added with disturbance (r(t)≠0 and d(t)≠0), the output curve of the injection molding system is as shown in . According to the results of and Figure 5 , it can be seen that the optimal set point tracking control method of the injection molding system based on output feedback of the embodiment can realize the optimal tracking of the set point of the reference trajectory, and in addition, the method has good inhibition effect on unknown constant disturbance. Figure 4 Figure 5 Further, for the injection molding system in the above example, when the system does not exist disturbance, the control method of the application is compared with the tracking performance (integral absolute error and mean square error) of the optimal strategy based on the model and the strategy with discount factor, and the comparison result is shown in Table 1.

[0068] Table 1

[0069]

[0070] From Table 1, it can be seen that the control method of the application realizes almost the same control precision as the optimal strategy based on the model, and the control effect of the strategy with discount factor is the worst. In addition, the control method of the application is completely based on data, while the model-based solution needs accurate known model information.

[0071] In addition, the control method of the application is compared with other data-driven methods in solving the optimal strategy performance, and based on 100 independent experiments and recorded experimental data, the statistical results (average error and average running time) are calculated and shown in Table 2.​​

[0072] Table 2

[0073] As can be seen from Table 2, the control method of the application is superior to the reinforcement learning method and the semi-definite programming method in terms of calculation accuracy and efficiency.

[0074] The above results show that the optimal set point tracking control method of the injection molding system based on output feedback of the application can achieve optimal steady-state tracking of the reference trajectory under unknown system model information and all state measurement information. In addition, the control strategy has a good inhibitory effect on unknown constant disturbances.

[0075] It should be noted that the size of the serial number of each step in the above embodiments does not mean the order of execution. The execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application. In actual application, all possible implementation manners described above can be combined in any combination to form possible embodiments of the application, which will not be described one by one here.

[0076] Based on the same inventive concept, the embodiments of the application also provide an electronic device including a processor and a memory, the memory storing a computer program, and the processor being configured to run the computer program to execute the optimal set point tracking control method of the injection molding system based on output feedback of any one of the above embodiments.

[0077] In an exemplary embodiment, an electronic device is provided, such as Figure 6 As shown in FIG. 6, Figure 6 The electronic device 600 shown in FIG. 6 includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, such as through a bus 602. Optionally, the electronic device 600 can also include a transceiver 604. It should be noted that in actual application, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation on the embodiments of the application.

[0078] The processor 601 can be a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure. The processor 601 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0079] The bus 602 can include a path that transmits information between the above-mentioned components. The bus 602 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 602 can be divided into an address bus, a data bus, a control bus, and the like. For convenience of representation, Figure 6 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0080] The memory 603 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, and the like), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0081] The memory 603 is configured to store computer program codes for implementing the solutions of the present application, and the processor 601 is configured to execute the computer program codes stored in the memory 603. The processor 601 is configured to execute the computer program codes stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0082] The electronic device includes, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 6 The electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0083] Based on the same inventive concept, the embodiments of the present application further provide a storage medium in which a computer program is stored, wherein the computer program is configured to execute the output feedback-based injection molding system optimal set point tracking control method of any one of the foregoing embodiments when running.

[0084] Those skilled in the art can clearly understand the specific working process of the system, device, and module described above, and can refer to the corresponding process in the foregoing method embodiments. For the sake of brevity, no further description is given here.

[0085] Those skilled in the art can understand that the technical solutions of the present application can be embodied in the form of a software product in essence or in whole or part of the technical solutions, and the computer software product is stored in a storage medium, and includes a plurality of program instructions for causing an electronic device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application when the program instructions are run. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0086] Alternatively, all or part of the steps of the foregoing method embodiments can be completed by program instruction-related hardware (such as an electronic device of a personal computer, a server, or a network device, etc.), and the program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the method described in the embodiments of the present application.

[0087] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application is described in detail with reference to the above embodiments, those skilled in the art should understand that, within the spirit and principle of the present application, the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of the present application.

Claims

1. An output feedback based optimal set point tracking control method for injection molding systems, characterized in that, The method comprises: establishing a discrete-time state-space equation of an injection molding system, the discrete-time state-space equation comprising a state vector, an input vector, and an output vector; subtracting the output vector from a corresponding set point in a reference trajectory to obtain a tracking error, and constructing an augmented system according to the state vector and the tracking error; constructing an infinite-domain cost function without a discount factor based on the augmented system; equivalently representing an augmented state vector of the augmented system as a non-minimal state vector composed of historical input data and output data by using a state parameterization method; defining an output feedback Q function based on the non-minimal state vector; collecting input data and output data of the injection molding system to design a zero-error controller, and executing an output feedback off-policy Q-learning algorithm with the zero-error controller as an initial stable control strategy, and converging the output feedback Q function by iterative solving to obtain an optimal output feedback control strategy that minimizes the infinite-domain cost function.

2. The output feedback based injection molding system optimal set point tracking control method of claim 1, wherein, The method further comprises: constructing an augmented system according to the state vector and the tracking error, comprising: calculating a difference value of the state vector at adjacent time points; defining an augmented state vector based on the difference value of the state vector and the tracking error; 3. The output feedback based injection molding system optimal set point tracking control method of claim 2, wherein, obtaining the augmented system based on the augmented state vector. ; wherein denotes the state vector difference, and , denotes the state vector at time instant denotes the state vector at time instant denotes the tracking error, and , , denote the output vector and the setpoint at time instant , , are the dimensions of the state vector and the output vector, respectively.

4. The output feedback based injection molding system optimal set point tracking control method of claim 3, wherein, The augmented state vector is: ; wherein , , denotes the augmented state vector, denotes the increment of the input vector, 、 、 denote the system matrix, the input matrix and the output matrix of the injection molding system, respectively, denotes the identity matrix of dimension , , , denote the dimension of the vector or matrix, respectively; wherein the input of the augmented system is the increment of the input vector , the state of the augmented system is an augmented state vector composed of the state vector difference and the tracking error , and the augmented system is adjusted to zero when the augmented state vector is adjusted to zero, the injection molding system is able to reach a steady state and optimal set point tracking.

5. The output feedback based injection molding system optimal set point tracking control method according to any one of claims 2-4, characterized in that, The augmented system is: equivalently representing the augmented state vector of the augmented system as a non-minimal state vector composed of historical input data and output data by using the state parameterization method, comprising:

6. The output feedback based injection molding system optimal set point tracking control method of claim 5, wherein, obtaining a state parameterization of the augmented state vector based on the state parameterization method, the state parameterization of the augmented state vector containing the non-minimal state vector. ; wherein denotes the non-minimal state vector, denotes a non-minimal state an increment in the time direction, denotes an observable indicator of the injection molding system, is a model parameter dependent row full rank transformation matrix, .

7. The output feedback based injection molding system optimal set point tracking control method of claim 5, wherein, The state parameterization of the augmented state vector is: defining an output feedback Q function based on the non-minimal state vector, comprising: determining a value function according to an optimal state feedback control strategy of the augmented system; constructing a state feedback Q function according to the value function; 8. The output feedback based injection molding system optimal set point tracking control method of claim 7, wherein, constructing an output feedback Q function based on the state parameterization of the augmented state vector and the state feedback Q function. The method further comprises: ; wherein is an auxiliary vector of the state feedback Q-function, is an auxiliary vector for constructing the output feedback Q-function, , denotes an identity matrix of dimension . constructing an output feedback Q function based on the state parameterization of the augmented state vector and the state feedback Q function, comprising: ; wherein , , denote the parameter matrices of the output feedback Q-function and the state feedback Q-function, respectively, , , denote the block matrices of the output feedback Q-function parameter matrix, , , , denote the block matrices of the state feedback Q-function parameter matrix, wherein is the solution of the augmented system algebraic Riccati equation, , , denotes a weight matrix set in advance.

9. The output feedback based injection molding system optimal set point tracking control method of claim 8, wherein, substituting the state parameterization of the augmented state vector into the state feedback Q function to obtain: obtaining the output feedback Q function based on the auxiliary vector as:

10. The output feedback based injection molding system optimal set point tracking control method of claim 1, wherein, The method further comprises: obtaining an optimal output feedback control strategy by minimizing the output feedback Q function after obtaining the output feedback Q function. executing an output feedback off-policy Q-learning algorithm with the zero-error controller as an initial stable control strategy, and converging the output feedback Q function by iterative solving to obtain an optimal output feedback control strategy that minimizes the infinite-domain cost function, comprising: processing input data and output data of the injection molding system to obtain a non-singular matrix Z and a matrix W ; Multiplying the state feedback Q function on the left by and on the right by yields the equation to be solved for the parameter matrix of the output feedback Q function: ; where , , denotes the gain of the output feedback Q function; if the equation to be solved needs iteration then the equation to be solved is: ; and in the first iteration to solve the equation to be solved, let and select an initial stable control strategy that makes the injection molding system asymptotically stable from an initial state; iteratively solving the equations to be solved , constantly updating the old policy to a new control policy , ; When is established, is a preset constant, the execution of the program is ended and the optimal output feedback control strategy of the injection molding system is output.

Citation Information

Patent Citations

  • Baxter mechanical arm trajectory tracking control method based on reinforcement learning

    CN113199477A

  • Engineering system fault-tolerant tracking control method based on data driving

    CN119847108A

  • Electromagnetic micromirror output feedback tracking control method and device based on adaptive Q learning

    CN120491471A

  • Adaptive dynamic programming method for aero-engine in optimal acceleration tracking control

    WO2021097696A1