Large Inertia Dynamic Flow Balance and Control Method for Molding Equipment Based on Knowledge Graph

By combining knowledge graph and reinforcement learning technology, the knowledge graph model and reinforcement learning model are constructed, and the problem of control accuracy and stability of industrial control systems in large inertia dynamic flow balance and precise control scenarios is solved, achieving high flexibility and stability of the system.

CN119717542BActive Publication Date: 2025-06-27GUIZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510196138.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

When existing industrial control systems face large inertia dynamic flow balance and precise control scenarios, the flow control accuracy is low and the control is unstable under large inertia loads.

Method used

Combining knowledge graph and reinforcement learning technology, a method is designed to build a knowledge graph model and reinforcement learning model by collecting and preprocessing the real-time running data of equipment, using knowledge graphs to provide initial training and real-time decision support, and optimize the control strategy of reinforcement learning model.

Benefits of technology

It realizes dynamic optimization of control signals under complex and variable operating conditions, and quickly adaptively adjusts control strategies, which significantly improves the flexibility and stability of the system and improves the accuracy and adaptability of molding equipment operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119717542B_ABST
    Figure CN119717542B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of industrial automation and control technology, and particularly relates to a large inertia dynamic flow balance and control method for forming equipment based on a knowledge graph. First, real-time sensor data and historical operation records are collected, and the moving average method is used for denoising and normalization processing; then a complex association model based on a multi-level knowledge graph is constructed to support intelligent reasoning and optimization decision-making; subsequently, a reinforcement learning model is constructed to define the state space, action space, and reward function, and the training efficiency is improved by combining the experience replay and target network mechanisms; finally, according to the real-time state of the equipment, the system generates an optimal control signal, and the control parameters are dynamically adjusted by means of the knowledge graph reasoning mechanism to ensure that the equipment works in the best operating state. The present invention can solve the problems of low flow control accuracy and unstable control under large inertia loads in the existing industrial control system when facing large inertia dynamic flow balance and precise control scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial automation and control, and particularly relates to a method for large inertia dynamic flow balance and control of forming equipment based on a knowledge graph. Background Art

[0002] In traditional industrial control systems, the regulation of parameters such as flow rate and pressure usually relies on classical control algorithms, such as PID control and fuzzy control. These methods work well in simple linear systems, but when faced with complex non-linear systems or large inertia loads, there are often significant limitations in control accuracy and response speed. Traditional control algorithms rely on preset rules and fixed parameters, lacking adaptability and being difficult to cope with the dynamic changes of equipment states. Especially in complex industrial environments, these algorithms often cannot quickly adapt to changes in equipment loads or working conditions, resulting in reduced efficiency or even equipment out of control. In addition, traditional methods still require manual intervention and lack intelligent decision-making and optimization capabilities, which is a significant bottleneck for modern automated systems.

[0003] In recent years, reinforcement learning, as a technology based on continuously optimizing decisions through interaction with the environment, has gradually been applied to control systems. Different from traditional control methods, reinforcement learning can dynamically adjust control strategies according to real-time feedback, enabling the system to have adaptability and automatically adjust operations according to equipment states and environmental changes. Through a reward mechanism, reinforcement learning guides the agent to find the optimal control actions, thereby improving system efficiency and stability. This technology is particularly suitable for dynamic and complex environments and helps to continuously improve control decisions.

[0004] Although reinforcement learning has made some progress in control systems, when faced with large-scale and complex systems, relying solely on its exploration mechanism still poses challenges. Especially during the training process, the trial-and-error cost is high and the learning progress is slow. To solve this problem, the research combining knowledge graphs and reinforcement learning has received attention. Knowledge graphs integrate a large amount of heterogeneous data through a graph structure, providing the relationships and rules between equipment, sensors, and control strategies, thereby reducing unnecessary exploration, accelerating the learning process, and improving decision-making accuracy. In control systems, knowledge graphs can not only store the relationships between equipment and control signals, but also provide real-time decision support for reinforcement learning models through an inference mechanism to optimize control strategies.

[0005] Focusing on the scenarios of large inertia dynamic flow balance and precise control, such scenarios widely exist in many industrial fields such as heavy machinery manufacturing and large-scale chemical production. In these fields, the equipment has a huge moment of inertia, and the requirements for dynamic balance regulation of the flow are extremely high. Even a slight flow deviation may cause product quality problems or even equipment failures. Traditional control methods are unable to cope with such stringent requirements. Although reinforcement learning has potential, it is limited by its own defects and is difficult to take on the heavy responsibility. Although the research on the combination of knowledge graph and reinforcement learning has started, there are still many technical difficulties in effectively integrating the two and applying them to such complex scenarios.

[0006] Based on this, this solution tightly combines reinforcement learning and knowledge graph reasoning, and designs a new method to ensure that the system can still dynamically optimize the control signal even under complex and changing working conditions, and at the same time quickly and adaptively adjust the control strategy according to the real-time state of the equipment, comprehensively improving the flexibility and stability of the system and filling the current technical gap. Summary of the Invention

[0007] The technical problem solved by the present invention is to provide a method for large inertia dynamic flow balance and control of a forming equipment based on a knowledge graph, so as to solve the problems of low flow control accuracy and unstable control under large inertia load in the existing industrial control system when facing the scenarios of large inertia dynamic flow balance and precise control.

[0008] The basic solution provided by the present invention: A method for large inertia dynamic flow balance and control of a forming equipment based on a knowledge graph, including:

[0009] S1: Collect the real-time operation data and historical operation data of the forming equipment, and preprocess the real-time operation data to generate preprocessed real-time operation data;

[0010] S2: Based on the preprocessed real-time operation data and historical operation data, call the database to construct a knowledge graph model based on entity nodes and relationships;

[0011] S3: Construct a reinforcement learning model. In the reinforcement learning model, define and construct the state space with the real-time operation data of the forming equipment, define and construct the action space with the control signal of the forming equipment, and define and construct the reward function with the preset performance standard of the forming equipment;

[0012] S4: Interact the reinforcement learning model with the scenarios of large inertia dynamic flow balance and precise control of the forming equipment, initialize the state according to the state space, select actions according to the current state of the forming equipment, feedback and optimize the strategy through the reward function, and provide control strategy guidance for the initial stage of the training of the reinforcement learning model through the knowledge graph model to train the reinforcement learning model to generate the optimal control signal and obtain the trained reinforcement learning model;

[0013] S5: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the forming equipment. The control system uses the knowledge graph model to infer the control strategy of the reinforcement learning model, and calls the reinforcement learning model to generate an action control signal according to the inferred control strategy to control the operating state of the forming equipment. The reinforcement learning model then continuously optimizes the control strategy of the forming equipment based on the feedback from the forming equipment.

[0014] The principle and advantages of the present invention are as follows: In the technical solution of this application, first collect the real-time operation data and historical operation data of the forming equipment. Among them, preprocessing the real-time operation data helps to provide high-quality data support for subsequent model input. Subsequently, based on the preprocessed real-time operation data, historical operation data, and the NE04J database, construct a knowledge graph model. At this time, the construction of the knowledge graph model provides efficient storage and query of data, and provides a basis for subsequent intelligent reasoning and optimization decision-making.

[0015] In the subsequent control strategy design, the reinforcement learning technology is used to generate a reinforcement learning model. The state space, action space, and reward function are defined, and the knowledge graph model is used to provide control strategy guidance in the initial training stage of the reinforcement learning model, solving the low-efficiency problem of traditional reinforcement learning relying on a large number of trials and errors, thereby accelerating convergence to the optimal control strategy and quickly obtaining a trained reinforcement learning model.

[0016] Finally, when deploying the model to the control system of the forming equipment, deeply combine the reinforcement learning model with knowledge graph reasoning, so that the control system still has high flexibility and intelligent capabilities in the scenarios of large inertia dynamic flow balance and precise control of the forming equipment. On the one hand, it can optimize the control signal according to the real-time state of the equipment, and on the other hand, it can continuously optimize the control strategy through feedback, thereby significantly improving the accuracy and adaptability of the operation of the forming equipment.

[0017] Further, the S1 includes:

[0018] S1-1: Collect the real-time operation data and historical operation data of the forming equipment;

[0019] S1-2: Use the moving average operation to denoise and smooth the collected real-time operation data. The expression is:

[0020]

[0021] Among them, represents the moving average value at time , represents the window size, represents the th data point of the time series data, represents the current time;

[0022] S1-3: Scale the denoised and smoothed real-time operation data to a specific interval through normalization processing. The expression is:

[0023]

[0024] where represents the normalized data point, represents a data point in the real-time operation data, represents the minimum value in the real-time operation data, represents the maximum value in the real-time operation data.

[0025] Beneficial effects: The real-time operation data collected from the forming equipment may contain noise, i.e., random errors or meaningless fluctuations. These noises originate from hardware failures, environmental changes, or other factors, and may cause the equipment control system to make misleading decisions, thus affecting the training and inference results of the model. Therefore, using the moving average operation for denoising and smoothing can effectively reduce the short-term fluctuations in the data, capture the long-term trends, and ensure that the control signals of the equipment are more stable; the normalization operation aims at the possible differences in the measurement units and magnitudes of the data sources, which may cause some features to dominate in the model training. Therefore, the differences are eliminated through the normalization operation.

[0026] Furthermore, the S2 includes:

[0027] S2-1: Extract the historical operation parameters, standard production process parameters, case experiences in the actual production process, physical laws, control theories, and manufacturing site processing records and experience notes from the preprocessed real-time operation data and historical operation data to obtain data features;

[0028] S2-2: Invoke the NEO4J database, convert the data features into a graph structure, and construct a knowledge graph model based on entity nodes and relationships; among them, the entity nodes represent key entities, and the relationships are represented by different types of edges to indicate the mutual relationships between entities.

[0029] Beneficial effects: By constructing a knowledge graph model and leveraging the graph structure, the associations between the equipment status, control parameters, and historical data can be efficiently stored and queried, supporting subsequent reasoning and decision-making.

[0030] Furthermore, in the S2-2, the entity nodes include equipment nodes, sensor nodes, control signal nodes, and historical data nodes;

[0031] The relationships in the S2-2 are represented by different types of edges to indicate the mutual relationships between entities specifically as follows:

[0032] The equipment node and the sensor node are connected through the HAS relationship;

[0033] The sensor node and the control signal node are connected through the INFLUENCES relationship;

[0034] The control signal node and the historical data node are connected through the BASED_ON relationship.

[0035] Beneficial effects: By representing the mutual relationships between entities with edges, the graph can comprehensively and accurately represent the complex relationships among equipment, sensors, control signals, and historical data, facilitating quick query and reasoning in subsequent reasoning processes, automatically selecting appropriate control signals, and thus optimizing the control strategy.

[0036] Further, the S3 includes:

[0037] S3-1: Construct a reinforcement learning model based on the Deep Q-Learning algorithm;

[0038] S3-2: Determine the state space, and define and construct the state space with the real-time operation data of the molding equipment. The expression is:

[0039]

[0040] where represents the state at time , , and respectively represent the parameters of the real-time operation of the equipment;

[0041] S3-3: Determine the action space, and define and construct the action space with the control signals of the molding equipment. The expression is:

[0042]

[0043] where represents the action at time , , , , respectively represent the control signal parameters;

[0044] S3-4: Determine the reward function, and define and construct the reward function with the preset performance standard of the molding equipment. The reward function includes positive rewards and negative rewards;

[0045] S3-5: Invoke the Deep Q-Learning algorithm to approximate the Q function, and process the state space and the action space through a deep neural network to continuously optimize the Q value to generate the optimal control strategy for the equipment. The expression is:

[0046]

[0047] Among them, represents the value of performing the action in the state , the value of represents the immediate reward, represents the discount factor, represents the target network parameters.

[0048] Beneficial effects: By constructing multi-dimensional states in the state space, the reinforcement learning model can comprehensively capture the dynamic characteristics of the equipment operation and make corresponding control decisions based on the current state; the definition of the action space can ensure that the reinforcement learning model can select the optimal control strategy under different working environments; positive and negative rewards are set in the reward function, enabling the reinforcement learning model to adjust the control strategy according to real-time feedback and optimize the control effect of the equipment; the Deep Q-Learning algorithm approximates the Q function, enabling the reinforcement learning model to continuously adjust the control signal to adapt to the dynamic changes of the equipment and ensure the optimal operation of the model.

[0049] Furthermore, the S4 includes:

[0050] S4-1: Initially train the reinforcement learning model through the historical operation data and control strategies in the knowledge graph;

[0051] S4-2: Interact the initially trained reinforcement learning model in real time with the control system of the large inertia dynamic flow balance and precise control scenario of the formed equipment. During the real-time interaction process, perform state initialization according to the state space, select actions according to the current state of the formed equipment, and optimize the strategy through the feedback of the reward function;

[0052] S4-3: Invoke the Deep Q-Learning algorithm and update the value, the value update formula is:

[0053]

[0054] Among them, represents the value corresponding to the current state and the action , represents the immediate reward obtained from the environment, represents the discount factor, represents the target network parameters, represents the learning rate.

[0055] Beneficial effects: The reinforcement learning model can be initially trained through the historical data and control strategies in the graph reasoning of the knowledge graph model to quickly adapt to the real-time state and changing working conditions of the equipment, thus avoiding the trial-and-error process in traditional reinforcement learning methods; and by continuously updating the value, the reinforcement learning model can select the control signal most suitable for the current state, thereby optimizing the control efficiency and response speed of the equipment.

[0056] Further, the S5 includes:

[0057] S5-1: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the molding equipment. The reinforcement learning model generates an optimal control signal based on the real-time collected operation data of the molding equipment, and the control system selects the optimal control action through the policy. The expression is:

[0058]

[0059] where, represents the control action selected at the current moment , represents the current equipment state, represents the target network parameters, the value represents the expected reward of this action in the current state;

[0060] S5-2: Then, based on the inference of the knowledge graph model, support the learning decision of the reinforcement learning model and adjust the control signal of the molding equipment in real time;

[0061] S5-3: Calculate the immediate reward according to the feedback of the molding equipment, and update the value through the Bellman equation to optimize the control signal. The update equation is:

[0062]

[0063] where, represents the current state and the action corresponding value, represents the immediate reward obtained from the environment, represents the discount factor, represents the target network parameters, represents the learning rate, represents the maximum value of the next state.

[0064] Beneficial effects: During the implementation of equipment dynamic control, by dynamically generating control signals and adjusting the key parameters of the equipment in real time, the operation efficiency and stability of the system under complex working conditions are ensured. The deep combination of the reinforcement learning model and knowledge graph reasoning endows the system with high flexibility and intelligent capabilities. It can not only optimize the control signals according to the real-time state of the equipment but also continuously optimize the control strategy through feedback, thus significantly improving the accuracy and adaptability of equipment operation. Brief Description of the Drawings

[0065] Figure 1 It is a flowchart of an embodiment of the present invention;

[0066] Figure 2 It is a logic block diagram of an embodiment of the present invention;

[0067] Figure 3 It is a schematic diagram of the graph structure of the knowledge graph model in an embodiment of the present invention;

[0068] Figure 4 It is a schematic diagram of the Deep Q-Learning algorithm process in an embodiment of the present invention. Detailed Embodiment

[0069] The following is a more detailed description through specific embodiments:

[0070] The embodiment is basically as shown in the attached Figure 1 and Figure 2 figures: The large-inertia dynamic flow balance and control method for a forming equipment based on a knowledge graph includes:

[0071] S1: Collect the real-time operation data and historical operation data of the forming equipment, and preprocess the real-time operation data to generate preprocessed real-time operation data; where S1 includes:

[0072] S1-1: Collect the real-time operation data and historical operation data of the forming equipment;

[0073] S1-2: Use a moving average operation to denoise and smooth the collected real-time operation data. The expression is:

[0074]

[0075] where represents the moving average value at time , represents the window size, represents the th data point of the time series data, represents the current time;

[0076] S1-3: Scale the real-time operation data that has been denoised and smoothed to a specific interval through normalization processing. The expression is as follows:

[0077]

[0078] Among them, represents the normalized data point, represents a data point in the real-time operation data, represents the minimum value in the real-time operation data, represents the maximum value in the real-time operation data.

[0079] In this embodiment, the real-time operation data of the equipment collected is the key parameters such as the flow rate, pressure, and temperature of the equipment monitored in real time by sensors, and the historical operation data collected at the same time includes the historical records of key parameters, fault records, and previous control strategies.

[0080] In the process of collecting real-time operation data, because it is monitored in real time by sensors, the sensor data may contain noise, that is, random errors or meaningless fluctuations. These noises may come from hardware failures, environmental changes, or other factors, and may cause the equipment control system to make misleading decisions, thus affecting the training and inference results of the model. Therefore, using a moving average operation for denoising and smoothing can effectively reduce the short-term fluctuations in the data, capture the long-term trend, and ensure that the control signal of the equipment is more stable; while the normalization operation aims at the possible differences in the measurement units and magnitudes of the data sources, which will cause some features to dominate in model training. Therefore, the differences are eliminated through the normalization operation.

[0081] S2: Based on the preprocessed real-time operation data and historical operation data, call the NEO4J database to construct a knowledge graph model based on entity nodes and relationships; among them, S2 includes:

[0082] S2-1: Extract the historical operation parameters, standard production process parameters, case experiences in the actual production process, physical laws, control theories, and manufacturing site processing records and experience notes from the preprocessed real-time operation data and historical operation data to obtain data features.

[0083] S2-2: Call the NEO4J database to convert the data features into a graph structure and construct a knowledge graph model based on entity nodes and relationships; among them, the entity nodes represent key entities, and the relationships are represented by different types of edges to indicate the mutual relationships between entities.

[0084] In this embodiment, the data used to construct the knowledge graph model includes historical operation parameters, standard production process parameters, physical laws, control theories, manufacturing thread processing records, and experience notes. In addition, it also includes case experiences in the actual production process. Among them, case experiences can not only avoid the difficulty of reasoning in case-based reasoning when facing new problems but also avoid the inability to handle abnormal situations in rule-based reasoning.

[0085] Based on the above data characteristics, the NEO4J database is called to construct a knowledge graph model based on entity nodes and relationships, and then the data characteristics are transformed into a graph structure. Specifically, entity nodes represent key entities, such as equipment nodes, sensor nodes, control signal nodes, and historical data nodes. Relationships are represented by different types of edges to indicate the mutual relationships between entities, such as "acquire", "influence", "based on", etc. Then, on the key entity nodes, the relationships are specifically manifested as follows:

[0086] The equipment node is connected to the sensor node through the HAS relationship;

[0087] The sensor node is connected to the control signal node through the INFLUENCES relationship;

[0088] The control signal node is connected to the historical data node through the BASED_ON relationship.

[0089] The graph structure of the generated knowledge graph model is as Figure 3 shown. Therefore, the above graph structure design of this solution can comprehensively and accurately represent the complex relationships among equipment, sensors, control signals, and historical data, enabling the system to quickly query and reason in the subsequent reasoning process, automatically select appropriate control signals, and thus optimize the control strategy.

[0090] S3: Construct a reinforcement learning model. In the reinforcement learning model, the real-time operation data of the forming equipment is used to define and construct the state space, the control signal of the forming equipment is used to define and construct the action space, and a preset forming equipment performance standard is used to define and construct the reward function. Among them, S3 includes:

[0091] S3-1: Construct a reinforcement learning model based on the Deep Q-Learning algorithm;

[0092] S3-2: Determine the state space. The real-time operation data of the forming equipment is used to define and construct the state space, and the expression is:

[0093]

[0094] Among them, represents the state at time , , and respectively represent the parameters of the equipment during real-time operation; in this embodiment, represents the flow rate, represents the pressure, represents the temperature. In other embodiments of this embodiment, other equipment parameters may also be included.

[0095] S3-3: Determine the action space, and define the construction action space with the control signal of the forming equipment. The expression is:

[0096]

[0097] where, represents the action at time , , , respectively represent the control signal parameters; in this embodiment, represents the action of increasing the flow rate, represents the action of decreasing the flow rate, represents the action of increasing the pressure, represents the action of decreasing the pressure. In addition, actions such as increasing the temperature and decreasing the temperature are also included, and the embodiments of this application do not make limitations.

[0098] S3-4: Determine the reward function, and define the construction reward function with the preset performance standard of the forming equipment. The reward function includes positive rewards and negative rewards; in this embodiment, the reward function is used to evaluate the system performance after the reinforcement learning model takes a certain action in the current state. Therefore, the reward function guides the learning process of the reinforcement learning model by feedback of the system effectiveness. Specifically, in this solution, the reward function is associated with the equipment performance, such as flow control accuracy, pressure stability, etc. The corresponding reward function expression is:

[0099]

[0100] Therefore, in the above reward function expression, positive rewards and negative rewards are included. Positive rewards indicate that the equipment operates well, while negative rewards indicate that the system state is unstable or the efficiency is low. Therefore, by designing positive rewards and negative rewards, the model can adjust the control strategy in real time through feedback and optimize the control effect of the equipment.

[0101] S3-5: Call the Deep Q-Learning algorithm to approximate the Q function, and process the state space and action space through a deep neural network to continuously optimize the Q value to generate the optimal control strategy of the equipment. The expression is:

[0102]

[0103] where, Indicates the state to perform an action of value, represents the immediate reward, represents the discount factor, represents the target network parameters.

[0104] In this embodiment, the schematic diagram of the Deep Q-Learning algorithm process is as Figure 4 shown. By using the Deep Q-Learning algorithm, the model can continuously adjust the control signal to adapt to the dynamic changes of the equipment, ensuring the optimized operation of the system.

[0105] S4: Interact the reinforcement learning model with the large inertia dynamic flow balance and precise control scenario of the forming equipment, initialize the state according to the state space, select an action according to the current state of the forming equipment, feedback and optimize the strategy through the reward function, and provide control strategy guidance for the initial stage of the training of the reinforcement learning model through the knowledge graph model to train the reinforcement learning model to generate the optimal control signal and obtain the trained reinforcement learning model; where S4 includes:

[0106] S4-1: Initially train the reinforcement learning model through the historical operation data and control strategy in the knowledge graph; in this embodiment, the reinforcement learning model can be initially trained through the historical data and control strategy in the graph reasoning, quickly adapt to the real-time state and changing working conditions of the equipment, thus avoiding the trial-and-error process in traditional reinforcement learning methods. The historical data and rules provided by the graph reasoning help the agent avoid unnecessary exploration, quickly converge to the optimal control strategy, and enhance the adaptability and flexibility of the system.

[0107] S4-2: Interact the initially trained reinforcement learning model with the control system of the large inertia dynamic flow balance and precise control scenario of the forming equipment in real time, and during the real-time interaction process, initialize the state according to the state space, select an action according to the current state of the forming equipment, and feedback and optimize the strategy through the reward function; in addition, to improve the learning effect and stability, the reinforcement learning model adopts the experience replay and target network mechanisms, learns from historical experiences, and reduces the variance in the update process.

[0108] S4-3: Call the Deep Q-Learning algorithm and update the value, The value update formula is:

[0109]

[0110] where, represents the current state and action corresponding value representing the immediate reward obtained from the environment representing the discount factor representing the target network parameters representing the learning rate; by continuously updating the value, the reinforcement learning model can select the control signal most suitable for the current state, thereby optimizing the control efficiency and response speed of the equipment.

[0111] S5: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the forming equipment. The control system uses the knowledge graph model to infer the control strategy of the reinforcement learning model, and calls the reinforcement learning model to generate an action control signal according to the inferred control strategy to control the operating state of the forming equipment. The reinforcement learning model then continuously optimizes the control strategy of the forming equipment according to the feedback of the forming equipment; where S5 includes:

[0112] S5-1: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the forming equipment. The reinforcement learning model generates an optimal control signal based on the real-time collected operating data of the forming equipment, and the control system selects the optimal control action through the policy. The expression is:

[0113]

[0114] where represents the control action selected at the current time represents the current equipment state represents the target network parameters the value represents the expected reward of this action in the current state;

[0115] S5-2: Then, based on the inference of the knowledge graph model, support the learning decision of the reinforcement learning model and adjust the control signal of the forming equipment in real time;

[0116] S5-3: Calculate the immediate reward based on the feedback of the forming equipment, and update the value through the Bellman equation to optimize the control signal. The update equation is:

[0117]

[0118] where represents the current state and action corresponding value representing the immediate reward obtained from the environment representing the discount factor Denote the target network parameters, Denote the learning rate, Denote the maximum of the next state value.

[0119] In this embodiment, the reinforcement learning model can be combined with the historical data and inference rules in the knowledge graph, and provide optimized decisions for the control signals through graph reasoning. Specifically, based on the historical equipment operation data and equipment performance, the graph reasoning engine can deduce the control strategy in advance under similar working conditions, thereby accelerating the training process of the reinforcement learning model and improving the convergence speed and stability of the model.

[0120] Therefore, in this solution, the proposed method for large inertia dynamic balance and control of forming equipment based on the knowledge graph protects the accurate control of the key parameters of the forming equipment by means of intelligence, ensuring that the equipment can maintain an efficient and stable operation even in complex and changeable working conditions.

[0121] First, the data collection and preprocessing link is extremely delicate. The system captures in real time the data such as flow rate, pressure, temperature, etc. covered by the equipment sensors, as well as precious historical operation data, and preprocesses them by professional means. On the one hand, the moving average method is used to cleverly eliminate short-term fluctuations, just like brushing away the dust on the surface of the data, making the data curve smoother; on the other hand, the normalization technology is used to solve the problems brought by the differences in units and magnitudes of different sensors, laying a solid and stable data foundation for the subsequent modeling work, ensuring that each piece of data can accurately "speak".

[0122] Second, the construction of the knowledge graph is of great significance. Based on the data processed in the early stage, the system builds a multi-level knowledge graph, just like weaving a large intelligent information network for the equipment, closely associating the equipment, sensors, control signals and historical data. Relying on the powerful NEO4J graph structure, integrating historical operation parameters, standard process parameters, physical laws and rich production experience, not only realizes the high-speed storage and convenient query of data, but more importantly, its intelligent reasoning and decision support functions are activated, injecting a "smart brain" into the system operation, and can give scientific and reasonable decision guidance at critical moments.

[0123] Thirdly, the reinforcement learning model is exquisitely designed. To achieve the ultimate goal of precise control, the system carefully constructs a reinforcement learning model based on Deep Q-Learning. This model is precisely defined, with the real-time states of the equipment, such as flow rate, pressure, etc., embodied as the state space, and control signals, such as flow rate adjustment, pressure adjustment, becoming the action space. At the same time, a reward function is introduced, giving a clear optimization direction for improving the stability and efficiency of the equipment. During the training process, the model continuously refines the control strategy through in-depth interaction with the environment, and innovatively applies the experience replay and target network mechanisms, like pressing the "acceleration button" for the learning process, cleverly avoiding the problems of time and resource waste caused by a large number of trials and errors in traditional reinforcement learning.

[0124] Fourthly, deployment and optimization work together. The trained reinforcement learning model successfully enters the control system, generating control signals in real time and accurately, and working in perfect harmony with the knowledge graph reasoning function to dynamically adjust the operating state according to the equipment feedback. Whether it is flow rate or pressure, the control signals can flexibly change according to the real-time needs of the equipment and continuously optimize, making the equipment like a well-trained athlete, always maintaining the best operating state.

[0125] In summary, the excellent dynamic control ability demonstrated by this system under complex working conditions is remarkable, greatly improving the intelligent level of the equipment, taking the control accuracy to a new height, and significantly enhancing the operation flexibility. It undoubtedly has extremely high application and promotion value in the industrial field and is expected to reshape the new pattern of the control of forming equipment.

[0126] The above are only the embodiments of the present invention. Specific structures and common knowledge such as characteristics well known in the art are not described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the filing date or the priority date, can know all the existing technologies in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several modifications and improvements can still be made, and these should also be regarded as the protection scope of the present invention, which will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to explain the content of the claims.

Claims

1. A large inertia dynamic flow balance and control method for forming equipment based on knowledge graph, characterized by: include: S1: collect real-time operation data and historical operation data of the molding equipment, and pre-process the real-time operation data to generate pre-processed real-time operation data; S2: Based on the preprocessed real-time operation data and historical operation data, call the database to build a knowledge graph model based on entity nodes and relationships; S3: constructing a reinforcement learning model, in which the state space is defined by the real-time operation data of the molding equipment, the action space is defined by the control signal of the molding equipment, and the reward function is defined by the preset molding equipment performance standard; S4: The reinforcement learning model interacts with the large inertia dynamic flow balance and precise control scenario of the molding equipment, initializes the state according to the state space, selects actions according to the current state of the molding equipment, optimizes the strategy through reward function feedback, and provides control strategy guidance for the initial stage of the reinforcement learning model training through the knowledge graph model, so as to train the reinforcement learning model to generate the optimal control signal and obtain a trained reinforcement learning model; S5: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the molding equipment. The control system uses the knowledge graph model to infer the control strategy of the reinforcement learning model, and calls the reinforcement learning model to generate action control signals according to the inferred control strategy to control the operating state of the molding equipment. The reinforcement learning model then continuously optimizes the control strategy of the molding equipment based on the feedback from the molding equipment. The S2 includes: S2-1: Extract the historical operation parameters, standard production process parameters, case experience, physical laws, control theory, and manufacturing site processing records and experience notes from the pre-processed real-time operation data and historical operation data to obtain data features; S2-2: Call the NEO4J database, convert data features into a graph structure, and build a knowledge graph model based on entity nodes and relationships; the entity nodes represent key entities, and the relationships represent the relationships between entities through different types of edges.

2. The method for balancing and controlling large inertia dynamic flow of forming equipment based on knowledge graph according to claim 1 is characterized in that: The S1 includes: S1-1: Collect real-time operation data and historical operation data of molding equipment; S1-2: Use the moving average operation to perform denoising and smoothing on the collected real-time running data. The expression is: in, Indicates at time The moving average of Indicates the window size, Represents the time series data data points, Indicates the current moment; S1-3: The real-time running data that has been denoised and smoothed is normalized and scaled to a specific range In the expression: in, represents the normalized data points, Represents a data point in real-time running data. Indicates the minimum value in real-time running data. Indicates the maximum value in real-time running data.

3. The large inertia dynamic flow balance and control method of forming equipment based on knowledge graph according to claim 1 is characterized by: In said S2-2, the entity nodes include equipment nodes, sensor nodes, control signal nodes and historical data nodes; The relationship in S2-2 represents the relationship between entities through different types of edges: Equipment nodes and sensor nodes are connected through HAS relationships; The sensor nodes and control signal nodes are connected through the INFLUENCES relationship; The control signal node is connected to the historical data node through the BASED_ON relationship.

4. The large inertia dynamic flow balance and control method of forming equipment based on knowledge graph according to claim 3 is characterized by: The S3 includes: S3-1: Building a reinforcement learning model based on the Deep Q-Learning algorithm; S3-2: Determine the state space. Define the construction state space with the real-time operation data of the molding equipment. The expression is: in, Indicates time status, , and They respectively represent the parameters of the real-time operation of the equipment; S3-3: Determine the action space, and define the action space with the control signal of the molding equipment. The expression is: in, express The action of the moment, , , , Respectively represent each control signal parameter; S3-4: Determine a reward function, and define a reward function based on a preset molding equipment performance standard, wherein the reward function includes a positive reward and a negative reward; S3-5: Call the Deep Q-Learning algorithm to approximate the Q function, and process the state space and action space through the deep neural network to continuously optimize the Q value to generate the optimal control strategy for the equipment. The expression is: in, Indicates in status Next action of value, Indicates immediate reward, represents the discount factor, Represents the target network parameters.

5. The method for balancing and controlling large inertia dynamic flow of forming equipment based on knowledge graph according to claim 4 is characterized in that: The S4 includes: S4-1: Initial training of the reinforcement learning model using historical operation data and control strategies in the knowledge graph; S4-2: The reinforcement learning model that has completed initial training interacts in real time with the control system of the large inertia dynamic flow balance and precise control scenario of the molding equipment. During the real-time interaction, the state is initialized according to the state space, the action is selected according to the current state of the molding equipment, and the optimization strategy is fed back through the reward function; S4-3: Call the Deep Q-Learning algorithm and The value is updated, The value update formula is: in, Indicates the current status and actions Corresponding value, represents the immediate reward obtained from the environment, represents the discount factor, represents the target network parameters, Represents the learning rate.

6. The knowledge graph-based large inertia dynamic flow balance and control method for forming equipment according to claim 5 is characterized by: The S5 includes: S5-1: Deploy the knowledge graph model and the trained reinforcement learning model to the control system of the molding equipment. The reinforcement learning model generates the optimal control signal based on the real-time collected molding equipment operation data. The control system is controlled by The strategy selects the optimal control action, which is expressed as: in, Indicates the current time Select the control action, Indicates the current equipment status. represents the target network parameters, The value represents the expected reward of the action in the current state; S5-2: Then, based on the knowledge graph model reasoning, the learning decision of the reinforcement learning model is supported to adjust the molding equipment control signal in real time; S5-3: Calculate the immediate reward based on the feedback of the formed equipment and update it through the Bellman equation The value is used to optimize the control signal, and the update equation is: in, Indicates the current status and actions Corresponding value, represents the immediate reward obtained from the environment, represents the discount factor, represents the target network parameters, represents the learning rate, The maximum value of the next state value.

Citation Information

Patent Citations

  • Power business data auxiliary knowledge graph construction method based on reinforcement learning

    CN118245607A