Process intelligent optimization decision-making method based on frequency conversion data acquisition

Through the combination of variable frequency data acquisition and neural network model combined with deep reinforcement learning, the problems of large errors and redundancy in data acquisition by casting enterprises are solved, efficient process optimization decisions are achieved, and production efficiency and product quality are improved.

CN120544747AActive Publication Date: 2025-08-26HUAZHONG UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510620543.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The data collection methods of existing casting companies have large errors or redundant, and the data utilization rate is low, making it difficult to convert into effective process optimization decision support.

Method used

The variable frequency data acquisition method is adopted to adjust the acquisition frequency according to the change rate of process parameters, and optimize the process parameters in combination with neural network model and deep reinforcement learning algorithm.

Benefits of technology

It realizes the accuracy and low redundancy of data acquisition, improves data utilization, optimizes casting process, and improves production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544747A_ABST
    Figure CN120544747A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of casting production process optimization decision, and particularly discloses a process intelligent optimization decision method based on frequency conversion data acquisition, which comprises the following steps: performing data acquisition for each process parameter in the casting process, and adaptively adjusting the acquisition frequency of the corresponding process parameter based on the change rate of the process parameter, the acquired data comprises acquisition data of production parameters and acquisition data of process quality; based on the collected data, training a neural network model to obtain a trained neural network model, the neural network model being used for predicting process quality based on the production parameters; and based on the trained neural network model, process parameters are optimized through deep reinforcement learning. According to the method, the technological parameters are optimized through frequency conversion data collection and combination of the trained neural network model and deep reinforcement learning, the casting technology data can be reasonably collected, and the collected data are efficiently utilized to optimize the casting technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of casting production process optimization decision-making, and more specifically, relates to a process intelligent optimization decision-making method based on variable frequency data acquisition. Background Art

[0002] As modern high-precision equipment places ever-increasing performance demands on components, casting process optimization has become a crucial tool for improving production efficiency, reducing costs, and ensuring product quality. Traditional process optimization methods rely primarily on manual experience and experimentation. However, due to the complex nonlinear relationships between process parameters and the variability of the production environment, manual optimization often struggles to achieve the optimal solution. Furthermore, with the continuous advancement of data acquisition and sensor technology, the massive amounts of data generated by industrial systems are providing new possibilities for process optimization.

[0003] Data acquisition technology has emerged in recent years within industrial control and optimization. Its core objective is to capture dynamic changes in the production process in real time through high-frequency, refined data collection. This data includes not only physical parameters like temperature, pressure, and flow, but also more detailed variables like equipment vibration, power consumption, and chemical reaction rates. This changing data provides real-time, accurate raw information for intelligent decision-making, enabling companies to promptly adjust process parameters, optimize production processes, and improve product quality and efficiency. Currently, most data collection processes within foundries use fixed time intervals. If the collection frequency is lower than the frequency of data changes, errors can easily occur in the collected data, leading to the loss of some information. Alternatively, if the collection frequency is significantly higher than the frequency of data changes, data redundancy can lead to a waste of storage space.

[0004] Meanwhile, despite significant progress in data acquisition technology, transforming massive amounts of dynamic data into effective decision-making support information remains a challenge for many industrial enterprises. Traditional optimization methods based on rules or statistical analysis are no longer sufficient to meet the demands of process control in complex production environments. In particular, the relationships between multivariable, high-dimensional, and nonlinear process parameters require more intelligent and adaptive optimization algorithms to ensure decision-making and adjustments within complex and ever-changing production processes.

[0005] Existing foundry data collection methods use a fixed frequency approach, resulting in large data errors and redundancy. Furthermore, the data utilization rate is low, making it difficult to translate into production benefits. Therefore, how to rationally collect casting process data and efficiently utilize this data to optimize the casting process is a pressing technical challenge in this field. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the purpose of this application is to reasonably collect casting process data and efficiently use the collected data to optimize the casting process.

[0007] To achieve the above objectives, in a first aspect, the present application provides a process intelligent optimization decision-making method based on variable frequency data acquisition, the method comprising: Data is collected for various process parameters (production parameters or process quality) during the casting process. The collection frequency of the corresponding process parameters is adjusted based on the rate of change of the process parameters. The collected data includes production parameter data and process quality data. Based on the collected data, a neural network model is trained to obtain the trained neural network model, and the neural network model is used to predict process quality based on production parameters; Based on the trained neural network model, process parameters are optimized through deep reinforcement learning.

[0008] It is understandable that by collecting variable frequency data and optimizing process parameters in combination with trained neural network models and deep reinforcement learning, it is possible to reasonably collect casting process data and efficiently use the collected data to optimize the casting process.

[0009] In a possible implementation, the adaptive adjustment of the acquisition frequency of the corresponding process parameters based on the change rate of the process parameters includes: When the change rate of the process parameter is less than the first change rate threshold, the acquisition frequency of the process parameter is kept unchanged; When the rate of change of the process parameter is greater than or equal to the first rate of change threshold and the rate of change of the process parameter is less than the second rate of change threshold, increasing the frequency of collecting the process parameter (decreasing the time interval for collecting the process parameter) according to the first adjustment ratio; When the change rate of the process parameter is greater than or equal to the second change rate threshold, increasing the acquisition frequency of the process parameter (reducing the acquisition time interval) according to the second adjustment ratio; The first change rate threshold is smaller than the second change rate threshold, and the first adjustment ratio is smaller than the second adjustment ratio.

[0010] For example, the first adjustment ratio is (e.g. 25%), accordingly, , increasing the acquisition frequency of process parameters can be ,in, is the original acquisition frequency, In accordance with right The new acquisition frequency obtained after the increase.

[0011] For example, the second adjustment ratio is (e.g. 50%), accordingly, , increasing the acquisition frequency of process parameters can be ,in, is the original acquisition frequency, In accordance with right The new acquisition frequency obtained after the increase.

[0012] In one possible implementation, the process parameter optimization through deep reinforcement learning based on the trained neural network model includes: Build a deep reinforcement learning environment, which includes a state space, an action space, and a reward function. The state space includes production parameters and process quality. The process quality in the state space is determined based on the production parameters and the trained neural network model. The action space is a set of actions targeting the production parameters. The reward function is used to characterize the mapping relationship between states and rewards in the state space. Based on the deep reinforcement learning environment, the deep Q network (DQN, including Q value network and target network) model is trained using the epsilon-greedy strategy; Based on the Q-value network in the trained deep Q-network, the optimal process parameters are obtained.

[0013] In one possible implementation, the action space includes a decrease action, a keep-alive action, and an increase action corresponding to each production parameter; Decrease action, used to reduce production parameters according to the preset change ratio; Keep the action unchanged, used to keep the production parameters unchanged; The increase action is used to increase the production parameters according to the preset change ratio.

[0014] Here we illustrate the reduction action, assuming that the preset change ratio is (For example, 0.01, 0.05, 0.1, etc.), according to Original production parameters If the new production parameters are reduced Determined by the following formula .

[0015] Here we illustrate the increase action, assuming that the preset change ratio is (For example, 0.01, 0.05, 0.1, etc.), according to Original production parameters If the new production parameters are increased Determined by the following formula .

[0016] In a possible implementation, the production parameters in the state space are configured as multiple parameter groups, and one parameter group includes multiple production parameters; The actions of production parameters in the same parameter group remain synchronized; At least one negative correlation is configured among the plurality of parameter groups, where the negative correlation indicates that, for a first parameter group and a second parameter group, an action of the first parameter group is negatively correlated with an action of the second parameter group.

[0017] Here we provide an example of how to synchronize the actions of the production parameters in the same parameter group. , which includes three parameters, namely 、 and When performing actions on these three parameters, the actions of the three parameters remain synchronized. For example, the actions of the three parameters are all decreasing, or the actions of the three parameters are all remaining unchanged, or the actions of the three parameters are all increasing.

[0018] Here we illustrate the negative correlation relationship by example. For two parameter groups in multiple parameter groups and ,exist 、 and When these three parameters execute the action, 、 and These three parameters perform opposite actions. For example, in 、 and When these three parameters perform reduction actions, 、 and These three parameters perform the increment action; 、 and When these three parameters perform the increase action, 、 and These three parameters perform the reduction action; 、 and When these three parameters are kept constant, 、 and These three parameters perform the same action.

[0019] In one possible implementation, the deep reinforcement learning environment described above trains the deep Q network model using an epsilon-greedy strategy, including: Initialize the buffer experience pool; Initialize the deep reinforcement learning environment; Continue to perform iterative training operations using the epsilon-greedy strategy until the maximum number of iterative training times is reached; The iterative training operation includes: Get the current state of the deep reinforcement learning environment ; Perform a first action selection operation with a first probability, perform a second action selection operation with a second probability, and select the action based on the selected and the trained neural network model to determine the next state , and based on the next state and reward function, determine the action Corresponding rewards ; The first probability is 1-epsilon, and the second probability is epsilon; The first action selection operation includes obtaining the state through the Q value network The optimal action under , and determine the optimal action is The second action selection operation includes randomly selecting an action from actions other than the action selected by the first action selection operation. ; Will The experience tuples are stored in the buffer experience pool, and the number of tuples in the buffer experience pool is controlled to be less than the upper limit by removing the earliest experience tuple put into the buffer experience pool. Extract a preset number of experience tuples from the buffer experience pool, and update the Q value network based on the extracted experience tuples; If the number of iterations between the last time the target value network was updated and the current time is greater than a preset value, the parameters of the Q value network are copied to the target network to update the target network.

[0020] Here, the above actions are selected and the trained neural network model to determine the next state Provide an example to illustrate the current status Execute actions with production parameters in , can get the next state The production parameters in the next state The production parameters in are input into the trained neural network model, and the output of the neural network model is used as the next state The quality of workmanship.

[0021] In one possible implementation, the Q-value network in the trained deep Q network is used to obtain optimal process parameters, including: Initialize the deep reinforcement learning environment; Continue to perform process parameter exploration operations until the maximum number of explorations is reached. The process parameter exploration operation includes selecting the current state through the Q value network The best action under , according to the action and the trained neural network model to determine the next state , and based on the next state and reward function, determine the action Rewards ; Select the process parameter with the highest reward value to explore the corresponding state of the operation as the optimal process parameters.

[0022] In a second aspect, the present application provides a process intelligent optimization decision-making device based on variable frequency data acquisition, comprising: The data acquisition module is used to collect data on various process parameters during the casting process and adjust the collection frequency of the corresponding process parameters based on the rate of change of the process parameters. The collected data includes production parameter data and process quality data; A neural network model training module is used to train a neural network model based on the collected data and obtain the trained neural network model. The neural network model is used to predict process quality based on production parameters; The process parameter optimization module is used to optimize process parameters through deep reinforcement learning based on the trained neural network model.

[0023] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0024] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.

[0025] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0026] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies: (1) By adjusting the acquisition frequency of the corresponding process parameters based on the rate of change of the process parameters during the data acquisition process, variable frequency data acquisition can be achieved, so that the acquisition frequency of the process parameters can adapt to the rate of change of the process parameters, effectively avoiding the acquisition frequency being lower than the frequency of data change, ensuring the accuracy of data acquisition, and effectively avoiding the acquisition frequency being significantly higher than the frequency of data change, reducing data redundancy, and taking into account both the accuracy of data acquisition and the low redundancy of data acquisition, so as to achieve reasonable acquisition of casting process data.

[0027] (2) After collecting the data, the neural network model is trained based on the collected data. The trained neural network model can be used to characterize the intrinsic relationship between production parameters and process quality. Then, combined with the deep reinforcement learning algorithm, intelligent process optimization decisions can be made to obtain optimized process parameters, thereby achieving efficient use of the collected data to optimize the casting process. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Schematic diagram of a process intelligent optimization decision-making method based on variable frequency data acquisition provided by an embodiment of the present application; Figure 2 This is a flow chart of deep reinforcement model training using the epsilon-greedy strategy provided in an embodiment of the present application; Figure 3 Schematic diagram of the structure of a process intelligent optimization decision-making device based on variable frequency data acquisition provided by an embodiment of the present application; Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0030] In the specification and claims of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, a first change rate threshold and a second change rate threshold are used to distinguish different change rate thresholds, rather than to describe a specific order of the change rate thresholds.

[0031] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0032] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0033] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0034] Figure 1 This is a flow chart of a process intelligent optimization decision-making method based on variable frequency data acquisition provided by an embodiment of the present application, such as Figure 1 As shown, the method includes the following steps S101 to S103.

[0035] Step S101: Data is collected for various process parameters (production parameters or process quality) during the casting process, and the collection frequency of the corresponding process parameters is adjusted based on the rate of change of the process parameters. The collected data includes the collected data of the production parameters and the collected data of the process quality; Step S102: training a neural network model based on the collected data to obtain a trained neural network model, where the neural network model is used to predict process quality based on production parameters; Step S103: Optimize process parameters through deep reinforcement learning based on the trained neural network model.

[0036] Specifically, by adaptively adjusting the acquisition frequency of corresponding process parameters based on the rate of change of process parameters during the data acquisition process, variable frequency data acquisition can be achieved, so that the acquisition frequency of process parameters can adapt to the rate of change of process parameters, effectively avoiding the acquisition frequency being lower than the frequency of data changes, ensuring the accuracy of data acquisition, and effectively avoiding the acquisition frequency being significantly higher than the frequency of data changes, reducing data redundancy, and taking into account both the accuracy of data acquisition and the low redundancy of data acquisition, thereby realizing the reasonable acquisition of casting process data.

[0037] After collecting the data, the neural network model is trained based on the collected data. The trained neural network model can be used to characterize the intrinsic relationship between production parameters and process quality. Then, combined with the deep reinforcement learning algorithm, intelligent process optimization decisions can be made, and optimized process parameters can be obtained, so as to efficiently use the collected data to optimize the casting process.

[0038] It is worth noting that by reasonably collecting casting process data, the accuracy and low redundancy of data collection can improve the training effect of the neural network model, and then the trained neural network model can more accurately characterize the intrinsic relationship between production parameters and process quality, thereby improving the effect of deep reinforcement learning algorithm on process parameter optimization.

[0039] Therefore, by acquiring variable frequency data and optimizing process parameters by combining trained neural network models and deep reinforcement learning, it is possible to reasonably acquire casting process data and efficiently utilize the collected data to optimize the casting process.

[0040] The following example illustrates the process intelligent optimization decision-making method based on variable frequency data acquisition provided by this application. The method includes the following steps 1 to 6.

[0041] Step 1: Collect data on various process parameters during the casting process.

[0042] Built-in sensors and built-in PLCs are sensors and PLCs that come with the casting equipment. Additional sensors and additional PLCs are sensors and PLCs that are attached to the casting equipment. By using built-in sensors, built-in PLCs, additional sensors, and additional PLCs, the relevant equipment in the core making, smelting, pouring, heat treatment, and quality inspection processes of the casting process are modified. First, the degree of data change is calculated: assuming that at the previous time point The data value is , current time point The data value is , calculate the data change value at two moments Relative to The ratio of the data to the original data is used as the rate of change. If the rate of change is less than 5%, the original collection frequency remains unchanged. If the rate of change is ≥5% and less than 50%, the collection interval is reduced by 25% based on the original collection interval. If the rate of change is ≥50%, the collection interval is reduced by 50% based on the original collection interval. The collected data is stored in a large-scale relational database, SQL Server.

[0043] It should be noted that the above-mentioned variable frequency data acquisition process does not take into account the situation where the time interval increases, because the initial acquisition frequency will be set slightly smaller (that is, the initial acquisition time interval is larger). If the change rate of the collected data is large, the acquisition frequency will be gradually increased according to the above-mentioned variable frequency data acquisition method (that is, the acquisition time interval will be gradually reduced). Since the acquisition frequency increases because there are large fluctuations in the collected data within the corresponding time interval, and there is no guarantee that similar fluctuations in the collected data will not occur in the future, in order to ensure that the collected data is more in line with the actual situation, the acquisition frequency will no longer be reduced (that is, the acquisition time interval will no longer be increased) to ensure accuracy.

[0044] During the collection process, some process parameters are collected at a high frequency and a large amount of data is collected, while some process parameters are collected at a low frequency and a small amount of data is collected. When constructing the data set later, in order to reduce the difference in the amount of collected data between different process parameters, the process parameters with smaller data volumes can be interpolated to keep the amount of collected data between different process parameters the same.

[0045] Step 2: Build a dataset based on the collected data.

[0046] Based on the multiple collected process parameters, the production time recorded during data collection (i.e., the collection time), and the casting number, a SQL Server database association query is used for association mapping to form a data set. The data set includes the collected data of production parameters and process quality. The collected production parameter data includes: production parameters of various processes, such as the sandblasting pressure, sandblasting time, solidification temperature, and solidification time for core making; the melting volume, refining temperature, and refining agent dosage for smelting; the melt temperature, filling pressure, filling time, holding pressure, and holding time for the casting process; the solution temperature, solution time, aging temperature, and aging time for the heat treatment process; and process quality data: the composition, performance, internal quality, and geometric dimensions of the casting.

[0047] Step 3: Train the neural network model based on the dataset.

[0048] Based on the above data set, a neural network model is established through training. The model includes an input layer, a hidden layer, and an output layer. The number of nodes in the input layer is the total number of production parameters, and the number of hidden layers and neurons is determined by empirical rules.

[0049] ; Where h represents the number of hidden layer nodes, n represents the number of input layer nodes, m represents the number of output layer nodes, and a is a constant with a value range of [1, 10]. The hyperbolic tangent function with a range of (-1, 1) is used as the normalization and activation function.

[0050] It can be understood that the neural network model is used to predict process quality based on production parameters. The input of the neural network model is production parameters (production parameters in core making, smelting, pouring, and heat treatment), and the output of the neural network model is process quality (parameters of the quality inspection process, such as the composition, performance, internal quality, and geometric dimensions of the casting).

[0051] Step 4: Model the process optimization problem as a deep reinforcement learning environment.

[0052] Based on the established neural network model, an optimization algorithm, the DQN algorithm, is used to optimize production parameters to improve process quality. This algorithm models the process optimization problem as a deep reinforcement learning environment consisting of a state space, an action space, and a reward function.

[0053] The state space includes all the information describing the current process.

[0054] ; Represents the process parameters of core making {sand shooting pressure, sand shooting time, curing temperature, curing time}, Represents the process parameters of the smelting process {melting amount, refining temperature, refining agent dosage}, Represents the casting process parameters {melt temperature, filling pressure, filling time, holding pressure, holding time}, Represents the process parameters of heat treatment process {solution temperature, solution time, aging temperature, aging time}, Indicates whether the ingredients are qualified under the current process parameters. Represents the performance data under the current process parameters, Represents the internal defect data of the casting, Represents the geometric dimensions of the casting. Among them, the production parameters of the current process include , the process quality of the current process includes .

[0055] The state space needs to be normalized, and the normalization formula is as follows: ; in, is a parameter The normalized value, is the lower limit of parameter CP, is the upper limit of the parameter CP.

[0056] The action space is the trend of production parameter changes that the agent can choose. Since the value range of production parameters is continuous, the parameters must be discretized before action selection. The specific process is as follows: (1) According to actual production experience, the production parameters are grouped into 8 groups. The first group is {sand blasting pressure, solidification temperature}, the second group is {solution temperature, aging temperature}, the third group is {melting amount, refining agent dosage}, the fourth group is {refining temperature, melt temperature}, the fifth group is {holding pressure, filling pressure}, the sixth group is {sand blasting time, solidification time}, the seventh group is {filling time, holding time}, and the eighth group is {solution time, aging time}.

[0057] (2) Set the action space to {slightly decrease, remain unchanged, slightly increase}, represented by the numbers {-1, 0, 1}. In the grouping in (1), the first five groups of production parameters can independently and randomly select one action in the action space. In the last three groups, the action of group 6 is opposite to that of group 1, the action of group 7 is opposite to that of group 5, and the action of group 8 is opposite to that of group 2. If groups 1, 2, and 5 remain unchanged, then groups 6, 7, and 8 also remain unchanged.

[0058] (3) The change ratio for a small decrease or increase is a value in {0.01, 0.05, 0.1}. The specific value will vary according to the range of each production parameter. For production parameters with a large range (≥100), the change amplitude is 0.1 each time. For production parameters with a small range (≤15), the change amplitude is 0.01. For other ranges, the change amplitude is 0.05. Among them, the unit of production parameters related to temperature is ℃, the unit of production parameters related to time is s, and the unit of production parameters related to weight is g. When the value of the production parameter after the change exceeds the upper or lower limit, the value is assigned to the limit.

[0059] In this invention, the reward function is set as: ; Among them, set Content_t to the reward related to the casting component. If it is unqualified, then , if qualified, it is 0; Properties_t is a reward related to performance quantitative parameters, such as tensile strength. The specific calculation formula is: actual tensile strength Lf - qualified tensile strength Lo; Quality_t is the reward related to internal quality. If there are internal defects, Quality_t = number of defects Nd 10; Size_t is the reward related to the quality of the geometric size, similar to performance, and its specific value is: actual size - specified size; ~ is the weight coefficient, which can be modified according to actual production needs. This application sets .

[0060] It can be understood that the above-mentioned neural network model is used to determine the process quality data in each state based on the production parameter data in each state; based on the reward function and the process quality data in the next state, the reward obtained by performing the specified action in the current state can be determined.

[0061] Step 5: Perform deep reinforcement model training using the epsilon-greedy strategy.

[0062] Figure 2 This is a flow chart of the deep reinforcement model training using the epsilon-greedy strategy provided by the embodiment of the present application, such as Figure 2 As shown, the state space is initialized first. The initial value of the state is the lower limit of all parameters. The Q value network parameters are randomly initialized and the parameters of the initialized Q value network are assigned to the target network (TargetNetwork). After initialization, the training process is carried out as follows: Step 51, initialize the replay buffer: for the initial state , generate experience tuples through random strategies Then from the state Start, generate the next experience tuple, and run until the initial experience pool with a capacity of 10,000 is reached.

[0063] Step 52: reinitialize the environment state, with the initial value of the state being the lower limit value of all parameters.

[0064] Step 53: Get the current state of the environment .

[0065] Step 54: Get the sample state through the Q value network The best action under , and select the action with probability 1-epsilon, randomly select another action with probability epsilon, and get the reward value based on the selected action and the next state .

[0066] Step 55, Store in the buffer experience pool. If the buffer experience pool capacity reaches the upper limit (100000) before storage, the oldest experience (the experience tuple first put into the experience pool) will be removed and then Deposit into the buffer experience pool.

[0067] In step 56, a mini-batch of 64 empirical tuple samples is randomly sampled from the replay buffer, and each sample is fed into the Q-value network and the target network. The value obtained by the following formula is used as the variance, and the batch average variance is calculated. The average variance is used as the loss function to update the Q-value network.

[0068] ; in, is the reward value of the current experience tuple, is the next state in the experience tuple The target network prediction value corresponding to the optimal action of is the Q value of the network for the current state and actions The predicted value of is the discount factor, which is set to 0.9 in this application.

[0069] Step 57: If the number of iterations between the last time the target value network was updated and the current time is greater than 100, the parameters of the Q value network are copied to the target network.

[0070] Step 58: When the number of iterations is ≤ T, repeat steps 53 to 57, where T is 100,000.

[0071] Step 6: Obtain the optimal process parameters based on the Q-value network in the trained deep Q network.

[0072] After the training is completed, the state is reinitialized. The initialization state can be any specified value, and then the process parameter exploration operation is continuously performed until the maximum number of explorations is reached. Specifically, the process parameter exploration operation includes: selecting the current state through the Q value network The best action under , based on the predicted action Interact with the process environment to obtain new states and rewards After executing the maximum number of explorations, the state corresponding to the process parameter exploration operation with the highest reward value is selected. Recommended process.

[0073] To sum up, the intelligent optimization decision-making method for production parameters based on data frequency conversion acquisition proposed in this application can perform frequency conversion data collection on the core making, molding, smelting, heat treatment, and quality inspection processes of casting in the actual production scenario of casting, and then make intelligent optimization decisions on the optimal process based on historical data to provide guidance for production, save costs, and improve production efficiency.

[0074] The process intelligent optimization decision-making device based on variable frequency data acquisition provided in this application is described below. The process intelligent optimization decision-making device based on variable frequency data acquisition described below and the process intelligent optimization decision-making method based on variable frequency data acquisition described above can be referenced to each other.

[0075] Figure 3 Schematic diagram of the structure of the process intelligent optimization decision-making device based on variable frequency data acquisition provided by the embodiment of the present application, such as Figure 3 As shown, the device includes: The data acquisition module 10 is used to collect data on various process parameters during the casting process and adjust the collection frequency of the corresponding process parameters based on the rate of change of the process parameters. The collected data includes the collection data of production parameters and the collection data of process quality; A neural network model training module 20 is used to train a neural network model based on the collected data and obtain a trained neural network model, wherein the neural network model is used to predict process quality based on production parameters; The process parameter optimization module 30 is used to optimize the process parameters through deep reinforcement learning based on the trained neural network model.

[0076] It is understandable that the detailed functional implementation of each of the above units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.

[0077] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.

[0078] Based on the method in the above embodiment, an embodiment of the present application provides an electronic device, Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 4 As shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830, and a communication bus 840. The processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the method in the above embodiment.

[0079] In addition, the logic instructions in the aforementioned memory 830 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0080] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0081] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0082] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0083] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.

[0084] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0085] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0086] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A process intelligent optimization decision-making method based on variable frequency data acquisition, characterized in that: include: Data is collected for various process parameters during the casting process, and the collection frequency of the corresponding process parameters is adjusted based on the rate of change of the process parameters. The collected data includes production parameter data and process quality data; Based on the collected data, a neural network model is trained to obtain the trained neural network model, and the neural network model is used to predict process quality based on production parameters; Based on the trained neural network model, process parameters are optimized through deep reinforcement learning.

2. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 1 is characterized in that: Adaptively adjusting the acquisition frequency of the corresponding process parameters based on the rate of change of the process parameters includes: When the change rate of the process parameter is less than the first change rate threshold, the acquisition frequency of the process parameter is kept unchanged; When the rate of change of the process parameter is greater than or equal to the first rate of change threshold and the rate of change of the process parameter is less than the second rate of change threshold, increasing the frequency of collecting the process parameter according to the first adjustment ratio; When the rate of change of the process parameter is greater than or equal to the second rate of change threshold, increasing the frequency of collecting the process parameter according to the second adjustment ratio; The first change rate threshold is smaller than the second change rate threshold, and the first adjustment ratio is smaller than the second adjustment ratio.

3. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 1 or 2 is characterized in that: The process parameters are optimized through deep reinforcement learning based on the trained neural network model, including: Build a deep reinforcement learning environment, which includes a state space, an action space, and a reward function. The state space includes production parameters and process quality. The process quality in the state space is determined based on the production parameters and the trained neural network model. The action space is a set of actions targeting the production parameters. The reward function is used to characterize the mapping relationship between states and rewards in the state space. Based on the deep reinforcement learning environment, the deep Q network model is trained using the epsilon-greedy strategy; Based on the Q-value network in the trained deep Q-network, the optimal process parameters are obtained.

4. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 3 is characterized in that: The action space includes the reduction action, the unchanged action and the increase action corresponding to each production parameter; Decrease action, used to reduce production parameters according to the preset change ratio; Keep the action unchanged, used to keep the production parameters unchanged; The increase action is used to increase the production parameters according to the preset change ratio.

5. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 4 is characterized in that: The production parameters in the state space are configured into multiple parameter groups, and one parameter group includes multiple production parameters; The actions of production parameters in the same parameter group remain synchronized; At least one negative correlation is configured among the plurality of parameter groups, where the negative correlation indicates that, for a first parameter group and a second parameter group, an action of the first parameter group is negatively correlated with an action of the second parameter group.

6. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 3 is characterized in that: The deep reinforcement learning environment is based on which the deep Q network model is trained using the epsilon-greedy strategy, including: Initialize the buffer experience pool; Initialize the deep reinforcement learning environment; Continue to perform iterative training operations using the epsilon-greedy strategy until the maximum number of iterative training times is reached; The iterative training operation includes: Get the current state of the deep reinforcement learning environment ; Perform a first action selection operation with a first probability, perform a second action selection operation with a second probability, and select the action based on the selected and the trained neural network model to determine the next state , and based on the next state and reward function, determine the action Corresponding rewards ; The first probability is 1-epsilon, and the second probability is epsilon; The first action selection operation includes obtaining the state through the Q value network The optimal action under , and determine the optimal action is The second action selection operation includes randomly selecting an action from actions other than the action selected by the first action selection operation. ; Will The experience tuples are stored in the buffer experience pool, and the number of tuples in the buffer experience pool is controlled to be less than the upper limit by removing the earliest experience tuple put into the buffer experience pool. Extract a preset number of experience tuples from the buffer experience pool, and update the Q value network based on the extracted experience tuples; If the number of iterations between the last time the target value network was updated and the current time is greater than a preset value, the parameters of the Q value network are copied to the target network to update the target network.

7. The process intelligent optimization decision-making method based on variable frequency data acquisition according to claim 3 is characterized in that: The method of obtaining the optimal process parameters based on the Q value network in the trained deep Q network includes: Initialize the deep reinforcement learning environment; Continue to perform process parameter exploration operations until the maximum number of explorations is reached. The process parameter exploration operation includes selecting the current state through the Q value network The best action under , according to the action and the trained neural network model to determine the next state , and based on the next state and reward function, determine the action Rewards ; Select the process parameter with the highest reward value to explore the corresponding state of the operation as the optimal process parameters.

8. A process intelligent optimization decision-making device based on variable frequency data acquisition, characterized in that: include: The data acquisition module is used to collect data on various process parameters during the casting process and adjust the collection frequency of the corresponding process parameters based on the rate of change of the process parameters. The collected data includes production parameter data and process quality data; A neural network model training module is used to train a neural network model based on the collected data and obtain the trained neural network model. The neural network model is used to predict process quality based on production parameters; The process parameter optimization module is used to optimize process parameters through deep reinforcement learning based on the trained neural network model.

9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optimization method, device and equipment for blade casting process and storage medium

    CN119784019A

  • Data monitoring systems and methods to update input channel routing in response to an alarm state

    US20190324439A1