A method and system for predicting corrosion resistance of paint based on reinforcement learning model
By combining reinforcement learning models and multi-objective optimization algorithms with YOLO v5 and Double DQN algorithms, the problems of low data extraction efficiency and single optimization strategy in the prediction of coating corrosion resistance are solved, realizing intelligent, automated and efficient multi-objective optimization of coating corrosion resistance.
Patent Information
- Application Number
- CN202511280046.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing technologies struggle to efficiently and accurately predict and optimize the corrosion resistance of coatings, especially in complex corrosive environments. Traditional methods rely on experiments and experience, lack theoretical support, and suffer from insufficient multi-objective optimization strategies, resulting in high costs and low efficiency.
A reinforcement learning-based approach was adopted, using the YOLO v5 algorithm to generate a structured corrosion performance data table. Combined with the Double DQN algorithm and the NSGA-II algorithm, a multi-objective optimization strategy for adjusting the corrosion resistance of coatings was generated, and the adjustment strategy was optimized through a closed-loop feedback mechanism.
It achieves intelligent, automated, and efficient optimization of coating corrosion resistance, improves data processing efficiency and prediction accuracy, and can adaptively adjust in complex environments to meet the needs of balancing multiple objectives.
Smart Images

Figure CN120808978B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coating corrosion resistance prediction technology, and in particular to a coating corrosion resistance prediction method and system based on reinforcement as the initial population input learning model of the NSGA-II algorithm. Background Technology
[0002] With the continuous development of industry, coatings are widely used on metal surfaces to improve the corrosion resistance of materials. However, the complex corrosive environment and the uncontrollability under long-term use make accurate prediction and effective optimization of coating corrosion resistance a research hotspot in materials science and industrial engineering. In traditional technologies, the study of coating corrosion resistance mainly relies on laboratory experiments and empirical models for prediction and optimization. These methods usually require a large number of manual experiments and chemical composition adjustments, which are costly and inefficient. In addition, because corrosion is a complex physicochemical process, its influencing factors are diverse, such as corrosive media, time, temperature, and coating composition. Results under different experimental conditions are difficult to generalize, and the predictive ability for long-term corrosion durability is very limited. In recent years, with the rapid development of artificial intelligence and big data technologies, deep learning and reinforcement learning technologies have emerged in industrial data processing, making it possible to extract key characteristics from complex experimental data and make predictions. However, directly applying these emerging technologies to the dynamic prediction and optimization of coating corrosion resistance still faces certain technical obstacles, such as how to fully utilize the time-series characteristics of corrosion data, how to implement adaptive optimization strategies for coating performance, and how to provide multi-objective optimization schemes that can guide actual industrial production.
[0003] While some breakthroughs have been achieved in intelligent research on the analysis of coating corrosion resistance, existing technologies have made some progress. For example, corrosion area detection and corrosion rate calculation based on traditional image processing techniques can initially extract corrosion characteristic data. However, traditional image algorithms have limited ability to identify corrosion features, especially in complex corrosive environments, and cannot adapt well to different types of noisy data, resulting in low efficiency. Furthermore, existing intelligent prediction methods rarely combine with physicochemical kinetic mechanisms, leaving prediction results without theoretical support and failing to accurately reflect the gradual process of corrosion behavior over time. Moreover, for most optimization strategy generation techniques, existing optimization systems often focus only on a single indicator, such as changes in corrosion area or coating thickness, failing to consider the trade-offs between multiple objectives such as corrosion resistance, cost, and resource efficiency. This limits their practicality in actual industrial production. Therefore, existing technologies struggle to form a complete closed-loop system for corrosion performance prediction and optimization, meaning that quantitative optimization and performance improvement of coating corrosion resistance still rely heavily on manual trial-and-error experiments. Summary of the Invention
[0004] In view of the problems existing in the above-mentioned methods and systems for predicting the corrosion resistance of coatings based on reinforcement learning models, this invention is proposed.
[0005] Therefore, this invention provides a method and system for predicting the corrosion resistance of coatings based on a reinforcement learning model, which solves the problems of low data extraction efficiency, inaccurate time series modeling, and single optimization strategy in the prediction of coating corrosion resistance.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a method for predicting the corrosion resistance of coatings based on a reinforcement learning model, which includes collecting data on coatings in a corrosive environment and using the YOLO v5 algorithm to form a structured corrosion performance data table.
[0008] Based on the corrosion performance data table, the corrosion state of time series is generated by interpolation method combined with kinetic equation. The corrosion state of time series is used to train a reinforcement learning model with the improved Double DQN algorithm. The corrosion resistance adjustment strategy is obtained through the reinforcement learning model. Pareto optimal initial population is generated based on the reinforcement learning model. The adjustment strategy is optimized by combining the NSGA-II algorithm.
[0009] The experience data generated during the process of generating and optimizing adjustment strategies is uploaded to the database to form a closed-loop feedback.
[0010] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the reinforcement learning model is trained using a time-series corrosion state combined with an improved Double DQN algorithm. The corrosion resistance adjustment strategy is obtained through the reinforcement learning model, specifically by extracting a state space vector from the time-series corrosion state and environmental state. ;
[0011] Determine the online network and the network parameters it uses. and the target network and the network parameters used. Using a deep neural network as a function approximator, the Q-value of the online network is calculated. Q-value of the target network ;
[0012] Calculate the Q-value Y of the target network at the next time step using the online network. Select the optimal action for the next time step, target network Evaluate the Q value at the next time step;
[0013] Based on the current Q-value and the Q-value at the next time step of the online network, the loss calculation function is obtained. ;
[0014] Set the threshold for the loss function. :
[0015] like Greater than or equal to Calculate the gradient of the loss function. Update parameters based on the gradient of the loss function. ;
[0016] Update the target network parameters every K training iterations. ;
[0017] like Less than Once the reinforcement learning model has been trained, it outputs an adjusted policy based on the current state vector. Final output To adjust the strategy.
[0018] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the method of forming a structured corrosion performance data table using the YOLO v5 algorithm specifically involves using the Labellmg annotation tool to annotate corrosion spots, corrosion areas formed by numerous connections, and zinc-covered areas based on the collected image data, and calculating the corrosion area based on the area of the annotation boxes. With zinc coverage area ;
[0019] The corrosion area ratio of the coating is calculated using the corrosion area, and the final output is structured data containing coating thickness, corrosion time, adhesion, corrosion ratio area, and zinc coverage area at each time point.
[0020] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the step of generating a time-series corrosion state based on a corrosion performance data table by combining an interpolation method with a kinetic equation specifically involves calculating the corrosion area corresponding to time t using a linear interpolation method based on the time points and state variables provided in the corrosion performance data table. With zinc coverage area ;
[0021] Corrosion rate calculated based on kinetic equations With protective capabilities ;
[0022] Based on the calculation results of corrosion area, zinc coverage area, corrosion rate, and protective capability, the corrosion state at a single time point is expressed as follows: ;
[0023] Representing a time series as a set of states at multiple points in time. Final output This represents the erosion state of the time series.
[0024] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the step of collecting data on the coating in a corrosive environment specifically involves collecting coating thickness, corrosion time, adhesion, and image data, and uploading the collected data to a database.
[0025] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the step of generating an initial population based on the reinforcement learning model and optimizing the adjustment strategy using the NSGA-II algorithm specifically involves generating a set of k multi-objective optimization solutions as the initial solution set. ;
[0026] Define the maximum service life of the coating. As the first optimization objective of the adjustment strategy, the erosion area is used. As the second optimization objective of the adjustment strategy, the adhesion measurement value is defined. As the third optimization objective of the adjustment strategy, the initial solution set will be... As the initial population input for the NSGA-II algorithm;
[0027] Initialize the solution set The solutions are sorted in a non-dominated order, and the crowding distance is calculated for all non-dominated solutions after sorting. ;
[0028] Set crowding distance threshold ,like Less than Then the e-th non-dominated solution Add it to the preferred strategy set; otherwise, do not use it.
[0029] Based on the selected set of preferred strategies, crossover and mutation operations are used to generate new sub-strategies. The generated sub-strategy is then integrated with the original optimal strategy to form a new solution set. ;
[0030] Set a maximum number of iterations, Iter. After the maximum number of iterations, Iter, output the strategy with the minimum crowding distance in the solution set. This is the optimal adjustment strategy.
[0031] As a preferred embodiment of the coating corrosion resistance prediction method based on reinforcement learning model described in this invention, the step of uploading the empirical data generated during the generation and optimization of adjustment strategies to the database to form a closed-loop feedback specifically involves recording the coating life, corrosion area, and adhesion data of different coatings under different environments, and uploading the empirical data of generating and optimizing adjustment strategies based on the data to the database to form a closed-loop feedback.
[0032] Secondly, this invention provides a coating corrosion resistance prediction system based on a reinforcement learning model, comprising:
[0033] The data acquisition module is used to collect various data of the coating samples in a corrosive environment;
[0034] The target detection module is used to detect and identify the acquired coating corrosion images using the YOL v5 algorithm and extract key corrosion features;
[0035] The time series generation module is used to generate corrosion states over time using interpolation methods and kinetic equations;
[0036] The reinforcement learning model training module is used to train the reinforcement learning module using the DQN algorithm and output the adjustment policy after training is completed.
[0037] Pareto's initial population generation module is used to generate multiple adjustment policies using a reinforcement learning model;
[0038] The NSGA-II optimization module is used to optimize the policy set and output the optimal adjustment policy;
[0039] The experience storage and learning module is used to save the experience generated during the generation and optimization of strategies and form a closed-loop feedback.
[0040] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the coating corrosion resistance prediction method based on a reinforcement learning model as described in the first aspect of the present invention.
[0041] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the coating corrosion resistance prediction method based on a reinforcement learning model as described in the first aspect of the present invention.
[0042] The beneficial effects of this invention are as follows: It achieves accurate extraction and structured representation of corrosion data through an improved YOLO v5 algorithm, and then generates time-series corrosion state data by combining interpolation methods and kinetic equations. This solves the problem of dynamic complexity and multi-objective trade-offs in optimizing the corrosion resistance performance of coatings. These data are then input into an improved Double DQN algorithm to train a reinforcement learning model to obtain dynamic corrosion resistance adjustment strategies, achieving intelligent, automated, and efficient corrosion performance adjustment. Furthermore, it combines the NSGA-II algorithm to achieve multi-objective optimization, and iteratively optimizes the adjustment strategy through a database closed-loop feedback mechanism, forming a highly efficient data-driven strategy optimization system. This realizes a cyclical system from data acquisition to strategy generation, optimization, and feedback, significantly improving the efficiency and reliability of optimization decisions. Attached Figure Description
[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating the coating corrosion resistance prediction method based on a reinforcement learning model in Example 1.
[0045] Figure 2 This is a schematic diagram of the coating corrosion resistance prediction system based on a reinforcement learning model in Example 1. Detailed Implementation
[0046] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0047] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0048] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0049] Example 1, referring to Figure 1 and Figure 2This is the first embodiment of the present invention, which provides a method for predicting the corrosion resistance of coatings based on a reinforcement learning model. The method for predicting the corrosion resistance of coatings based on a reinforcement learning model includes the following steps:
[0050] S1: Collect data on the coating in a corrosive environment and use the YOLO v5 algorithm to generate a structured corrosion performance data table.
[0051] Specifically, collecting data on coatings in corrosive environments refers to collecting data on coating thickness, corrosion time, adhesion, and image data, and then uploading the collected data to a database.
[0052] By collecting data on coatings in corrosive environments and uploading it to a database, centralized acquisition and standardized storage of coating data information are achieved, providing a foundation for subsequent data-driven optimization. This enhances the transparency and accuracy of coating usage in corrosive environments, thereby improving the responsiveness of corrosion data analysis and strategy development.
[0053] Furthermore, a structured corrosion performance data table is generated using the YOLO v5 algorithm. Based on the acquired image data, the Labellmg annotation tool is used to annotate corrosion spots, corrosion areas formed by numerous connections, and zinc-covered areas. The corrosion area is calculated based on the area of the annotation boxes. With zinc coverage area ;
[0054] The corrosion area is calculated based on the area of the labeled box. :
[0055] ,
[0056] in, The number of corrosion spots is indicated. The number of markings for the corroded areas. Let i be the labeled area of the i-th corrosion spot. Let i be the labeled area of the i-th eroded region;
[0057] The zinc coverage area is calculated based on the area of the labeled box. :
[0058] ,
[0059] in, The number of zinc-covered areas to be marked. Let i be the labeled area of the i-th zinc-covered region;
[0060] The corrosion area ratio of the coating is calculated using the corrosion area, and the final output is structured data containing coating thickness, corrosion time, adhesion, corrosion ratio area, and zinc coverage area at each time point.
[0061] By using the YOLO v5 algorithm to identify targets in the image data of coatings under corrosive environments, the raw and complex unstructured corrosion data is transformed into a structured corrosion performance data table that is easy to analyze, which greatly improves the data processing efficiency and accuracy.
[0062] S2: Based on the corrosion performance data table, the corrosion state of the time series is generated by interpolation method combined with the kinetic equation. The corrosion state of the time series is used to train the reinforcement learning model with the improved Double DQN algorithm. The corrosion resistance adjustment strategy is obtained through the reinforcement learning model. Pareto optimal initial population is generated based on the reinforcement learning model. The adjustment strategy is optimized by combining the NSGA-II algorithm.
[0063] Specifically, based on corrosion performance data tables, a time-series corrosion state index is generated using interpolation methods combined with kinetic equations. This index is then used to calculate the corrosion area corresponding to time t using linear interpolation based on the time points and state variables provided in the corrosion performance data tables. With zinc coverage area :
[0064] Calculate the corrosion area corresponding to time t :
[0065] ,
[0066] Where t is time. For the i-th time, Let i be the next time step after the i-th time step.
[0067] Calculate the zinc coverage area corresponding to time t :
[0068] ,
[0069] Corrosion rate calculated based on kinetic equations With protective capabilities The details are as follows:
[0070] Calculate the corrosion rate based on the dynamic changes in the corrosion area. :
[0071] ,
[0072] in, For time step;
[0073] Calculate protective capacity based on dynamic changes in zinc coverage area. :
[0074] ,
[0075] in, This represents the initial zinc coverage area.
[0076] The corrosion state at a single point in time can be represented as :
[0077] ,
[0078] in, Let be the corrosion area at time t. For corrosion rate, Let be the zinc coverage area at time t. The protection capability at time t;
[0079] Representing a time series as a set of states at multiple points in time. :
[0080] ,
[0081] Output This represents the erosion state of the time series.
[0082] By using structured corrosion performance data as a foundation, interpolation methods are used to generate corrosion state data with time-series characteristics. This solves the problem of discontinuous time dimension in corrosion experimental data, making the corrosion data more complete and able to more accurately reflect the dynamic corrosion process. At the same time, by using kinetic equations, the time-series data is made more consistent with the corrosion mechanism, providing key input variables for subsequent reinforcement learning models.
[0083] Furthermore, a reinforcement learning model is trained using time-series corrosion states combined with an improved Double DQN algorithm. The model then extracts the state space vector from the time-series corrosion states and environmental states to obtain the corrosion resistance adjustment strategy. :
[0084] ,
[0085] in, For ambient temperature, For ambient humidity, Environmental salt concentration;
[0086] Define action space :
[0087] ,
[0088] in, The strategy has been adjusted as follows:
[0089] ,
[0090] in, This represents the zinc powder content value. This is the coating thickness value. This refers to the proportion of specific additives, which include, but are not limited to, dispersants, corrosion inhibitors, antioxidants, and stabilizers.
[0091] Define reward function :
[0092] ,
[0093] in, This is the maximum service life of the coating. For the corrosion area, This is the adhesion measurement value. , The weighting coefficients are set;
[0094] Construct an online network and a target network, and define the neural network parameters. and ,and Using deep neural networks as function approximators, the Q-values of the two networks are calculated. and :
[0095] The method uses a deep neural network as a function approximator to calculate the Q-value of the online network. :
[0096] ,
[0097] Here, DNN refers to a multi-layered neural network that maps input variables to Q-values. Network parameters used for online networks;
[0098] The method uses a deep neural network as a function approximator to calculate the Q value of the target network. :
[0099] ,
[0100] in, Network parameters used for the target network;
[0101] Calculate the Q-value Y of the target network at the next time step: ,
[0102] in, For instant rewards, As a discount factor, This is the state space vector for the next time step. The action used by the Q network in the next time step is determined by the online network. Select the optimal action for the next time step, target network Evaluate the Q value at the next time step;
[0103] Based on the current Q-value and the Q-value at the next time step of the online network, the loss calculation function is obtained. :
[0104] ,
[0105] In each training step, set the probability... Explore by selecting random actions from the action space, with probability. Choose the action with the highest Q value;
[0106] Set the threshold for the loss function. :
[0107] like Greater than or equal to Calculate the gradient of the loss function. :
[0108] ,
[0109] Then update the parameters based on the gradient of the loss function:
[0110] ,
[0111] in, The learning rate weights are set;
[0112] Update the target network parameters every K training iterations. :
[0113] ,
[0114] Repeatedly train the reinforcement learning model until the loss function is less than the threshold.
[0115] like Less than This indicates that the reinforcement learning model training is complete, and the latest online network parameters are being used. Adjust the strategy based on the current state vector output. :
[0116] ,
[0117] Final output To adjust the strategy.
[0118] By using the erosion state of the time series as the state input of the reinforcement learning model, the reinforcement learning model is trained. The Double DQN algorithm is used to overcome the defect of the traditional DQN algorithm that is prone to overestimating the Q value, making the policy evaluation more accurate. By separating the process of choosing the action and evaluating the target value, the problem of complex state transition and high noise in the erosion scenario is solved. Through multiple rounds of training, the reinforcement learning model can explore autonomously and generate high-quality adjustment policies, solving the problem of traditional methods relying on human experience for adjustment.
[0119] Furthermore, based on the reinforcement learning model, a Pareto-optimal initial population is generated. Combined with the NSGA-II algorithm, the policy is optimized and adjusted through multiple interactions using the reinforcement learning model to generate a set of k multi-objective optimization solutions as the initial solution set. :
[0120] ,
[0121] Maximum service life of the coating As the first optimization objective of the adjustment strategy, the erosion area is used. As the second optimization objective of the adjustment strategy, adhesion measurement values were used. As the third optimization objective of the adjustment strategy, the solution set As the initial population input for the NSGA-II algorithm;
[0122] The solution set Perform a non-dominated sort on the solutions in the given set, and for any two solutions... , ,like:
[0123] ,
[0124] but non-dominance h is the h-th optimization objective;
[0125] Calculate the congestion distance for all non-dominated solutions. :
[0126] ,
[0127] in, For the e-th solution among all non-dominated solutions, , Let f(x) be the solution to the left and the solution to the right of the non-dominated solution that is adjacent to the e-th solution. Let be the value of the solution to the left of the e-th solution in the non-dominated solution, in the h-th optimization objective. Let be the value of the solution to the right of the e-th solution in the non-dominated solution at the h-th optimization objective, and g be the total number of optimization objectives. , These are the maximum and minimum values of the h-th optimization objective, respectively;
[0128] Set crowding distance threshold ,like Less than Then the e-th non-dominated solution Add it to the preferred strategy set; otherwise, do not use it.
[0129] Based on the selected set of optimal strategies, crossover and mutation operations are performed to randomly select two strategies from the set of optimal strategies. , By swapping parts of the two policies and adjusting their values, a new sub-policy is generated. Then, an indefinite number of optimal strategies are randomly selected, and the original strategy values are randomly adjusted to generate new sub-strategies. The generated sub-strategies are then integrated with the original optimal strategies to form a new solution set. ;
[0130] Set the maximum number of iterations Iter for the solution set. The adjustment strategy in the algorithm performs non-dominated sorting. Crowding distance is calculated for the sorted non-dominated strategies. Policies with crowding distances less than a threshold are added to the new solution set. Before the iteration count reaches the threshold, new sub-policies are repeatedly generated and non-dominated sorted. After the iteration count reaches the upper limit (Iter), the strategy with the smallest crowding distance in the solution set is output. This is the optimal adjustment strategy.
[0131] By combining reinforcement learning models with multi-objective optimization techniques, this approach not only absorbs the predictive ability of reinforcement learning for dynamic erosion adjustment but also utilizes multi-objective optimization algorithms to further enhance the ability of decision-making schemes to meet complex practical needs. Multiple initial adjustment schemes are generated through Pareto optimal sets, and the adjustment strategy is further optimized by combining the NSGA-II algorithm, enabling the system to find multiple trade-off adjustment schemes and overcoming the bottleneck of limitations of single-objective schemes.
[0132] S3: Upload the experience data generated during the process of generating and optimizing adjustment strategies to the database to form a closed-loop feedback.
[0133] Specifically, the experience data generated during the process of generating and optimizing adjustment strategies will be uploaded to the database to form a closed-loop feedback. This means recording the coating life, corrosion area, and adhesion data of different coatings under different environments, and uploading the experience data for generating and optimizing adjustment strategies based on the data to the database to form a closed-loop feedback.
[0134] By uploading key data such as adjustment strategies, reward values, and erosion state evolution generated during the optimization process to the database, and forming a knowledge base from the experience in reinforcement learning and multi-objective optimization through a closed-loop feedback mechanism, the system reduces redundant training and computation, enabling more efficient iteration in subsequent optimizations.
[0135] This embodiment also provides a coating corrosion resistance prediction system based on a reinforcement learning model, including:
[0136] The data acquisition module is used to collect various data of the coating samples in a corrosive environment;
[0137] The target detection module is used to detect and identify the acquired coating corrosion images using the YOL v5 algorithm and extract key corrosion features;
[0138] The time series generation module is used to generate corrosion states over time using interpolation methods and kinetic equations;
[0139] The reinforcement learning model training module is used to train the reinforcement learning module using the DQN algorithm and output the adjustment policy after training is completed.
[0140] Pareto's initial population generation module is used to generate multiple adjustment policies using a reinforcement learning model;
[0141] The NSGA-II optimization module is used to optimize the policy set and output the optimal adjustment policy;
[0142] The experience storage and learning module is used to save the experience generated during the generation and optimization of strategies and form a closed-loop feedback.
[0143] This embodiment also provides a computer device applicable to a coating corrosion resistance prediction method based on a reinforcement learning model, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the coating corrosion resistance prediction method based on a reinforcement learning model as proposed in the above embodiment.
[0144] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0145] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the reinforcement learning model-based method for predicting the corrosion resistance of coatings as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0146] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for predicting the corrosion resistance of coatings based on a reinforcement learning model, characterized in that: include, Data on coatings in corrosive environments were collected, and a structured corrosion performance data table was generated using the YOLO v5 algorithm. Based on the corrosion performance data table, the corrosion state of time series is generated by interpolation method combined with kinetic equation. The corrosion state of time series is used to train a reinforcement learning model with the improved Double DQN algorithm. The adjustment strategy of corrosion resistance is obtained through the reinforcement learning model. Pareto optimal initial population is generated based on the reinforcement learning model. The adjustment strategy is optimized by combining the NSGA-II algorithm. The experience data generated during the process of generating and optimizing adjustment strategies is uploaded to the database to form a closed-loop feedback. The method of using time-series corrosion states combined with an improved Double DQN algorithm to train a reinforcement learning model, and obtaining corrosion resistance adjustment strategies through the reinforcement learning model, refers to extracting state space vectors from the time-series corrosion states and environmental states. ; Define action space With reward function ; Construct an online network (Online) and a target network (Target), and use them respectively. and As network parameters, a deep neural network is used as a function approximator to calculate the Q-values of the two networks. and ; Calculate the Q-value Y of the target network at the next time step using the online network. Select the optimal action for the next time step, target network Evaluate the Q value at the next time step; Based on the current Q-value and the Q-value at the next time step of the online network, the loss calculation function is obtained. ; In each training step, set the probability... Explore by selecting random actions from the action space, with probability. Choose the action with the highest Q value; Set the threshold for the loss function. : like Greater than or equal to Calculate the gradient of the loss function. Update parameters based on the gradient of the loss function. ; Update the target network parameters every K training iterations. ; Repeatedly train the reinforcement learning model until the loss function is less than the threshold. like Less than This indicates that the reinforcement learning model training is complete, and the latest online network parameters are being used. Adjust the strategy based on the current state vector output. Final output To adjust the strategy; The process of generating a Pareto-optimal initial population based on a reinforcement learning model and optimizing the adjustment strategy using the NSGA-II algorithm involves multiple interactions through the reinforcement learning model to generate a set of k multi-objective optimization solutions as the initial solution set. ; Will Corrosion area , These were respectively used as the first, second, and third optimization objectives of the adjustment strategy. ,Will As the initial population input for NSGA-II; The solution set The solutions are sorted in a non-dominated order, and the crowding distance is calculated for all non-dominated solutions after sorting. ; Set crowding distance threshold ,like Less than Then the plan Add it to the preferred strategy set; otherwise, do not use it. Based on the selected set of optimal strategies, crossover and mutation operations are performed to randomly select two strategies from the set of optimal strategies. , By swapping parts of the two policies and adjusting their values, a new sub-policy is generated. Then, an indefinite number of optimal strategies are randomly selected, and the original strategy values are randomly adjusted to generate new sub-strategies. The generated sub-strategies are then integrated with the original optimal strategies to form a new solution set. ; Set the maximum number of iterations Iter for the solution set. The adjustment strategy in the algorithm performs non-dominated sorting. Crowding distance is calculated for the sorted non-dominated strategies. Policies with crowding distances less than a threshold are added to the new solution set. Before the iteration count reaches the threshold, new sub-policies are repeatedly generated and non-dominated sorted. After the iteration count reaches the upper limit (Iter), the strategy with the smallest crowding distance in the solution set is output. This is the optimal adjustment strategy.
2. The coating corrosion resistance prediction method based on reinforcement learning model as described in claim 1, characterized in that: The aforementioned use of the YOLO v5 algorithm to form a structured corrosion performance data table refers to using the Labellmg annotation tool to annotate corrosion spots, corrosion areas formed by numerous connections, and zinc-covered areas based on the acquired image data, and calculating the corrosion area based on the area of the annotation boxes. With zinc coverage area ; The corrosion area ratio of the coating is calculated using the corrosion area, and the final output is structured data containing coating thickness, corrosion time, adhesion, corrosion ratio area, and zinc coverage area at each time point.
3. The coating corrosion resistance prediction method based on reinforcement learning model as described in claim 2, characterized in that: The corrosion state data generated based on the corrosion performance data table, using interpolation methods combined with kinetic equations, refers to calculating the corrosion area corresponding to time t using linear interpolation based on the time points and state variables provided in the corrosion performance data table. With zinc coverage area ; Corrosion rate calculated based on kinetic equations With protective capabilities ; Based on the calculation results of corrosion area, zinc coverage area, corrosion rate, and protective capability, the corrosion state at a single time point is expressed as follows: ; Representing a time series as a set of states at multiple points in time. Final output This represents the erosion state of the time series.
4. The coating corrosion resistance prediction method based on reinforcement learning model as described in claim 3, characterized in that: The data collected on the coating in the corrosive environment refers to the collection of coating thickness, corrosion time, adhesion, and image data, and the collected data is uploaded to the database.
5. The coating corrosion resistance prediction method based on reinforcement learning model as described in claim 4, characterized in that: The process of generating and optimizing adjustment strategies by uploading the experience data generated during the process to the database to form a closed-loop feedback refers to recording the coating life, corrosion area, and adhesion data of different coatings under different environments, and uploading the experience of generating and optimizing adjustment strategies based on the data to the database to form a closed-loop feedback.
6. A coating corrosion resistance prediction system based on a reinforcement learning model, based on the coating corrosion resistance prediction method based on a reinforcement learning model as described in any one of claims 1 to 5, characterized in that: include, The data acquisition module is used to collect various data of the coating samples in a corrosive environment; The target detection module is used to detect and identify the acquired coating corrosion images using the YOL v5 algorithm and extract key corrosion features; The time series generation module is used to generate corrosion states over time using interpolation methods and kinetic equations; The reinforcement learning model training module is used to train the reinforcement learning module using the DQN algorithm and output the adjustment policy after training is completed. Pareto's initial population generation module is used to generate multiple adjustment policies using a reinforcement learning model; The NSGA-II optimization module is used to optimize the policy set and output the optimal adjustment policy; The experience storage and learning module is used to save the experience generated during the generation and optimization of strategies and form a closed-loop feedback.
7. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the coating corrosion resistance prediction method based on the reinforcement learning model as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the coating corrosion resistance prediction method based on the reinforcement learning model as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Aluminum product quality multi-scale detection method and system
CN120489987A
Building control system using reinforcement learning
US20230168649A1