Model-driven lithium hydroxide production autonomous control method and system

The autonomous control system based on the model-driven multi-agent DRL algorithm solves the problems of low efficiency, high energy consumption and unstable quality in traditional lithium hydroxide production, and achieves efficient and stable production process optimization and product quality improvement.

CN118113005BActive Publication Date: 2025-11-11江西协成锂业有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410241547.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-11-11
Estimated Expiration
2044-03-04

AI Technical Summary

Technical Problem

Traditional lithium hydroxide production processes suffer from problems such as low efficiency, high energy consumption, and unstable product quality, mainly due to reliance on manual experience for control and the limitations of reactors.

Method used

A model-driven multi-agent deep reinforcement learning (DRL) algorithm is used to construct an autonomous control system for lithium hydroxide production. By training agents in a simulated environment, the system can adjust input information such as reaction conditions and raw material ratios in real time to achieve autonomous decision-making and optimization.

Benefits of technology

It improves production efficiency and stability, reduces human error, optimizes resource utilization, lowers production costs, enhances product quality and consistency, increases adaptability and flexibility, and reduces risks and losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118113005B_ABST
    Figure CN118113005B_ABST
Patent Text Reader

Abstract

This application discloses a model-driven autonomous control method and system for lithium hydroxide production. The method includes: constructing a system framework; adaptively setting a multi-agent DRL algorithm according to preset requirements, the multi-agent DRL algorithm comprising several agents, a decision-making mechanism for each agent, and interaction strategies between different agents; conducting centralized training of the agents in a simulated environment; the simulated environment being a virtual lithium hydroxide production environment; deploying each trained agent to various key stages of the lithium hydroxide production line; collecting initial data in real time during the production process and sending this initial data to the agents; and the corresponding agents adjusting the input information in lithium hydroxide production in real time based on the multi-agent DRL algorithm. This method can effectively improve the efficiency and stability of lithium hydroxide production, reduce human error, and achieve autonomous real-time adjustment and decision-making, thereby improving production quality and economic benefits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of lithium hydroxide production technology, and in particular to a model-driven autonomous control method and system for lithium hydroxide production. Background Technology

[0002] Lithium hydroxide, as an important chemical, has wide applications in battery manufacturing, aerospace, and pharmaceuticals. Traditional lithium hydroxide production processes are typically based on experience-driven control, with operators making adjustments based on experience and observation. This method is subjective and limited by the operator's experience. Furthermore, traditional reactors are often restrictive and energy-intensive. Therefore, traditional lithium hydroxide production processes suffer from low efficiency, high energy consumption, and unstable product quality. Summary of the Invention

[0003] In view of this, the present disclosure provides a model-driven autonomous control method and system for lithium hydroxide production, which can solve the problems of low efficiency, high energy consumption and unstable product quality in existing lithium hydroxide production processes.

[0004] In a first aspect, embodiments of this disclosure provide a model-driven autonomous control method for lithium hydroxide production, including:

[0005] Build the system framework;

[0006] In the system framework, a multi-agent DRL algorithm is adaptively set according to preset requirements. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents.

[0007] Several of the aforementioned agents are trained in a simulated environment, which is a virtual lithium hydroxide production environment.

[0008] Each trained agent is deployed to the lithium hydroxide production line;

[0009] Several initial data points are collected in real time during the production process, and the initial data are sent to the corresponding intelligent agent;

[0010] The corresponding intelligent agent is based on the multi-agent DRL algorithm and adjusts the input information in lithium hydroxide production in real time.

[0011] Optionally, the construction of the system framework includes establishing a model of the lithium hydroxide production process:

[0012] The establishment of the lithium hydroxide production process model includes:

[0013] Set preset control targets; the preset control targets include one or more of the following: production efficiency indicators, energy consumption indicators, and product quality indicators;

[0014] A lithium hydroxide production process model is established based on the preset control objectives.

[0015] Optionally, establishing the lithium hydroxide production process model based on the preset control target includes:

[0016] Collect first data; the first data includes one or more of the following: raw material characteristic data, reaction condition data, and product quality index data;

[0017] The first data is preprocessed to obtain the second data;

[0018] The second data is analyzed according to a preset strategy to obtain the third data;

[0019] The preset mathematical model is trained based on the third data, and the trained model is the established lithium hydroxide production process model.

[0020] Optionally, the step of analyzing the second data according to a preset strategy to obtain the third data includes:

[0021] A descriptive statistical analysis strategy was used to conduct a preliminary analysis of the second data to obtain data distribution information;

[0022] The Pearson correlation coefficient strategy is used to evaluate the data distribution information and obtain the strength of the linear relationship between different data.

[0023] The strength of the linear relationship was analyzed using statistical testing methods to identify key factors that significantly affect production results; these key factors are those whose impact on the production process is no less than a preset threshold.

[0024] The key factors are filtered using machine learning algorithms to obtain the third data;

[0025] The third set of data is a subset of features that meets the model prediction criteria.

[0026] Optionally, the step of training a preset mathematical model based on the third data, wherein the trained model is the established lithium hydroxide production process model, includes:

[0027] Analyze the data characteristics of the first data;

[0028] Based on the data characteristics and prediction objectives, an initial regression model is selected;

[0029] The initial regression model is optimized and trained using historical data to obtain the first model;

[0030] The first model is validated and trained using preset indicators. The model that meets the validation requirements is the established lithium hydroxide production process model.

[0031] Optionally, the preset indicators include one or more of the following: cross-validation indicators, mean square error indicators, and coefficient of determination indicators.

[0032] Optionally, the step of adaptively setting a multi-agent DRL algorithm according to preset requirements in the system framework includes:

[0033] Several intelligent agents can be flexibly set up according to preset requirements; each of the intelligent agents corresponds one-to-one with the functional agents required in the lithium hydroxide production process.

[0034] A decision-making mechanism is set for each of the aforementioned agents;

[0035] The interaction strategies between the different intelligent agents are set; the interaction strategies include communication strategies and cooperation strategies.

[0036] Optionally, the intelligent agents include a reaction control intelligent agent, a temperature monitoring intelligent agent, and a pressure control intelligent agent;

[0037] The decision-making mechanism includes a reaction control mechanism, a temperature monitoring mechanism, and a pressure control intelligent mechanism.

[0038] Optionally, the communication strategy includes communicating via a local area network or via a cloud platform.

[0039] Secondly, embodiments of this disclosure also provide a model-driven autonomous control system for lithium hydroxide production, comprising:

[0040] The build module is configured to build the system framework;

[0041] The configuration module is configured to adaptively set a multi-agent DRL algorithm according to preset requirements in the system framework. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents.

[0042] The training module is configured to perform centralized training on several of the aforementioned agents in a simulated environment; the simulated environment is a virtual lithium hydroxide production environment.

[0043] The deployment module is configured to deploy each trained agent to the lithium hydroxide production line.

[0044] The data acquisition module is configured to collect several initial data points during the production process in real time and send the initial data points to the corresponding intelligent agent.

[0045] The execution module is configured to adjust the input information in lithium hydroxide production in real time based on the multi-agent DRL algorithm; the input information includes reaction conditions and optimized raw material ratio.

[0046] Thirdly, this disclosure also provides an electronic device that adopts the following technical solution:

[0047] The electronic device includes:

[0048] At least one processor; and,

[0049] A memory communicatively connected to the at least one processor; wherein,

[0050] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to execute any of the model-driven autonomous control methods for lithium hydroxide production described above.

[0051] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the model-driven autonomous control methods for lithium hydroxide production described above.

[0052] The model-driven autonomous control method for lithium hydroxide production disclosed in this application, through real-time adjustment of a multi-agent DRL algorithm, can automatically adjust the input information in the lithium hydroxide production line according to actual conditions, such as reaction conditions and optimized raw material ratios, thereby optimizing the production process and improving production efficiency. Using an agent model for autonomous control reduces the impact of human factors on the production process, lowers the possibility of human error, and improves production stability and consistency. By centrally training the agents in a simulation environment, they can provide feedback and adjustments based on real-time collected production data, thus adapting to changes in the production environment and making corresponding decisions, improving the adaptability and flexibility of production. Based on its own decision-making mechanism and interaction strategies with other intelligent agents, the intelligent agent can make autonomous decisions according to the system framework and preset requirements, thereby achieving autonomous control and reducing reliance on human intervention. By optimizing reaction conditions and raw material ratios in real time, it can utilize resources and raw materials more effectively, reduce waste, and improve production utilization, thus saving production costs. The intelligent agent can make intelligent decisions based on real-time data, promptly identify potential problems, and take corresponding measures, thereby reducing risks and possible losses in the production process. By adjusting the input information in the production process in real time, it can more accurately control reaction conditions and raw material ratios, which helps to improve product quality and consistency and meet customer needs and standards.

[0053] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a schematic diagram of the model-driven autonomous control method for lithium hydroxide production in this application.

[0056] Figure 2 A flowchart illustrating the method for establishing a lithium hydroxide production process model provided in this embodiment of the disclosure.

[0057] Figure 3 A schematic diagram illustrating the process of establishing a lithium hydroxide production process model based on a preset control target, provided for embodiments of this disclosure.

[0058] Figure 4 for Figure 3 A flowchart illustrating the method for obtaining third-party data.

[0059] Figure 5 for Figure 3 A schematic diagram of the process of training a preset mathematical model based on third-party data.

[0060] Figure 6 for Figure 1 A schematic diagram of the adaptive setting method of the multi-agent DRL algorithm.

[0061] Figure 7 A schematic diagram of the model-driven autonomous control system for lithium hydroxide production provided in this embodiment of the disclosure.

[0062] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0063] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0064] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0065] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0066] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0067] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0068] Reference Figure 1 The first aspect of this application discloses a model-driven autonomous control method for lithium hydroxide production, which includes the following steps:

[0069] S100, constructing the system framework.

[0070] In this embodiment, building the system framework includes establishing a model of the lithium hydroxide production process.

[0071] S200, within the system framework, adaptively sets a multi-agent DRL algorithm according to preset requirements. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents.

[0072] The S300 performs centralized training on several agents in a simulated environment, which is a virtual lithium hydroxide production environment.

[0073] The S400 deploys each trained agent to key stages of the lithium hydroxide production line.

[0074] The S500 collects several initial data points during the production process in real time and sends these initial data points to the corresponding intelligent agents.

[0075] The S600, whose corresponding agent is based on the multi-agent DRL algorithm, adjusts the input information in lithium hydroxide production in real time.

[0076] The input information includes reaction conditions and optimized raw material ratios.

[0077] The model-driven autonomous control method for lithium hydroxide production disclosed in this application, through real-time adjustment using a multi-agent DRL (deep reinforcement learning) algorithm, can automatically adjust input information in the lithium hydroxide production line according to actual conditions, such as reaction conditions and optimized raw material ratios, thereby optimizing the production process and improving production efficiency. Using an agent model for autonomous control reduces the impact of human factors on the production process, lowers the possibility of human error, and improves production stability and consistency. By training the agents intensively in a simulated environment, they can provide feedback and adjustments based on real-time collected production data, thus adapting to changes in the production environment and making corresponding decisions, improving the adaptability and flexibility of production. Activity; Each agent, based on its own decision-making mechanism and interaction strategies with other agents, can make autonomous decisions according to the system framework and preset requirements, thereby achieving autonomous control and reducing reliance on human intervention; By optimizing reaction conditions and raw material ratios in real time, resources and raw materials can be utilized more effectively, reducing waste and improving production utilization, thus saving production costs; Agents can make intelligent decisions based on real-time data, promptly identify potential problems and take corresponding measures, thereby reducing risks and possible losses in the production process; By adjusting input information in the production process in real time, reaction conditions and raw material ratios can be controlled more precisely, helping to improve product quality and consistency and meet customer needs and standards.

[0078] This application, through the configuration of a multi-agent DRL algorithm, enables the deployment of agents at different operation points, allowing the system to make more accurate and efficient decisions in complex production environments.

[0079] Reference Figure 2 The method for establishing a lithium hydroxide production process model includes the following steps:

[0080] S110, Set preset control targets; preset control targets include one or more of the following: production efficiency indicators, energy consumption indicators, and product quality indicators.

[0081] S120, establish a lithium hydroxide production process model based on preset control targets.

[0082] In this embodiment, by defining preset control targets and establishing corresponding production process models, it is possible to better understand the relationship between various indicators of the production process and control targets, thus helping to optimize production control strategies. Establishing models can help identify potential production efficiency bottlenecks and provide targeted improvement solutions, thereby improving the achievement of production efficiency indicators. Optimizing models can help reduce energy consumption indicators, guide energy-saving measures in the production process, optimize energy utilization, and reduce energy waste. By establishing models, it is possible to better monitor factors that may affect product quality during the production process and make timely adjustments and optimizations, thereby improving the compliance rate and stability of product quality indicators. Establishing models can help utilize historical and real-time data for analysis, providing data-supported decision-making basis and making production process decisions more scientific and reliable.

[0083] Reference Figure 3 The method for establishing a lithium hydroxide production process model based on preset control targets includes the following steps:

[0084] S121, Collect first data; the first data includes one or more of the following in the lithium hydroxide production process: raw material characteristic data, reaction condition data, and product quality index data.

[0085] The reaction conditions include temperature, pressure, and time.

[0086] S122, preprocess the first data to obtain the second data.

[0087] Preprocessing preferably includes cleaning and formatting the collected data, removing outliers and noisy data, and ensuring data quality.

[0088] Perform data standardization or normalization to facilitate subsequent analysis.

[0089] S123, Analyze the second data according to the preset strategy to obtain the third data.

[0090] In this embodiment, the focus is on extracting features that significantly impact the production process, such as reactant concentration and reaction temperature changes. Specifically, statistical analysis methods or machine learning techniques can be used for feature selection to identify the most influential factors.

[0091] S124, The preset mathematical model is trained based on the third data, and the trained model is the established lithium hydroxide production process model.

[0092] The preset mathematical models include regression models, neural networks, etc.

[0093] In this embodiment, by collecting data on raw material characteristics, reaction conditions, and product quality indicators, the established model can more accurately reflect the actual situation and key parameters of the production process, improving the model's reliability and predictive ability. Preprocessing the collected data, such as removing noise, filling missing values, and handling outliers, can improve data quality and reduce interference and errors during model training. Analyzing the preprocessed data can reveal potential correlations and influencing factors between data points, helping to determine suitable mathematical models and algorithms for better simulation and prediction of the lithium hydroxide production process. Based on this third data, appropriate mathematical models and machine learning algorithms can be used for model training, enabling the model to learn and adapt to the characteristics and changes of the production process, improving its accuracy and predictive ability. The established lithium hydroxide production process model can be used for real-time prediction and optimization decisions. By inputting real-time data, the model can output corresponding prediction results and optimization suggestions, helping to make real-time adjustments and decisions, thereby improving production efficiency and product quality.

[0094] In this embodiment, selecting a suitable preset mathematical model to simulate the lithium hydroxide production process is crucial. This process involves considering the characteristics of the data, the applicability of the model, and the accuracy of the prediction.

[0095] Reference Figure 4 The method for obtaining third-party data includes the following steps:

[0096] S1231, Descriptive statistical analysis strategy is used to conduct preliminary analysis of the second data to obtain data distribution information.

[0097] Descriptive statistical analysis strategies include those based on the mean, standard deviation, and quantiles, with the aim of understanding the basic distribution of the data.

[0098] S1232 uses the Pearson correlation coefficient strategy to evaluate the data distribution information and obtain the strength of the linear relationship between different data.

[0099] This step is used to identify factors that are highly correlated with key indicators such as production efficiency and product quality.

[0100] S1233, statistical test methods are used to analyze the strength of linear relationship and obtain key factors that significantly affect production results; key factors that significantly affect production results are factors whose degree of influence on the production process is not less than the preset degree threshold.

[0101] Among them, statistical testing methods include ANOVA (analysis of variance) or chi-square, which aim to assess the degree of influence of different characteristics on the production process.

[0102] S1234, using machine learning algorithms to screen key factors to obtain third data; the third data is a subset of features that meet the model prediction criteria.

[0103] Among them, machine learning algorithms include random forests, gradient boosting trees, etc., with the aim of selecting the subset of features that are most helpful for model prediction.

[0104] In this embodiment, descriptive statistical analysis, Pearson correlation coefficient strategy, and statistical testing methods can comprehensively understand the distribution of the original data, the linear relationships between different data, and the key factors that significantly affect production results. Analyzing the strength of linear relationships using statistical testing methods can accurately identify factors whose impact on the production process is no less than a preset threshold, helping to precisely identify key factors affecting production results. Machine learning algorithms are used to screen key factors, obtaining third data—a subset of features that meet the model's prediction criteria—which can optimize the original data and extract the most influential feature subset. By obtaining this third data, more accurate production and prediction models can be established, providing more reliable data support for enterprise decision-making, thereby improving production efficiency and quality. The acquisition of third data can also help enterprises discover new business opportunities, optimize production processes, and enhance competitiveness.

[0105] Specifically, suppose we have a set of data, including parameters such as raw material ratio, reaction temperature, reaction time, and product purity.

[0106] First, we identify the parameters that are highly correlated with product purity by calculating the Pearson correlation coefficient between each parameter and product purity.

[0107] Then, ANOVA was used to test whether the relationship between these parameters and product purity was significant.

[0108] Finally, a random forest model is used to rank all features by importance, and the most important features are selected for subsequent model training.

[0109] Through this process, we can effectively identify the factors that have the greatest impact on the lithium hydroxide production process, providing support for building more accurate predictive models and optimizing production processes.

[0110] Reference Figure 5 The method for training a pre-defined mathematical model based on third-party data includes the following steps:

[0111] S1241, Analyze the data characteristics of the first data.

[0112] In this embodiment, the characteristics of the data collected during the lithium hydroxide production process are first analyzed, including the dimensions, distribution, and missing values ​​of the data.

[0113] For example, suppose we collect data such as raw material ratios, reaction temperature, pressure, and the final yield and purity as response variables.

[0114] S1242, Select the initial regression model based on the data characteristics and prediction objectives.

[0115] Common choices include linear regression, multinomial regression, and decision tree regression.

[0116] For example, if the data shows a linear relationship between raw material ratios and output, a linear regression model can be tried first.

[0117] S1243 uses historical data to optimize and train the initial regression model to obtain the first model.

[0118] For example, in a linear regression model, overfitting can be prevented by adjusting the regularization term.

[0119] S1244, Use preset indicators to verify and train the first model. The model that meets the verification requirements is the established lithium hydroxide production process model.

[0120] The preset indicators include one or more of the following: cross-validation indicators, mean square error indicators, and coefficient of determination indicators.

[0121] In this embodiment, a lithium hydroxide production process model is established by training on third-party data, enabling accurate prediction of future production outcomes. This helps enterprises make rational decisions during production, improving production efficiency and output. Optimizing the initial regression model using historical data enhances its accuracy and reliability. This makes the final production process model more closely reflect reality and better explain and predict production results. Validation training using preset indicators assesses the model's quality and reliability. Using preset indicators such as cross-validation, mean squared error, and coefficient of determination ensures the model fits and predicts data well, improving accuracy and reliability. After establishing the lithium hydroxide production process model, enterprises can make decisions based on the model's predictions. This reduces the influence of subjective factors on decision-making, making it more scientific and data-driven, thus improving production efficiency and reducing risks.

[0122] Furthermore, if the predictive performance of the linear model is found to be poor, more complex models, such as random forest regression or support vector machine regression, can be tried.

[0123] Specifically, suppose our goal is to predict the yield under specific feedstock ratios and reaction conditions.

[0124] First, attempt to predict using a linear regression model. If the model predicts R... 2 A lower value may mean that the relationship is not entirely linear.

[0125] Next, we will try using a multinomial regression model to capture more complex nonlinear relationships.

[0126] The new model was trained using the same historical dataset and its performance was evaluated using cross-validation.

[0127] If the performance of the multinomial regression model is better than that of the linear model, it means that it is more suitable for these data.

[0128] Reference Figure 6 The method for adaptive setting of the multi-agent DRL algorithm includes the following steps:

[0129] S210 allows for the flexible configuration of several intelligent agents based on preset requirements; each intelligent agent corresponds one-to-one with the functional agents required in the lithium hydroxide production process.

[0130] S220 sets up a decision-making mechanism for each agent.

[0131] S230 sets the interaction strategy between different intelligent agents; the interaction strategy includes communication strategy and cooperation strategy.

[0132] In this embodiment, multiple agents can be flexibly configured according to preset requirements, each corresponding to a function required in the lithium hydroxide production process. This flexible configuration adapts to different production environments and needs, improving the system's adaptability and flexibility. Setting a decision-making mechanism for each agent allows them to play different roles within the system, making decisions and taking actions according to predetermined strategies, thus improving collaborative work and overall system performance. Setting interaction strategies between different agents, including communication and collaboration strategies, promotes information sharing and collaboration, leading to more efficient overall decision-making and actions, which enhances system synergy and efficiency. The adaptive configuration method of the multi-agent DRL algorithm continuously adjusts and optimizes the agent settings, decision-making mechanisms, and interaction strategies based on actual conditions and needs, enabling adaptive management and optimization of the system. This helps the system respond flexibly to different situations and requirements, improving its robustness and effectiveness.

[0133] This solution enables flexible, optimized, collaborative, and adaptive management of multi-agent systems, which helps improve the overall performance and adaptability of the system and brings more efficient intelligent management and control to the lithium hydroxide production process.

[0134] In this embodiment, the plurality of intelligent agents preferably includes a reaction control intelligent agent, a temperature monitoring intelligent agent, and a pressure control intelligent agent; the decision-making mechanism preferably includes a reaction control mechanism, a temperature monitoring mechanism, and a pressure control intelligent mechanism.

[0135] In this embodiment, the communication strategy includes communicating via a local area network or via a cloud platform to achieve data sharing and command transmission.

[0136] For collaboration strategies between different agents, for example, in certain situations, a temperature monitoring agent needs to work in conjunction with a pressure control agent to ensure production safety.

[0137] In computer science and artificial intelligence, an agent is generally defined as an entity that can perceive its environment and make decisions based on the perceived information to achieve a specific goal. Agents can learn to optimize their decision-making processes to better accomplish their tasks.

[0138] Furthermore, the specific settings for several intelligent agents include:

[0139] S211 is a smart body structure designed for lithium hydroxide production processes.

[0140] Specifically, based on the characteristics of the lithium hydroxide production process, the role of each agent is defined, such as a reaction control agent and a temperature monitoring agent. The function and responsibility of each agent are clearly defined; for example, the temperature monitoring agent is responsible for monitoring and adjusting the reaction temperature in real time.

[0141] S212, Design the decision-making mechanism of the intelligent agent.

[0142] Decision-making algorithms are designed for each agent to enable it to react based on environmental information. For example, reinforcement learning algorithms can be used to enable agents to learn optimal temperature regulation strategies.

[0143] Ensure that intelligent agents can handle uncertainty and environmental changes, and possess a certain degree of adaptability and flexibility.

[0144] S213, Communication and cooperation between intelligent agents.

[0145] In this embodiment, through a communication mechanism, agents can exchange and share information, including environmental status and decision results, which helps improve the information volume and accuracy of the entire system, enabling each agent to have a more comprehensive understanding of the entire production process. Communication and collaboration between agents can facilitate collective decision-making, with all agents jointly negotiating and making decisions to achieve the goal of overall optimization. Through collaborative decision-making, conflicts and incoordination between different agents can be avoided, improving the overall performance and effectiveness of the system. Each agent can have different roles and functions, making decisions and taking actions according to its own responsibilities and tasks. Through collaboration between agents, task division and cooperation can be achieved, allowing each agent to focus on its own task and improve efficiency and effectiveness. Through communication and collaboration, agents can adapt and flexibly adjust based on the information and actions of other agents. When the environment changes or other agents need support, agents can adjust and provide support according to the collaboration strategy to adapt to changes and achieve optimal decisions and actions.

[0146] Communication and collaboration among intelligent agents can enable information sharing, collaborative decision-making, division of labor, and adaptive flexibility, thereby improving the overall performance and effectiveness of the system. In model-driven autonomous control of lithium hydroxide production, this communication and collaboration mechanism can enhance the intelligence level and autonomous decision-making capability of the production process.

[0147] The design of the decision-making mechanism for an intelligent agent specifically includes the following steps:

[0148] S2121 defines the environment and the intelligent agent.

[0149] Taking temperature regulation strategies as an example, the environment is the lithium hydroxide production process, particularly the part related to temperature regulation. The environment needs to be able to provide feedback on the agent's current state and respond to the agent's behavior.

[0150] An intelligent agent is a system for controlling temperature, which needs to be able to observe the state of the environment and make decisions.

[0151] S2122, determine the status, action, and reward.

[0152] The state can include parameters such as current temperature, pressure, and reaction rate.

[0153] Actions are operations that an intelligent agent can perform, such as raising or lowering the temperature.

[0154] Rewards are feedback received by an agent based on the effects of its actions. For example, a positive reward is given if temperature regulation leads to higher yields or better product quality, and a negative reward is given if the temperature regulation results in lower yields or lower product quality.

[0155] S2123, Select a suitable reinforcement learning model.

[0156] Choose an appropriate reinforcement learning model based on the complexity of the problem and the characteristics of the data. For temperature regulation, Q-Learning, Deep Q-Network (DQN), or Proximal Policy Optimization (PPO) can be considered.

[0157] S2124, Model Training.

[0158] Using simulated or historical production data, the agent is trained to select the optimal action under different conditions.

[0159] Through continuous trial and error and learning, the agent learns how to choose the action that leads to the highest reward based on the current state.

[0160] S2125, testing and tuning.

[0161] Specifically, the performance of the intelligent agent is tested in a simulated environment or a real control system. Based on the test results, the model is adjusted and optimized to improve the accuracy and efficiency of decision-making.

[0162] Taking Deep Q-Network (DQN) as an example, the specific steps for training a model using DQN include:

[0163] S21241, Initialize the DQN model.

[0164] First, a deep Q-network is initialized. This network takes the state of the lithium hydroxide production process as input and outputs the expected reward for each possible action.

[0165] For example, a network can take into account current temperature, pressure, and other conditions, and output the expected reward value for actions that raise or lower the temperature.

[0166] S21242, Gathering experience.

[0167] The agent performs actions in a simulated lithium hydroxide production environment, collecting data on state, actions, rewards, and new states (state transitions).

[0168] This data (experience) is stored in a so-called "experience replay pool" for subsequent learning.

[0169] S21243, small-batch learning.

[0170] A small batch of experiences is randomly selected from the experience replay pool.

[0171] The DQN model is updated using this empirical data. Specifically, the update rules of Q-learning are used in conjunction with backpropagation and gradient descent methods from deep learning.

[0172] S21244, Exploration and Utilization.

[0173] In the early stages of training, the agent should be allowed to explore more (e.g., through an epsilon-greedy strategy) in order to gather diverse experiences.

[0174] As training progresses, exploratory behavior is gradually reduced, and more of the learned strategies are utilized.

[0175] S21245, Model Evaluation and Adjustment.

[0176] Regularly evaluate the agent's performance. This can be done by checking its performance in a simulated environment.

[0177] Based on the evaluation results, it may be necessary to adjust the architecture of DQN (such as increasing the number of layers or changing the activation function) or learning parameters (such as learning rate or discount factor).

[0178] Furthermore, the development of deep reinforcement learning algorithms specifically includes:

[0179] S221, Select a suitable reinforcement learning framework.

[0180] Specifically, based on the characteristics of the lithium hydroxide production process, suitable reinforcement learning algorithms are selected, such as Q-learning, deep Q-network (DQN), or proximal policy optimization (PPO).

[0181] Consider the algorithm's stability, convergence speed, and ability to handle complex problems.

[0182] S222, Construct the state and reward functions.

[0183] Specifically, the state space of the system is defined, including all observable production process parameters such as temperature, pressure, and raw material ratio.

[0184] Design reward functions to reflect key performance indicators such as production efficiency, product quality, and energy consumption.

[0185] S223, Network Architecture and Parameter Settings.

[0186] Specifically, design neural network architectures suitable for deep reinforcement learning, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs).

[0187] Adjust network parameters, including learning rate and discount factor, to optimize the learning process.

[0188] S224, Training and Testing.

[0189] Specifically, agents are trained in a simulated environment (each agent is trained individually) using a large amount of production data for iterative learning.

[0190] Conduct testing and optimization to ensure the reliability and robustness of the algorithm in real-world environments.

[0191] In this application, the intensive training specifically includes the following steps:

[0192] S311, Establish the simulation environment.

[0193] Specifically, a virtual lithium hydroxide production environment is constructed to simulate the real production process, including all aspects such as raw material input, reaction conditions, and product output.

[0194] Ensure that the simulated environment accurately reflects the physical and chemical reaction processes in actual production.

[0195] S312, deploy intelligent agents.

[0196] Specifically, pre-trained agents (already trained in S224) are deployed in a simulated environment, such as temperature control agents and pressure monitoring agents.

[0197] Ensure that each agent can function properly and make decisions in the simulation environment.

[0198] S313, Training Process Management.

[0199] Specifically, through extensive iterations and simulation experiments, the agent learns and optimizes in a simulated environment.

[0200] The model is continuously adjusted and optimized based on the agent's performance using methods such as supervised learning, unsupervised learning, or reinforcement learning.

[0201] S314, performance evaluation and optimization.

[0202] Specifically, the performance of the agent is evaluated through simulation results, including its impact on production efficiency, product quality, and energy consumption.

[0203] Adjust the agent's strategy and model parameters based on the evaluation results to achieve better control.

[0204] Furthermore, taking the application of reinforcement learning methods as an example, the following steps are provided in detail:

[0205] S3131, Interaction between the environment and the intelligent agent.

[0206] Specifically, the agent performs actions in a simulated or real environment, and the environment rewards or punishes the agent based on the actions.

[0207] For example, an intelligent agent adjusts a parameter in the lithium hydroxide production process, and the environment provides feedback based on production efficiency and product quality.

[0208] S3132, Gather experience.

[0209] Specifically, each action of the agent and the feedback from the environment (state, action, reward, new state) are stored to form learning experience.

[0210] This data is used in the subsequent learning process to help the agent understand which actions are beneficial.

[0211] S3133, Strategy Update.

[0212] Specifically, reinforcement learning algorithms (such as Q-learning, DQN, PPO, etc.) are used to update the agent's policy based on the collected experience.

[0213] This includes adjusting parameters in the decision-making process so that the agent can make better decisions in similar situations.

[0214] S3134, Exploration and Utilization.

[0215] Specifically, by striking a balance between exploration and exploitation, agents can not only utilize known good strategies but also explore new possibilities.

[0216] For example, using an epsilon-greedy strategy, an action is randomly selected with a certain probability, rather than always choosing the action currently considered optimal.

[0217] S3135, performance monitoring and tuning.

[0218] Specifically, the performance of intelligent agents in the environment is monitored regularly, such as production efficiency and resource utilization.

[0219] Adjust the learning process and model parameters based on the monitoring results, for example, by changing the reward function or fine-tuning the algorithm.

[0220] In this application, each trained agent model is deployed to various key stages of the lithium hydroxide production line, i.e., distributed execution, specifically including the following steps:

[0221] S321, Agent Deployment.

[0222] Specifically, the trained agent model is deployed to various key stages of the lithium hydroxide production line, such as raw material handling, reaction control, and product quality monitoring.

[0223] Ensure that each agent can independently perform its decision-making and control tasks.

[0224] S322, Real-time Data Integration.

[0225] Specifically, the trained agent model is deployed to various key stages of the lithium hydroxide production line, such as raw material handling, reaction control, and product quality monitoring.

[0226] Ensure that each agent can independently perform its decision-making and control tasks.

[0227] S323, Agent Decision Execution.

[0228] Specifically, the agent makes decisions based on real-time data and the strategies it has learned, such as adjusting reaction conditions and optimizing raw material ratios.

[0229] The agent's decisions should be fed back into the production process in real time to enable optimization and adjustment.

[0230] S324, Monitoring and Feedback Adjustment;

[0231] Specifically, the production process and the performance of intelligent agents are continuously monitored to assess the effectiveness of decision-making and production efficiency.

[0232] Based on monitoring results and actual production needs, the strategies of the intelligent agents are adjusted and optimized in a timely manner.

[0233] This solution combines multi-agent deep reinforcement learning (DRL) with a strategy of centralized training and distributed execution to achieve efficient and precise control of the lithium hydroxide production process. The system aims to optimize production efficiency, reduce costs, and ensure product quality.

[0234] The following detailed description, in conjunction with specific embodiments, demonstrates how multi-agent DRLs optimize reaction conditions for lithium hydroxide, such as temperature and pressure, to improve yield and quality.

[0235] Example 1: Optimize the production process

[0236] Step 1: Process Analysis and Goal Setting

[0237] Step 1.1: Analyze the existing production process

[0238] Review the current lithium hydroxide production process to identify key production steps and potential bottlenecks.

[0239] Step 1.2: Define the optimization goal

[0240] Set specific optimization goals, such as improving production efficiency, reducing energy consumption, and improving product quality.

[0241] Step 2: Data Collection and Preprocessing

[0242] Step 2.1: Data Collection

[0243] Collect relevant data from the production line, including raw material characteristics, machine parameters, and environmental conditions.

[0244] Step 2.2: Data Preprocessing

[0245] The collected data is cleaned, formatted, and standardized.

[0246] Step 3: Build a deep reinforcement learning model

[0247] Step 3.1: Design the intelligent agent structure

[0248] Design a multi-agent structure suitable for the lithium hydroxide production process.

[0249] Step 3.2: Develop reinforcement learning algorithms

[0250] Based on the collected data and the characteristics of the production process, we developed applicable deep reinforcement learning algorithms.

[0251] Step 4: Model Training and Validation

[0252] Step 4.1: Training in a simulated environment

[0253] Train the agents in a simulated environment to ensure they can understand and optimize production processes.

[0254] Step 4.2: Verify the model's performance

[0255] Test the agent's performance in a simulated environment to ensure it achieves the preset optimization goals.

[0256] Step 5: Application and Adjustment in Real-World Environments

[0257] Step 5.1: Deploy the intelligent agent

[0258] Deploy the trained agent into the actual production environment.

[0259] Step 5.2: Continuous Monitoring and Optimization

[0260] Continuously monitor the performance of the intelligent agent and the efficiency of the production process, and make adjustments and optimizations based on the actual situation.

[0261] Reference Figure 7 The first aspect of this application discloses a model-driven autonomous control system for lithium hydroxide production, comprising:

[0262] The build module is configured to build the system framework;

[0263] The configuration module is set to adaptively set the multi-agent DRL algorithm according to preset requirements in the system framework. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents.

[0264] The training module is configured to perform centralized training on several agents in a simulated environment; the simulated environment is a virtual lithium hydroxide production environment.

[0265] The deployment module is configured to deploy each trained agent model to various key stages of the lithium hydroxide production line.

[0266] The data acquisition module is configured to collect several initial data points in the production process in real time and feed these initial data points back to the corresponding intelligent agents.

[0267] The execution module is configured as a corresponding intelligent agent based on the multi-agent DRL algorithm to adjust the input information in lithium hydroxide production in real time; the input information includes reaction conditions and optimized raw material ratio.

[0268] An electronic device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, flash memory, etc.

[0269] The processor may be a central processing unit (CPU) or other processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the electronic device to perform all or part of the steps of the model-driven autonomous control method for lithium hydroxide production described in the foregoing embodiments of this disclosure.

[0270] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0271] like Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the electronic device in the embodiment of the present disclosure. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0272] like Figure 8 As shown, an electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from a storage device into random access memory (RAM). The RAM also stores various programs and data required for the operation of the electronic device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0273] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow electronic devices to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0274] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the model-driven autonomous control method for lithium hydroxide production according to embodiments of this disclosure are performed.

[0275] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0276] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the model-driven autonomous control method for lithium hydroxide production described in the foregoing embodiments of the present disclosure are performed.

[0277] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0278] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0279] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

Claims

1. A model-driven autonomous control method for lithium hydroxide production, characterized in that, include: Build the system framework; In the system framework, a multi-agent DRL algorithm is adaptively set according to preset requirements. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents. Several of the aforementioned agents are trained in a simulated environment, which is a virtual lithium hydroxide production environment. Each trained agent is deployed to the lithium hydroxide production line; Several initial data points are collected in real time during the production process, and the initial data are sent to the corresponding intelligent agent; The corresponding intelligent agent is based on the multi-agent DRL algorithm and adjusts the input information in lithium hydroxide production in real time; In the system framework, the multi-agent DRL algorithm is adaptively configured according to preset requirements, including: Several intelligent agents can be flexibly set up according to preset requirements; each of the intelligent agents corresponds one-to-one with the functional agents required in the lithium hydroxide production process. A decision-making mechanism is set for each of the aforementioned agents; Set interaction strategies between the different intelligent agents; the interaction strategies include communication strategies and cooperation strategies. The aforementioned intelligent agents include a reaction control intelligent agent, a temperature monitoring intelligent agent, and a pressure control intelligent agent; The decision-making mechanism includes a reaction control mechanism, a temperature monitoring mechanism, and a pressure control intelligent mechanism.

2. The model-driven autonomous control method for lithium hydroxide production according to claim 1, characterized in that, The system framework construction includes establishing a lithium hydroxide production process model: The establishment of the lithium hydroxide production process model includes: Set preset control targets; the preset control targets include one or more of the following: production efficiency indicators, energy consumption indicators, and product quality indicators; A lithium hydroxide production process model is established based on the preset control objectives.

3. The model-driven autonomous control method for lithium hydroxide production according to claim 2, characterized in that, The establishment of the lithium hydroxide production process model based on the preset control target includes: Collect first data; the first data includes one or more of the following: raw material characteristic data, reaction condition data, and product quality index data; The first data is preprocessed to obtain the second data; The second data is analyzed according to a preset strategy to obtain the third data; The preset mathematical model is trained based on the third data, and the trained model is the established lithium hydroxide production process model.

4. The model-driven autonomous control method for lithium hydroxide production according to claim 3, characterized in that, The step of analyzing the second data according to a preset strategy to obtain the third data includes: A descriptive statistical analysis strategy was used to conduct a preliminary analysis of the second data to obtain data distribution information; The Pearson correlation coefficient strategy is used to evaluate the data distribution information and obtain the strength of the linear relationship between different data. The strength of the linear relationship was analyzed using statistical testing methods to identify key factors that significantly affect production results; these key factors are those whose impact on the production process is no less than a preset threshold. The key factors are filtered using machine learning algorithms to obtain the third data; The third set of data is a subset of features that meets the model prediction criteria.

5. The model-driven autonomous control method for lithium hydroxide production according to claim 3, characterized in that, The step of training a preset mathematical model based on the third data, resulting in a trained model that is the established lithium hydroxide production process model, includes: Analyze the data characteristics of the first data; Based on the data characteristics and prediction objectives, an initial regression model is selected; The initial regression model is optimized and trained using historical data to obtain the first model; The first model is validated and trained using preset indicators. The model that meets the validation requirements is the established lithium hydroxide production process model.

6. The model-driven autonomous control method for lithium hydroxide production according to claim 5, characterized in that, The preset indicators include one or more of the following: cross-validation indicators, mean square error indicators, and coefficient of determination indicators.

7. The model-driven autonomous control method for lithium hydroxide production according to claim 1, characterized in that, The communication strategy includes communicating via a local area network or via a cloud platform.

8. A model-driven autonomous control system for lithium hydroxide production, used to execute the model-driven autonomous control method for lithium hydroxide production as described in any one of claims 1 to 7, characterized in that, include: The build module is configured to build the system framework; The configuration module is configured to adaptively set a multi-agent DRL algorithm according to preset requirements in the system framework. The multi-agent DRL algorithm includes several agents, the decision-making mechanism of each agent, and the interaction strategy between different agents. The training module is configured to perform centralized training on several of the aforementioned agents in a simulated environment; the simulated environment is a virtual lithium hydroxide production environment. The deployment module is configured to deploy each trained agent to the lithium hydroxide production line. The data acquisition module is configured to collect several initial data points during the production process in real time and send the initial data points to the intelligent agent. The execution module is configured to adjust the input information in lithium hydroxide production in real time based on the multi-agent DRL algorithm.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning scheduling method and system, and electronic device

    CN109947567A

  • Intelligent data detection method for battery-grade lithium hydroxide

    CN115114986A