Intelligent Optimization Control System and Method for the Recycling and Preparation Process of Battery Materials

Through the combination of multimodal sensing technology and deep reinforcement learning model, the battery material recycling and preparation process is automatically regulated, which solves the problems of cumbersome parameter regulation and high energy consumption in traditional methods, and improves production efficiency and resource utilization.

CN119395996BActive Publication Date: 2025-07-18常州厚丰新能源有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411513271.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-07-18
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

During the recycling and preparation of traditional battery materials, process parameters are cumbersome, product quality fluctuates greatly, energy consumption is high, manual control flexibility is poor, and it is difficult to deal with sudden changes in working conditions, resulting in low production efficiency.

Method used

Multimodal sensing technology is used to collect training sample data, train quality prediction models, and build the state space and action space of the deep reinforcement learning model. Through feature fusion, the fusion process characterization vector is generated, and the parameters of the preparation process are recycled in real time.

Benefits of technology

It realizes automated optimization control, reduces production energy consumption, improves resource utilization efficiency, and solves the problems of artificial experience dependence and inefficient parameter regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119395996B_ABST
    Figure CN119395996B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent optimization control system and method for the battery material recycling preparation process, which relates to the field of intelligent control technology. By collecting training sample data in the battery material recycling preparation process in the test environment, training a quality prediction model, then constructing the state space, action space and reward function of the deep reinforcement learning model, fusing the prediction results of the quality prediction model and the monitoring data of the sensor in the test environment through a feature fusion strategy to generate a fused process characterization vector, inputting the fused process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model, and based on real-time sensor data, the quality prediction model and the deep reinforcement learning model, obtaining the action output by the deep reinforcement learning model, and using this action as an optimization strategy to adjust the parameters of the recycling preparation process to be monitored; effectively reducing production energy consumption and improving resource utilization efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control, and specifically to an intelligent optimization control system and method for the battery material recycling and preparation process. Background Art

[0002] There are many deficiencies in the traditional battery material recycling and preparation process, such as cumbersome process parameter regulation, large fluctuations in product quality, high energy consumption, etc. There is an urgent need for an intelligent and automated solution to optimize and control the entire recycling and preparation process, improve resource utilization rate, and reduce production costs.

[0003] Currently, the battery material recycling and preparation process mainly relies on manual experience for parameter setting and process control. This traditional method has obvious defects. Due to the complexity and uncertainty of influencing factors, it is difficult for humans to accurately grasp the optimal parameter combination of each link, which often leads to unstable product quality and inefficient use of energy. In addition, the flexibility of manual control of reactions is poor, and it is difficult to respond promptly to sudden changes in working conditions, thus affecting production efficiency.

[0004] Chinese Patent with the authorization announcement number CN117810570B discloses a recycling discharge control method and related device for waste lithium batteries. Classification is carried out based on the characteristic tags generated by clustering analysis of the state parameters of waste lithium batteries to obtain several waste lithium batteries of each category; real-time images of several waste lithium batteries are acquired to determine the coordinate positions of their positive and negative electrodes; several waste lithium batteries of the corresponding category are placed at the corresponding positions of the battery discharge device for discharging based on the coordinate positions of the positive and negative electrodes; during the discharge process, corresponding monitoring data is acquired by a combination of several monitoring sensors, and the working parameters of the discharge device are regulated based on the monitoring data; at the same time, the remaining power of each waste lithium battery is continuously calculated based on the monitoring data, and when its remaining power reaches the preset interval, the discharge process is ended; however, this method fails to control the recycling parameters to improve the quality of the recycled materials;

[0005] Therefore, the present invention proposes an intelligent optimization control system and method for the battery material recycling and preparation process. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention proposes an intelligent optimization control system and method for the battery material recycling and preparation process, which effectively reduces production energy consumption and improves resource utilization efficiency.

[0007] To achieve the above object, an intelligent optimization control method for the battery material recycling and preparation process is proposed, including the following steps:

[0008] Step 1: Use multi-modal sensing technology to collect training sample data in the battery material recycling and preparation process in the test environment;

[0009] Step 2: Based on the training sample data, train a quality prediction model with the product quality index as the output;

[0010] Step 3: Based on the prediction results of the quality prediction model, construct the state space of the deep reinforcement learning model, and construct the action space and reward function of the deep reinforcement learning model;

[0011] Step 4: Through the feature fusion strategy, fuse the prediction results of the quality prediction model and the monitoring data of the sensors in the test environment to generate a fusion process characterization vector;

[0012] Step 5: Input the fusion process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model;

[0013] Step 6: For the recycling and preparation process to be monitored, collect real-time sensor data in real time through multi-modal sensing technology. Based on the real-time sensor data, quality prediction model and deep reinforcement learning model, obtain the action output by the deep reinforcement learning model; use this action as an optimization strategy to adjust the parameters of the recycling and preparation process to be monitored.

[0014] The collection of the training sample data includes the following steps:

[0015] Step 11: Prepare N batches of waste battery anode and cathode materials with different sources and compositions; N is the number of selected batches;

[0016] Step 12: In the test environment, cycle through the complete recycling and preparation process flow for each batch of waste battery anode and cathode materials, and record the monitoring data of each sensor in real time, and synchronously record the controllable key parameters during the operation process;

[0017] Step 13: Measure and record the index value of the quality index of the battery material obtained after extraction and recycling as the quality index label;

[0018] Step 14: Combine the monitoring data collected in real time by each sensor, the real-time controllable key parameters, and the quality index labels of all batches of waste battery anode and cathode materials to form the training sample data.

[0019] The training of the quality prediction model with the product quality index as the output based on the training sample data includes the following steps:

[0020] Step 21: Convert the training sample data into sample input features;

[0021] The conversion of the training sample data into sample input features includes the following steps:

[0022] Step 211: For the positive and negative electrode materials of waste batteries in each batch, preprocess the monitoring data of the sensors at each corresponding moment, and splice the preprocessed monitoring data and controllable key parameters to obtain a feature splicing vector at each moment;

[0023] The construction method of the feature splicing vector is as follows:

[0024] Perform vectorization operations on the monitoring data of each sensor at each moment to obtain corresponding parameter feature vectors;

[0025] For each controllable key parameter, use its respective parameter values to form a controllable parameter feature vector;

[0026] Splice the parameter feature vectors of all sensors and the controllable parameter feature vectors of the positive and negative electrode materials of waste batteries at each moment to obtain a feature splicing vector.

[0027] Step 212: Preset a time window t; for each moment during the operation of the positive and negative electrode materials of waste batteries in each batch, collect the monitoring parameter sequence of the monitoring data of each corresponding sensor and the controllable parameter sequence of each controllable key parameter at this moment; the monitoring parameter sequence is a sequence formed by arranging the parameter values collected by the corresponding sensor in chronological order within t moments before this moment; the controllable parameter sequence is a sequence of parameter values formed by arranging the parameter values of the corresponding controllable key parameters in chronological order within t moments before this moment;

[0028] The monitoring parameter sequence and the controllable parameter sequence at each moment form a sample input feature.

[0029] Step 22: Use the sample input feature at each moment as the input, and use the predicted value of the quality index at the subsequent t0 moment as the output to train a quality prediction model; t0 is the preset duration of the time period.

[0030] The method of training the quality prediction model is as follows:

[0031] Use the sample input feature at each moment as the input of the quality prediction model. The quality prediction model takes the predicted value of the quality index at the subsequent t0 moment of this moment as the output, takes the quality index label corresponding to the subsequent t0 moment of this moment as the prediction target, takes the difference between the predicted value of the quality index and the quality index label as the prediction error, and takes minimizing the sum of the prediction errors as the training target; train the quality prediction model until the sum of the prediction errors converges and then stop training; the quality prediction model is a time series prediction model.

[0032] The method for constructing the state space of the deep reinforcement learning model, the action space of the deep reinforcement learning model, and the reward function based on the prediction result of the quality prediction model is as follows:

[0033] The deep reinforcement learning model is set as an Actor-Critic model;

[0034] Mark the state space as S. The state space S reflects the real-time state of the entire recycling and preparation process. The construction method of the state space is as follows:

[0035] Segment the recycling and preparation process according to a fixed time window; the size of the fixed time window is t0;

[0036] For each time segment, use the set of characterization vectors composed of the characterization vectors of the fusion process at each moment in chronological order as the state space;

[0037] The construction method of the characterization vector of the fusion process is as follows:

[0038] Collect the feature splicing vectors corresponding to each moment in the recycling and preparation process of the positive and negative electrode materials of waste batteries;

[0039] Input the feature splicing vector at each moment into the quality prediction model to obtain the predicted value of the quality index at the subsequent t0 moments output by the quality prediction model;

[0040] Use a one-dimensional convolutional network to extract the local feature sequence from the monitoring parameter sequence of the monitoring data of each sensor corresponding to each moment;

[0041] Input the local feature sequence into the LSTM model to obtain the global context encoding vector Ct;

[0042] Concatenate the predicted value of the quality index and the global context encoding vector Ct to obtain the characterization vector of the fusion process.

[0043] Mark the action space as A. The action space A includes the adjustment amount of the controllable key parameters output by the Actor model. Mark the parameter adjustment action taken in the next time segment as a;

[0044] Mark the reward function as Q. The reward value function Q = r(s, a), which measures the effect of performing the action a in the current state s;

[0045] Design the agent of the deep reinforcement learning model as an RNN-based Actor-Critic structure;

[0046] The input of the agent is the state s;

[0047] The state value V(s) is output by the Critic network;

[0048] The Actor network outputs the action probability π(a|s).

[0049] The method for obtaining the action output by the deep reinforcement learning model based on real-time sensor data, a quality prediction model, and a deep reinforcement learning model is as follows:

[0050] Construct a feature concatenation vector through real-time sensor data and the parameter values of various controllable key parameters in the current recycling and preparation process;

[0051] Input the feature concatenation vector into the quality prediction model to obtain the predicted value of the quality index at the subsequent time t0 output by the quality prediction model;

[0052] Based on the predicted value of the quality index and the monitoring data of various sensors, construct a corresponding fusion process characterization vector;

[0053] Collect all the fusion process characterization vectors in the time segment where the current moment is located to form a corresponding characterization vector set;

[0054] Input the characterization vector set into the deep reinforcement learning model to obtain the action output by the Actor network of the deep reinforcement learning model.

[0055] An intelligent optimization control system for the battery material recycling and preparation process is proposed, including a sample data collection module, a prediction model training module, a reinforcement model training module, and an action decision module; among them, each module is connected electrically;

[0056] The sample data collection module uses multi-modal sensing technology to collect training sample data in the battery material recycling and preparation process in the test environment and sends the training sample data to the prediction model training module;

[0057] The prediction model training module trains a quality prediction model with the product quality index as the output based on the training sample data and sends the quality prediction model to the reinforcement model training module and the action decision module;

[0058] The reinforcement model training module constructs the state space of the deep reinforcement learning model based on the prediction results of the quality prediction model, constructs the action space and reward function of the deep reinforcement learning model, fuses the prediction results of the quality prediction model and the monitoring data of the sensors in the test environment through a feature fusion strategy to generate a fusion process characterization vector, inputs the fusion process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model, and sends the deep reinforcement learning model to the action decision module;

[0059] The action decision-making module, for the recycling preparation process to be monitored, collects real-time sensor data in real time through multi-modal sensing technology. Based on the real-time sensor data, the quality prediction model, and the deep reinforcement learning model, it obtains the action output by the deep reinforcement learning model; and uses this action as an optimization strategy to adjust the parameters of the recycling preparation process to be monitored.

[0060] There is provided an electronic device, including: a processor and a memory, wherein, a computer program that can be called by the processor is stored in the memory;

[0061] The processor executes the intelligent optimization control method for the battery material recycling preparation process described above by calling the computer program stored in the memory.

[0062] There is provided a computer-readable storage medium, on which a rewritable computer program is stored;

[0063] When the computer program runs on a computer device, it causes the computer device to execute the intelligent optimization control method for the battery material recycling preparation process described above.

[0064] Compared with the prior art, the beneficial effects of the present invention are:

[0065] The present invention collects training sample data in the battery material recycling preparation process in the test environment by using multi-modal sensing technology. Based on the training sample data, it trains a quality prediction model with the product quality index as the output. Based on the prediction result of the quality prediction model, it constructs the state space of the deep reinforcement learning model, and constructs the action space and the reward function of the deep reinforcement learning model. Through the feature fusion strategy, it fuses the prediction result of the quality prediction model and the monitoring data of the sensor in the test environment to generate a fused process characterization vector, and inputs the fused process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model. For the recycling preparation process to be monitored, it collects real-time sensor data in real time through multi-modal sensing technology. Based on the real-time sensor data, the quality prediction model, and the deep reinforcement learning model, it obtains the action output by the deep reinforcement learning model; and uses this action as an optimization strategy to adjust the parameters of the recycling preparation process to be monitored; through big data modeling and machine learning technology, it excavates the internal laws affecting the preparation process, and then automatically generates the optimal parameter configuration and seamlessly integrates it into the automatic control system, thus completely solving problems such as dependence on manual experience and low efficiency of parameter regulation, effectively reducing production energy consumption, and improving resource utilization efficiency. Description of the Drawings

[0066] Figure 1 It is a flowchart of the intelligent optimization control method for the battery material recycling preparation process in Embodiment 1 of the present invention;

[0067] Figure 2It is a module connection diagram of the intelligent optimization control system for the battery material recycling and preparation process in Embodiment 1 of the present invention;

[0068] Figure 3 It is a schematic structural diagram of an electronic device in Embodiment 3 of the present invention;

[0069] Figure 4 It is a schematic structural diagram of a computer-readable storage medium in Embodiment 4 of the present invention. Detailed implementation manners

[0070] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0071] Embodiment 1

[0072] As Figure 1 shown, the intelligent optimization control method for the battery material recycling and preparation process includes the following steps:

[0073] Step 1: Use multi-modal sensing technology to collect training sample data in the battery material recycling and preparation process in the test environment;

[0074] Step 2: Based on the training sample data, train a quality prediction model with the product quality index as the output;

[0075] Step 3: Based on the prediction result of the quality prediction model, construct the state space of the deep reinforcement learning model, and construct the action space and reward function of the deep reinforcement learning model;

[0076] Step 4: Through the feature fusion strategy, fuse the prediction result of the quality prediction model and the monitoring data of the sensor in the test environment to generate a fusion process characterization vector;

[0077] Step 5: Input the fusion process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model;

[0078] Step 6: For the recycling and preparation process to be monitored, collect real-time sensor data in real time through multi-modal sensing technology. Based on the real-time sensor data, quality prediction model and deep reinforcement learning model, obtain the action output by the deep reinforcement learning model; use this action as the optimization strategy to adjust the parameters of the recycling and preparation process to be monitored.

[0079] Specifically, the multi-modal sensing technology is to use various types of sensors to collect corresponding different forms of sensor data in real time during the battery material recycling and preparation process.

[0080] Specifically, the multimodal sensing technology includes, but is not limited to:

[0081] Spectral sensors, such as UV-Vis, NIR, Raman, etc., electrochemical sensor arrays, such as pH, ORP, ISE, etc., physical sensors, such as temperature sensor arrays, pressure sensors, flow meters, etc., imaging sensors, such as CCD, thermal imaging, etc., acoustic sensors, such as ultrasonic waves, etc.

[0082] Among them, in-situ Raman spectrometers monitor the changes in chemical components in solutions in real time;

[0083] X-ray fluorescence (XRF) analyzers detect the content of metal elements online;

[0084] Near-infrared spectroscopy (NIR) is used to monitor the content of organic substances and the purity of solvents;

[0085] Among them, ion-selective electrodes (ISE) are used to measure the concentration of specific ions (such as Li+, Co2+, Ni2+, etc.);

[0086] pH electrodes and oxidation-reduction potential (ORP) electrodes monitor the acidity, alkalinity, and oxidation-reduction state of solutions;

[0087] Among them, temperature sensor arrays monitor the temperature distribution at different positions in the reactor;

[0088] Pressure sensors are used to control the reaction pressure and detect abnormal situations;

[0089] Flow meters accurately control the addition amounts of various solutions and reagents;

[0090] Among them, high-speed cameras capture the changes in particle morphology and distribution during the reaction process.

[0091] Thermal imaging cameras monitor the hot spots of equipment to prevent potential failures;

[0092] Among them, ultrasonic sensors are used to monitor the solid content and particle size distribution in solutions.

[0093] Acoustic emission sensors detect abnormal vibrations of equipment for predictive maintenance.

[0094] Furthermore, the test environment refers to a simulated environment of the battery material recycling and preparation process constructed under laboratory or factory conditions to conduct tests during the recycling and preparation process, so as to collect training samples.

[0095] Furthermore, the collection of the training sample data includes the following steps:

[0096] Step 11: Prepare N batches of cathode and anode materials of waste batteries with different sources and compositions; N is the number of selected batches;

[0097] Step 12: In the test environment, run the complete recycling and preparation process flow for the positive and negative electrode materials of each batch of waste batteries in a cycle, record the monitoring data of each sensor in real time, and synchronously record various controllable key parameters during the operation process;

[0098] Specifically, the controllable key parameters are physical parameters that can be automatically adjusted during the recycling and preparation process, including but not limited to the temperature, pressure, extractant concentration, etc. during the extraction process.

[0099] Step 13: Measure and record the index values of the quality indicators of the battery materials obtained after extraction and recycling as quality indicator labels;

[0100] Specifically, the quality indicator can be the purity of the extracted and recycled substance, and the purity of this extracted battery raw material is related to the controllable key parameters during the extraction process.

[0101] Step 14: Combine the monitoring data collected in real time by each sensor, the real-time controllable key parameters, and the quality indicator labels of the positive and negative electrode materials of all batches of waste batteries to form training sample data.

[0102] Furthermore, training a quality prediction model with the product quality indicator as the output based on the training sample data includes the following steps:

[0103] Step 21: Convert the training sample data into sample input features;

[0104] Specifically, converting the training sample data into sample input features includes the following steps:

[0105] Step 211: For the positive and negative electrode materials of each batch of waste batteries, preprocess the monitoring data of the sensors at each corresponding moment, and perform feature splicing on the preprocessed monitoring data and controllable key parameters to obtain a feature splicing vector at each moment;

[0106] Specifically, the preprocessing includes but not limited to operations such as noise removal and feature dimensionality reduction;

[0107] Preferably, the construction method of the feature splicing vector is as follows:

[0108] Perform vectorization operations on the monitoring data of each sensor at each moment to obtain corresponding parameter feature vectors;

[0109] Specifically, for example: use the principal component analysis method for spectral data to obtain the spectral principal component vector, use sensor readings for electrochemical data to obtain the electrochemical vector, use a convolutional autoencoder to encode image data to obtain the image feature vector, and calculate the Mel frequency gradient for sound data to obtain the sound feature vector.

[0110] For each controllable key parameter, use their respective parameter values to form a controllable parameter feature vector;

[0111] Concatenate the parameter feature vectors of all sensors at each moment of the positive and negative electrode materials of the waste battery and the controllable parameter feature vector to obtain a feature concatenation vector; it can be understood that by vectorizing the feature of the monitoring data in each data form, the unity of multi-modal data is ensured.

[0112] Step 212: Preset a time window t; for each moment during the operation of each batch of positive and negative electrode materials of the waste battery, collect the monitoring parameter sequence of the monitoring data of each corresponding sensor and the controllable parameter sequence of each controllable key parameter at this moment; the monitoring parameter sequence is a sequence formed by arranging the parameter values collected by the corresponding sensor in chronological order within t moments before this moment; the controllable parameter sequence is a parameter value sequence formed by arranging the parameter values of the corresponding controllable key parameter in chronological order within t moments before this moment;

[0113] The monitoring parameter sequence and the controllable parameter sequence at each moment form a sample input feature.

[0114] Step 22: Use the sample input feature at each moment as the input, and use the predicted value of the quality index at the subsequent t0 moment as the output to train a quality prediction model; t0 is the preset duration of the time period.

[0115] Specifically, the method of training the quality prediction model is as follows:

[0116] Use the sample input feature at each moment as the input of the quality prediction model. The quality prediction model uses the predicted value of the quality index at the subsequent t0 moment of this moment as the output, uses the quality index label corresponding to the subsequent t0 moment of this moment as the prediction target, uses the difference between the predicted value of the quality index and the quality index label as the prediction error, and uses minimizing the sum of the prediction errors as the training target; train the quality prediction model until the sum of the prediction errors reaches convergence and then stop training; the quality prediction model is a time series prediction model; the time series prediction model is any one of the RNN model or the LSTM model.

[0117] Furthermore, the method of constructing the state space of the deep reinforcement learning model, constructing the action space of the deep reinforcement learning model, and constructing the reward function based on the prediction result of the quality prediction model is as follows:

[0118] The deep reinforcement learning model is set as an Actor-Critic model;

[0119] Mark the state space as S. The state space S reflects the real-time state of the entire recycling and preparation process. The state space is constructed as follows:

[0120] Segment the recycling and preparation process according to a fixed time window. The size of the fixed time window is t0. It can be understood that by setting the fixed time window to the same size as the preset time period duration, the quality indicators in the subsequent time segments can be predicted, and thus the adjustment amount of the controllable key parameters can be guided according to the subsequent quality indicators.

[0121] For each time segment, use the set of characterization vectors formed by arranging the characterization vectors of the fusion process at each moment in chronological order as the state space.

[0122] Specifically, the fusion process characterization vector is constructed as follows:

[0123] Collect the characteristic splicing vectors corresponding to each moment in the recycling and preparation process of the positive and negative electrode materials of waste batteries.

[0124] Input the characteristic splicing vector at each moment into the quality prediction model to obtain the predicted value of the quality indicator at the subsequent t0 moment output by the quality prediction model.

[0125] Use a one-dimensional convolutional network to extract the local feature sequence from the monitoring parameter sequence of the monitoring data of each sensor corresponding to each moment.

[0126] Input the local feature sequence into the LSTM model to obtain the global context encoding vector Ct.

[0127] Concatenate the predicted value of the quality indicator and the global context encoding vector Ct to obtain the fusion process characterization vector.

[0128] It should be noted that the fusion process characterization vector captures the internal semantic associations between different features. Therefore, it can enable the subsequent model training process to learn the correlations and coupling patterns between features, and the fused feature vector can be directly input into the subsequent reinforcement learning model. While simple concatenation usually requires a complex model architecture to process and is more difficult to converge in training.

[0129] Mark the action space as A. The action space A includes the adjustment amounts of the controllable key parameters output by the Actor model, such as the temperature adjustment amount, the pressure adjustment amount, and the extraction concentration adjustment amount, etc. Mark the parameter adjustment action taken in the next time segment as a.

[0130] Mark the reward function as Q. The reward value function Q = r(s, a), which measures the effect of executing action a in the current state s.

[0131] Specifically, the reward value function Q can be set as: Q = b1 × |P_t - P| - b2 × ∑ i hi; where both b1 and b2 are preset proportionality coefficients;

[0132] where P_t represents the actual value of the quality index at the current moment, P represents the predicted value of the quality index, which is used to measure the improvement of the quality index value after performing the action a. The greater the improvement, the greater the reward value Q;

[0133] i is the number of controllable key parameters, and hi is the specific adjustment amount of the i-th controllable key parameter. The smaller the adjustment amount, the smaller the perturbation to the recycling and preparation process, the better the execution effect, and the greater the reward Q.

[0134] The agent for designing the deep reinforcement learning model is an Actor-Critic structure based on RNN;

[0135] The input of the agent is the state s;

[0136] The state value V(s) is output by the Critic network;

[0137] The action probability π(a|s) is output by the Actor network.

[0138] Further, the method of inputting the fusion process characterization vector into the deep reinforcement learning model and training the deep reinforcement learning model is as follows:

[0139] Take the set of characterization vectors of each time segment of the recycling and preparation process to be monitored as the current state s; input the current state s into the Actor model of the deep reinforcement learning model, and the Actor model outputs an action in the action space A, so that the state-action value output by the Critic model is maximized, and an action decision on the adjustment amount of each controllable key parameter in the next time segment within the recycling and preparation process to be monitored is output;

[0140] Execute the action decision to obtain the next state s' and the actual reward value;

[0141] The Critic network outputs the state-action value after executing this action decision in the current state, calculates the TD error δ of the Critic network, updates the Critic network parameters according to δ, and minimizes the Critic loss function; the state-action value output by the Critic model is calculated by the state-action value function generated by the time difference optimization method.

[0142] Specifically, the state-action value function generated by the time difference optimization method can be:

[0143] Q(s, a) = r(s, a) + γ × V(s'), where γ is the discount factor and V(s') is the estimated value in state s'; generally, the estimated value can be replaced by Q(s', a'), that is, V(s') represents the estimated value of the state-action value function corresponding to the next state s' and action a', so as to ensure that Q(s, a) can consider more in the long term and not only consider the current state;

[0144] Then the TD error δ is δ = (r(s, a) + γ × V(s') - Q(s, a)) 2 ;

[0145] The Actor model takes maximizing the state-action value output by the Critic model as the output goal and updates the parameters of the Actor network model using the policy gradient update method.

[0146] Furthermore, the method for obtaining the action output by the deep reinforcement learning model based on real-time sensor data, quality prediction model, and deep reinforcement learning model is as follows:

[0147] Construct a feature concatenation vector through real-time sensor data and the parameter values of various controllable key parameters in the current recycling and preparation process;

[0148] Input the feature concatenation vector into the quality prediction model to obtain the predicted value of the quality index at the subsequent time t0 output by the quality prediction model;

[0149] Based on the predicted value of the quality index and the monitoring data of various sensors, construct a corresponding fusion process characterization vector;

[0150] Collect all the fusion process characterization vectors in the time segment where the current moment is located to form a corresponding characterization vector set;

[0151] Input the characterization vector set into the deep reinforcement learning model to obtain the action output by the Actor network of the deep reinforcement learning model.

[0152] Embodiment 2

[0153] As Figure 2 shown, the intelligent optimization control system for the battery material recycling and preparation process includes a sample data collection module, a prediction model training module, a reinforcement model training module, and an action decision module; among them, each module is connected electrically;

[0154] The sample data collection module uses multi-modal sensing technology to collect training sample data in the battery material recycling and preparation process in the test environment and sends the training sample data to the prediction model training module;

[0155] The prediction model training module trains a quality prediction model with product quality indicators as outputs based on training sample data, and sends the quality prediction model to the reinforcement model training module and the action decision module;

[0156] The reinforcement model training module constructs the state space of the deep reinforcement learning model based on the prediction results of the quality prediction model, constructs the action space and reward function of the deep reinforcement learning model, fuses the prediction results of the quality prediction model and the monitoring data of the sensor in the test environment through the feature fusion strategy to generate a fusion process characterization vector, inputs the fusion process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model, and sends the deep reinforcement learning model to the action decision module;

[0157] For the recycling and preparation process to be monitored, the action decision module collects real-time sensor data in real time through multi-modal sensing technology, and obtains the action output by the deep reinforcement learning model based on the real-time sensor data, the quality prediction model and the deep reinforcement learning model; uses this action as an optimization strategy to adjust the parameters of the recycling and preparation process to be monitored.

[0158] Embodiment 3

[0159] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, according to another aspect of the present application, an electronic device 100 is also provided. The electronic device 100 may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute the intelligent optimization control method for the battery material recycling and preparation process as described above.

[0160] The method or device according to the embodiment of the present application can also be implemented by means of Figure 3 the architecture of the electronic device shown. As Figure 3 shown, the electronic device 100 may include a bus 101, one or more CPUs 102, a ROM 103, a RAM 104, a communication port 105 connected to the network, an input / output component 106, a hard disk 107, etc. The storage device in the electronic device 100, such as the ROM 103 or the hard disk 107, can store the intelligent optimization control method for the battery material recycling and preparation process provided by the present application.

[0161] Furthermore, the electronic device 100 may further include a user interface 108. Of course, Figure 3 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the electronic device may be omitted according to actual needs. Figure 3

[0162] ​Embodiment 4

[0163] Figure 4 is a schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present application. As Figure 4 shown, it is a computer-readable storage medium 200 according to an embodiment of the present application. Computer-readable instructions are stored on the computer-readable storage medium 200. When the computer-readable instructions are run by a processor, the intelligent optimization control method for the battery material recovery and preparation process according to the embodiment of the present application described with reference to the above drawings can be executed. The computer-readable storage medium 200 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0164] In addition, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application. When this computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.

[0165] The method, apparatus, and device of the present application can be implemented in many ways. For example, the method, apparatus, and device of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present application are not limited to the above specific described order unless otherwise specifically stated. In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0166] In addition, in the above technical solutions provided by the embodiments of the present application, the parts that are the same as the corresponding technical solutions in the prior art in terms of implementation principle are not described in detail to avoid excessive elaboration.

[0167] As described above in the specific embodiments, the purpose, technical solution, and beneficial effects of the present invention are further described in detail. It should be understood that the above is only the specific embodiment of the present invention and is not used to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0168] The above preset parameters or preset thresholds are all set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0169] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. An intelligent optimization control method for the battery material recycling preparation process, characterized in that It includes the following steps: Step 1: Use multimodal sensing technology to collect training sample data during the recycling preparation process of battery materials in the test environment; Step 2: Based on the training sample data, train a quality prediction model with product quality indicators as the output; Step 3: Based on the prediction results of the quality prediction model, construct the state space of the deep reinforcement learning model, and construct the action space and reward function of the deep reinforcement learning model; Step 4: Through the feature fusion strategy, fuse the prediction results of the quality prediction model and the monitoring data of the sensors in the test environment to generate a fusion process characterization vector; Step 5: Input the fusion process characterization vector into the deep reinforcement learning model to train the deep reinforcement learning model; Step 6: For the recycling preparation process to be monitored, collect real-time sensor data in real time through multimodal sensing technology. Based on the real-time sensor data, quality prediction model, and deep reinforcement learning model, obtain the action output by the deep reinforcement learning model; Use this action as an optimization strategy to adjust the parameters of the recycling preparation process to be monitored; The method of obtaining the action output by the deep reinforcement learning model based on the real-time sensor data, quality prediction model, and deep reinforcement learning model is as follows: Construct a feature splicing vector through the real-time sensor data and the parameter values of the controllable key parameters of the current recycling preparation process; Input the feature splicing vector into the quality prediction model to obtain the predicted value of the quality indicator at the subsequent t0 moment output by the quality prediction model; Based on the predicted value of the quality indicator and the monitoring data of each sensor, construct a corresponding fusion process characterization vector; Collect all the fusion process characterization vectors in the time segment where the current moment is located to form a corresponding characterization vector set; Input this characterization vector set into the deep reinforcement learning model to obtain the action output by the Actor network of the deep reinforcement learning model.

2. The intelligent optimization control method for the preparation process of battery material recycling according to claim 1, wherein The collection of the training sample data includes the following steps: Step 11: Prepare N batches of waste battery anode and cathode materials with different sources and compositions; N is the number of selected batches; Step 12: In the test environment, cycle and run the complete recycling preparation process flow for each batch of waste battery anode and cathode materials, and record the monitoring data of each sensor in real time, and synchronously record the controllable key parameters during the operation process; Step 13: Measure and record the index value of the quality indicator of the battery material obtained after extraction and recycling as the quality indicator label; Step 14: Combine the monitoring data collected in real time by each sensor of all batches of waste battery anode and cathode materials, the real-time controllable key parameters, and the quality indicator label to form training sample data.

3. The intelligent optimization control method for the battery material recycling preparation process according to claim 2, wherein Training a quality prediction model with product quality indicators as the output based on the training sample data includes the following steps: Step 21: Convert the training sample data into sample input features; Step 22: Use the sample input features at each moment as the input and the predicted value of the quality indicator at the subsequent t0 moment as the output to train the quality prediction model; t0 is the preset time period duration.

4. The intelligent optimization control method for the battery material recycling preparation process according to claim 3, wherein The conversion of the training sample data into sample input features includes the following steps: Step 211: For the positive and negative electrode materials of waste batteries in each batch, preprocess the monitoring data of the sensors at each corresponding moment, and splice the preprocessed monitoring data and controllable key parameters to obtain a feature splicing vector at each moment; Step 212: Preset a time window t; for each moment during the operation of the positive and negative electrode materials of waste batteries in each batch, collect the monitoring parameter sequence of each sensor data and the controllable parameter sequence of each controllable key parameter corresponding to this moment; the monitoring parameter sequence is a sequence formed by arranging the parameter values collected by the corresponding sensor in chronological order within t moments before this moment; the controllable parameter sequence is a parameter value sequence formed by arranging the parameter values of the corresponding controllable key parameters in chronological order within t moments before this moment; Step 213: The monitoring parameter sequence and the controllable parameter sequence at each moment form the sample input features.

5. The intelligent optimization control method for the battery material recycling preparation process according to claim 4, characterized in that, The method for training the quality prediction model is as follows: Use the sample input features at each moment as the input of the quality prediction model. The quality prediction model takes the predicted value of the quality index at the subsequent t0 moment of this moment as the output, takes the quality index label corresponding to the subsequent t0 moment of this moment as the prediction target, takes the difference between the predicted value of the quality index and the quality index label as the prediction error, and takes minimizing the sum of the prediction errors as the training target; train the quality prediction model until the sum of the prediction errors reaches convergence and then stop training; the quality prediction model is a time series prediction model.

6. The intelligent optimization control method for the battery material recycling preparation process according to claim 5, wherein The construction method of the feature splicing vector is as follows: Perform vectorization operations on the monitoring data of each sensor at each moment to obtain the corresponding parameter feature vectors; For each controllable key parameter, use its respective parameter values to form a controllable parameter feature vector; Splice the parameter feature vectors of all sensors and the controllable parameter feature vectors at each moment of the positive and negative electrode materials of waste batteries to obtain a feature splicing vector.

7. The intelligent optimization control method for the battery material recycling preparation process according to claim 6, wherein The method for constructing the state space, action space, and reward function of the deep reinforcement learning model based on the prediction result of the quality prediction model is as follows: The deep reinforcement learning model is set as an Actor-Critic model; Mark the state space as S. The state space S reflects the real-time state of the entire recycling and preparation process. The construction method of the state space is as follows: Segment the recycling and preparation process according to a fixed time window; the size of the fixed time window is t0; For each time segment, use the set of representation vectors formed by arranging the fusion process representation vectors at each moment in chronological order as the state space; Mark the action space as A. The action space A includes the adjustment amount of the controllable key parameters output by the Actor model. Mark the parameter adjustment action taken in the next time segment as a; Mark the reward function as Q. The reward value function Q = r(s, a), which measures the effect of performing action a in the current state s; Design the agent of the deep reinforcement learning model as an RNN-based Actor-Critic structure; The input of the agent is the state s; The Critic network outputs the state value V(s); The Actor network outputs the action probability π(a|s).

8. The intelligent optimization control method for the battery material recycling preparation process according to claim 7, characterized in that The construction method of the fusion process characterization vector is as follows: Collect the feature splicing vectors corresponding to each moment during the recycling and preparation process of the positive and negative electrode materials of waste batteries; Input the feature splicing vector of each moment into the quality prediction model to obtain the predicted value of the quality index at the subsequent t0 moment output by the quality prediction model; Use a one-dimensional convolutional network to extract the local feature sequence from the monitoring parameter sequence of each item of sensor data corresponding to each moment; Input the local feature sequence into the LSTM model to obtain the global context encoding vector C_t; Splice the predicted value of the quality index and the context encoding C_t to obtain the fusion process characterization vector.

9. An intelligent optimization control system for the battery material recycling preparation process, which is used to implement the intelligent optimization control method for the battery material recycling preparation process described in any one of claims 1-8, and is characterized in that, It includes a sample data collection module, a prediction model training module, a reinforcement model training module, and an action decision-making module; among them, each module is connected electrically; The sample data collection module uses multi-modal sensing technology to collect training sample data during the battery material recycling and preparation process in the test environment, and sends the training sample data to the prediction model training module; The prediction model training module trains a quality prediction model with the product quality index as the output based on the training sample data, and sends the quality prediction model to the reinforcement model training module and the action decision-making module; The reinforcement model training module constructs the state space of the deep reinforcement learning model based on the prediction result of the quality prediction model, constructs the action space and the reward function of the deep reinforcement learning model, fuses the prediction result of the quality prediction model and the monitoring data of the sensor in the test environment through the feature fusion strategy to generate the fusion process characterization vector, inputs the fusion process characterization vector into the deep reinforcement learning model, trains the deep reinforcement learning model, and sends the deep reinforcement learning model to the action decision-making module; The action decision-making module, for the recycling and preparation process to be monitored, collects real-time sensor data in real time through multi-modal sensing technology, and obtains the action output by the deep reinforcement learning model based on the real-time sensor data, the quality prediction model, and the deep reinforcement learning model; uses this action as an optimization strategy to adjust the parameters of the recycling and preparation process to be monitored.

10. An electronic device, characterized in that, It includes: A processor and a memory, where, The memory stores a computer program that can be called by the processor; The processor executes the intelligent optimization control method for the battery material recycling and preparation process described in any one of claims 1-8 in the background by calling the computer program stored in the memory.

11. A computer-readable storage medium, characterized in that, It stores an erasable computer program on it; When the computer program runs on the computer device, it enables the computer device to execute the intelligent optimization control method for the battery material recycling and preparation process described in any one of claims 1-8 in the background.

Citation Information

Patent Citations

  • A recycling discharge control method and related device for waste lithium batteries

    CN117810570B

  • Nanoscale structure processing prediction and monitoring method and device based on deep learning

    CN117268539A