Lithium battery assembly line predictive maintenance method and system based on deep reinforcement learning

By using a multi-agent system based on deep reinforcement learning, combining a global coordinator and agents, predictive maintenance of lithium battery assembly lines was achieved, solving the complexity and conflict issues in equipment maintenance decisions and improving production efficiency and reliability.

CN121258490BActive Publication Date: 2026-03-24HEFEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

The complexity and randomness of continuous lithium battery production and assembly lines make it difficult to implement preventive maintenance accurately. This can lead to equipment maintenance decisions causing downstream blockages or upstream backlogs, and there are conflicts between decisions made by multiple equipment.

Method used

A deep reinforcement learning-based approach is adopted. By constructing a multi-agent system, a global coordinator and agents are used to dynamically correlate information such as equipment health and work-in-process load in the buffer zone, generating local predictive maintenance decisions. The equipment maintenance decisions are then made in combination with a greedy strategy and an attention mechanism.

Benefits of technology

It significantly reduces unplanned downtime, improves the production efficiency and operational reliability of lithium battery assembly lines, and avoids production line blockages and equipment dependency conflicts caused by local maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121258490B_ABST
    Figure CN121258490B_ABST
Patent Text Reader

Abstract

The application provides a lithium battery assembly line predictive maintenance method and system based on deep reinforcement learning, and relates to the technical field of deep reinforcement learning. According to the current health degree of the equipment on the lithium battery assembly line, the work-in-process loading degree, the running state and the remaining duration of the maintenance task of the upstream and downstream buffer zones, the state vector of the equipment is constructed; the state vectors of all the equipment are input into the global coordinator to obtain the global context vector; the state vector of each equipment and the global context vector are input into the agent corresponding to the equipment to obtain the local predictive action decision; and the local predictive action decision is converted into the equipment action for execution. The introduction of the global coordinator effectively eliminates the conflicts among multi-equipment decisions, makes the maintenance decision based on the whole line state rather than the isolated equipment, the agent makes the localized decision in combination with the global context and the local state, the individual characteristics of the equipment are reserved and blind maintenance is avoided, and the predictive maintenance is implemented at the optimal time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep reinforcement learning, in particular to a lithium battery assembly line predictive maintenance method and system based on deep reinforcement learning. BACKGROUND

[0002] Lithium-ion batteries have become the current mainstream electrochemical energy storage solution due to their excellent performance and cost advantage. With the continuous growth of the electric vehicle market, the production demand for lithium batteries is expanding, and higher requirements are placed on the efficiency and reliability of the production line.

[0003] In the manufacturing field, preventive maintenance, as a common strategy, reduces the risk of random failures by timely replacing or repairing aging equipment and parts. However, in a lithium battery continuous production assembly line with intermediate buffers, the complexity and randomness of the production line itself are more diverse, making it difficult to determine when and where preventive maintenance should be implemented. SUMMARY

[0004] The problem to be solved by the present application is the complexity and randomness of the lithium battery continuous production assembly line, making it difficult to accurately implement preventive maintenance.

[0005] To solve the above problems, in a first aspect, the present application provides a lithium battery assembly line predictive maintenance method based on deep reinforcement learning, comprising:

[0006] constructing a state vector of the equipment according to the current health degree of the equipment on the lithium battery assembly line, the work-in-process loading degree of the upstream buffer, the work-in-process loading degree of the downstream buffer, the running state and the remaining duration of the maintenance task, wherein the lithium battery assembly line comprises a plurality of equipment and a plurality of buffers;

[0007] inputting the state vectors of all equipment into a global coordinator of a trained deep reinforcement learning model to obtain a global context vector, wherein the deep reinforcement learning model comprises a plurality of agents equipped on the equipment and a global coordinator, and each equipment on the lithium battery assembly line is equipped with an agent corresponding thereto;

[0008] inputting the state vector of each equipment and the global context vector into the agent corresponding to the equipment to obtain a local predictive action decision, wherein the local predictive action decision comprises 0 and 1, 0 indicating no maintenance action, and 1 indicating a predictive maintenance action;

[0009] converting the local predictive action decision into an equipment maintenance action for execution.

[0010] Optionally, the step of inputting the state vectors of all equipment into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector comprises:

[0011] inputting the state vector of each device into the corresponding agent, generating a query vector, a key vector and a value vector through neural network mapping;

[0012] For the current agent, the query vector is calculated with the key vectors output by the agents of all devices, and a set of attention weights is obtained through function normalization;

[0013] The attention weights are used as coefficients to perform weighted summation on the value vectors of all devices to generate a global context vector, which contains the health degree of all devices, the work-in-process loading degree of all buffer zones, the maintenance activity distribution of all devices and the downtime distribution of all devices.

[0014] Optionally, the inputting of the state vector of each device and the global context vector into the agent corresponding to the device to obtain a local predictive action decision comprises:

[0015] The state vector of the current agent itself and the global context vector are spliced and input into the current agent to obtain an action probability distribution;

[0016] A greedy strategy is adopted to select the final action from the action probability distribution as the local predictive action decision.

[0017] Optionally, before the state vector of the device is constructed according to the current health degree of the device on the lithium battery assembly line, the work-in-process loading degree of the upstream buffer zone, the work-in-process loading degree of the downstream buffer zone, the running state and the remaining duration of the maintenance task, it further comprises:

[0018] determining whether the current health degree of the device is less than a health degree threshold or the random failure probability of the device is greater than a failure probability threshold;

[0019] If the current health degree is less than the health degree threshold or the random failure probability is greater than the failure probability threshold, corrective maintenance is performed;

[0020] If the current health degree is greater than or equal to the health degree threshold or the random failure probability is less than or equal to the failure probability threshold, it is predicted whether to perform predictive maintenance.

[0021] Optionally, the lithium battery assembly line comprises a positive electrode cutting and rolling device, a negative electrode cutting and rolling device, an ear pre-welding device, a connecting piece welding device, a cover plate welding device and a top cover full-welding device;

[0022] The health degree model of the positive electrode cutting and rolling device and the negative electrode cutting and rolling device is

[0023]

[0024] wherein, Indicates the first The equipment is in Health status at all times; The attenuation rate constant; In order to be in The winding alignment deviation of the positive and negative electrode cutting and winding equipment is measured at all times. The initial winding alignment deviation of the equipment is 0. This represents the critical deviation from the fault. The continuous operating time of the equipment refers to the time from the last maintenance until... The duration of continuous operation;

[0025] The health model of the electrode pre-welding equipment and the connecting piece welding equipment is as follows:

[0026]

[0027] in, Indicates that the device is in The wear degree of the welding head at any given time; the initial wear degree of the welding head is 0. This is the wear sensitivity coefficient;

[0028] The health model for the cover plate welding equipment and the top cover full welding equipment is as follows:

[0029]

[0030] in, This is the power instability penalty coefficient; For the first The equipment is in The actual output power at any given time; For the first The set power value of the device.

[0031] Optionally, the random failure probability of the device is

[0032]

[0033] in, Indicates the first The probability of random failure of a device; This represents the highest probability of random failure throughout the entire lifespan of the equipment. Indicates the first The equipment is in Health status at all times; For both positive and negative electrode slitting equipment, the failure probability shape parameter is... For electrode pre-welding equipment and connecting piece welding equipment, the failure probability shape parameter For cover plate welding equipment and top cover full welding equipment, the failure probability shape parameter .

[0034] Optionally, the predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning further includes: training the deep reinforcement learning model, wherein the deep reinforcement learning model further includes a centralized critic network;

[0035] Training methods for deep reinforcement learning models include:

[0036] Randomly sample from the stored device operation data to obtain the state data of all devices and the corresponding local predictive action decisions output by the intelligent agent;

[0037] The status data of all devices are input into a centralized critic network. With the goal of maximizing the global reward, the optimal action strategy is obtained. The action strategy refers to a rule system that decides to execute specific maintenance actions based on the current status information of the devices.

[0038] Based on the optimal action policy and local predictive action decisions, the parameters of the local decision network in the agent and the parameters of the attention network in the global coordinator are updated by minimizing the loss function composed of temporal difference errors and using gradient descent.

[0039] Optionally, the global reward is

[0040]

[0041] in, Indicates revenue per unit of product; This indicates the output of products produced by the lithium battery assembly line. Indicates the first Maintenance costs incurred from corrective maintenance of the equipment; Indicates the first Maintenance costs incurred from preventative maintenance of equipment; Indicator variables that are either 0 or 1. Indicates the first The equipment is in Whether to perform corrective maintenance at all times Indicates the first The equipment is in Whether preventative maintenance is performed at all times; This represents the production loss cost incurred per unit of equipment downtime; Indicates the first The operating status of the equipment Indicates that the first device is in It is constantly in a shutdown state; This indicates that the second device is in It is constantly in a shutdown state.

[0042] Optionally, the optimal action strategy is:

[0043]

[0044] in, Represents the strategy set for a lithium battery assembly line; Indicating in strategy Next deadline Total revenue generated by the production line.

[0045] Secondly, the present invention also provides a predictive maintenance system for lithium battery assembly lines based on deep reinforcement learning, comprising:

[0046] The state vector construction module is used to construct the state vector of the equipment based on the current health of the equipment on the lithium battery assembly line, the work-in-process loading degree of the upstream buffer, the work-in-process loading degree of the downstream buffer, the operating status, and the remaining duration of the maintenance task. The lithium battery assembly line includes multiple equipment and multiple buffers.

[0047] The coordination module is used to input the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent.

[0048] The prediction module is used to input the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions. The local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action.

[0049] The conversion and execution module is used to convert local predictive action decisions into equipment maintenance actions for execution.

[0050] This invention provides a predictive maintenance method and system for lithium battery assembly lines based on deep reinforcement learning. Compared with existing technologies, it has the following advantages:

[0051] By dynamically linking equipment health status with the work-in-process loading levels of upstream and downstream buffers, this method avoids downstream congestion or upstream backlog problems that may arise from localized maintenance decisions based solely on equipment health. Simultaneously, the introduction of a global coordinator effectively eliminates conflicts between multi-device decisions, ensuring that maintenance decisions are based on the overall line status rather than isolated equipment, thus resolving the maintenance cascading effects caused by equipment interdependence in continuous production lines. Furthermore, the agent combines its own state with the global context to make localized decisions, preserving individual equipment characteristics while avoiding blind maintenance, thereby enabling predictive maintenance to be implemented at the optimal time. This scheme significantly reduces unplanned downtime through these mechanisms, improving the production efficiency and operational reliability of lithium battery assembly lines. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A flowchart illustrating a predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning, provided in an embodiment of the present invention.

[0054] Figure 2 This is a schematic diagram of a lithium battery production and assembly line provided in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of a predictive maintenance system for a lithium battery assembly line based on deep reinforcement learning. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application are described clearly and completely. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0058] like Figure 1 As shown in the embodiment of this application, a predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning is provided, comprising:

[0059] S1: Construct the state vector of the equipment based on the current health of the equipment on the lithium battery assembly line, the work-in-process loading of the upstream buffer, the work-in-process loading of the downstream buffer, the operating status, and the remaining duration of the maintenance task. The lithium battery assembly line includes multiple equipment and multiple buffers.

[0060] S2: Input the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent.

[0061] S3: Input the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions, wherein the local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action.

[0062] S4: Transform local predictive action decisions into equipment maintenance actions for execution.

[0063] Specifically, for example, the lithium battery can be a lithium-ion power battery, and the lithium battery assembly line includes positive electrode slitting equipment, negative electrode slitting equipment, tab pre-welding equipment, connecting piece welding equipment, cover plate welding equipment, and top cover full welding equipment. When the current health of the tab pre-welding equipment is 0.82, the upstream buffer work-in-process loading is 60%, and the downstream buffer work-in-process loading is 40%, its state vector is constructed as [0.82, 0.60, 0.40, 1, 0]. This vector, along with the state vectors of other equipment, is input into the global coordinator to generate a global context vector. This context vector is then concatenated with the device's own state vector and input to the agent of the tab pre-welding equipment, outputting a local decision such as "perform preventative maintenance." Here, the health reflects the degree of equipment degradation, while the buffer work-in-process loading reflects the production line's buffering capacity. This combination is used to accurately quantify the impact of equipment maintenance on the smoothness of the entire production line.

[0064] In this optional embodiment, by dynamically linking the equipment health status with the work-in-process loading levels of upstream and downstream buffers, this method avoids downstream congestion or upstream backlog problems that may be caused by localized maintenance decisions based solely on equipment health. Simultaneously, the introduction of a global coordinator effectively eliminates conflicts between multi-device decisions, ensuring that maintenance decisions are based on the overall line status rather than isolated equipment, thus resolving the maintenance chain reaction problem caused by equipment interdependence in continuous production lines. Furthermore, the agent combines its own state with the global context to make localized decisions, preserving individual equipment characteristics while avoiding blind maintenance, thereby enabling predictive maintenance to be implemented at the optimal time. This scheme significantly reduces unplanned downtime through the above mechanisms, improving the production efficiency and operational reliability of lithium battery assembly lines.

[0065] The following is a detailed description of each step.

[0066] S1: Construct a state vector for each device based on its current health, the work-in-process load in the upstream buffer, the work-in-process load in the downstream buffer, its operating status, and the remaining duration of its maintenance task. The lithium battery assembly line comprises multiple devices and multiple buffers. Each device's current health, the work-in-process load in the upstream buffer, and the work-in-process load in the downstream buffer construct its own state vector. Devices that have not performed maintenance actions have a remaining maintenance task duration of 0.

[0067] Specifically, such as Figure 2 As shown, the lithium battery assembly line is a production assembly line containing 6 machines and 4 buffer zones. The lithium battery assembly line includes positive electrode slitting and coiling equipment, negative electrode slitting and coiling equipment, tab pre-welding equipment, connector welding equipment, cover plate welding equipment, and top cover full welding equipment. The positive electrode slitting and coiling equipment is denoted as... The negative electrode cutting and rolling equipment is denoted as The electrode pre-welding equipment is denoted as The connecting piece welding equipment is denoted as The cover plate welding equipment is denoted as The equipment for fully welding the top cover is recorded as The intermediate buffer zones are designated as Buffer Zone 1, Buffer Zone 2, Buffer Zone 3, and Buffer Zone 4 according to the battery production and assembly sequence. The positive electrode cutting and coiling equipment and the negative electrode cutting and coiling equipment operate in parallel, cutting and coiling the incoming materials to serve as the positive and negative electrode materials for lithium battery work-in-process, and inputting the work-in-process into Buffer Zone 1 of the production assembly line. The tab pre-welding equipment removes the work-in-process from Buffer Zone 1 and performs tab pre-welding processing, placing it into Buffer Zone 2. The connecting piece welding equipment removes the work-in-process from Buffer Zone 2, firmly welding the tabs to the metal connecting pieces, and placing it into Buffer Zone 3. The cover plate welding equipment removes the work-in-process from Buffer Zone 3 and firmly welding the connecting pieces to the cover plate, placing it into Buffer Zone 4. The top cover full welding equipment removes the work-in-process from Buffer Zone 4 and performs sealing welding between the battery cover plate and the aluminum shell, producing the finished product.

[0068] A state space is constructed, encompassing the parameters of each parameter state of the lithium-ion battery assembly line equipment and the buffer zones between them. Based on the assembly line's operational rules, the maintenance action space of each equipment within its respective state space is determined, and a global reward function is established that comprehensively considers overall equipment downtime, production throughput, and maintenance costs. Based on the state space, maintenance action space, and global reward function, a multi-agent deep reinforcement learning model employing a "centralized training and distributed execution" framework is constructed. This deep reinforcement learning model includes an agent assigned to each equipment, a global coordinator based on an attention mechanism, and a centralized critic network. The critic network is used to evaluate the global Q-value of joint actions during the training phase to guide policy optimization for all agents. A local decision network is set within each agent, comprising three fully connected layers. The first layer has an input neuron count equal to the sum of the dimensions of the device's own state vector and the global context vector; the second layer has 64 neurons; and the third layer has an output neuron count equal to the spatial dimension of the actions.

[0069] During the implementation of predictive maintenance methods, it is necessary to calculate the health status of each piece of equipment in the lithium battery assembly line in real time. The health status model of each piece of equipment in the production assembly line is analyzed separately to obtain the health status parameters for each piece of equipment. The logic is as follows:

[0070] The winding alignment deviation of the positive electrode cutting and winding equipment and the negative electrode cutting and winding equipment were obtained respectively. The wear of the welding head of the electrode tab pre-welding equipment and the connecting piece welding equipment, as well as the stability of the laser output power of the cover plate welding equipment and the top cover full welding equipment, were obtained. The real-time continuous working time of each machine was also obtained.

[0071] The health models for the positive electrode winding equipment and the negative electrode winding equipment are based on the following formulas:

[0072]

[0073] in, Indicates the first The equipment is in Health status at all times; This is the decay rate constant, which determines the rate at which health decreases over time; the specific value is provided by the equipment manufacturer. In order to be in The real-time winding alignment deviation of the positive and negative electrode cutting and winding equipment is measured. The initial winding alignment deviation of the equipment is 0. As the equipment runs, the actual real-time winding deviation... Increase; The fault threshold deviation refers to the random fault that occurs when the real-time winding alignment deviation of the equipment exceeds the fault threshold deviation. The specific value is provided by the equipment manufacturer. The continuous operating time of the equipment refers to the time from the last maintenance until... The duration of continuous operation.

[0074] The health model for the tab pre-welding equipment and the connecting piece welding equipment is based on the following formula:

[0075]

[0076] in, Indicates that the device is in The wear degree of the welding head at any given time; the initial wear degree of the welding head is 0. The wear sensitivity coefficient, A higher value indicates a greater impact of wear and tear on health, determined based on historical data.

[0077] The health model for the cover plate welding equipment and the top cover full welding equipment is based on the following formula:

[0078]

[0079] in, This is the power instability penalty factor. The larger the value, the greater the impact of power fluctuations on health. For the first The equipment is in The standard deviation of the actual output power at any given time will gradually increase as the equipment operates. For the first The power settings for the equipment are provided by the equipment manufacturer.

[0080] In addition, prior to S1, predictive maintenance methods also included:

[0081] S11: Determine whether the current health status of the device is less than the health status threshold or whether the random failure probability of the device is greater than the failure probability threshold.

[0082] Specifically, all devices start with a health level of 1, which gradually decreases as the devices operate. Device health level Determines the probability of random failure of the equipment As the health level of the equipment decreases, its probability of random failure gradually increases. The probability of random failure of the equipment is as follows:

[0083]

[0084] in, Indicates the first The probability of random failure of a device; This represents the highest probability of random failure throughout the entire lifecycle of the equipment, calculated based on historical data. Indicates the first The equipment is in Health status at all times; The failure probability shape parameter is given. For the positive and negative electrode slitting equipment, due to mechanical wear and fatigue accumulation effects, the value range of the failure probability shape parameter is as follows: Characterizing typical loss failure modes; due to transducer aging and electrical component degradation, the failure probability shape parameter values ​​for electrode pre-welding equipment and connecting piece welding equipment range from [value missing]. The failure probability shape parameter range for cover plate welding equipment and top cover full welding equipment is affected by the attenuation of the optical system and the performance drift of the cooling system. The shape parameters of the aforementioned equipment types are all obtained based on the actual failure time data recorded during the historical operation of these equipment types, and are identified through statistical methods such as maximum likelihood estimation and least squares fitting.

[0085] S12: If the current health status is less than the health status threshold or the random failure probability is greater than the failure probability threshold, then corrective maintenance is performed.

[0086] Specifically, when equipment health drops to a health threshold or the probability of random failure exceeds the failure probability threshold, the equipment cannot continue processing and requires corrective maintenance. During assembly line operation, equipment with declining health can be proactively shut down and preventive maintenance performed. Corrective maintenance takes time and incurs maintenance costs. After corrective maintenance, all equipment performance is restored to near-new levels, continuous operating time returns to 0, and health level returns to 1. After preventive maintenance, the winding alignment deviation of the positive / negative electrode switching equipment is partially corrected. , The preventative maintenance efficiency coefficient is provided by the equipment manufacturer. The corresponding value corresponds to the recovery of continuous working time to 0 and an increase in health. After preventative maintenance is completed on the tab pre-welding equipment and connecting piece welding equipment, the wear on their welding heads partially recovers. When continuous working time returns to 0, the health level increases by the corresponding value; after preventative maintenance is completed on the cover plate welding equipment and the top cover full welding equipment, the standard deviation of their actual output power... reduce, Continuous working time is restored to 0.

[0087] S13: If the current health level is greater than or equal to the health level threshold or the random failure probability is less than or equal to the failure probability threshold, then predict whether to perform predictive maintenance, i.e., start executing the predictive maintenance method of S1-S4. The above maintenance action space consists of active actions that the agent can choose. When the device health level is lower than the health level threshold, the device will be forced into a corrective maintenance state, which is not an optional action of the agent.

[0088] S2: Input the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent. This step specifically involves:

[0089] S21: Input the state vector of each device into the corresponding agent, and generate query vector, key vector and value vector through neural network mapping.

[0090] Specifically, the local state vector of each device agent... Through a shared fully connected network, corresponding query vectors are generated by mapping. Key vector Sum value vector .

[0091] S22: For the current agent, calculate the similarity between the query vector and the key vectors output by the agents of all devices, and normalize them using a function to obtain a set of attention weights. These weights represent the impact of the state of each other device on the agent in the current state. The relative importance of making decisions.

[0092] Specifically, for a specific device intelligent agent that needs to make a decision at present... Compare its query vector with the key vectors of all devices in the assembly line. Perform dot product calculations to assess its correlation with each device state; then transmit the dot product results via... Function normalization yields a set of attention weights. This weight dynamically reflects the device's state under the current system conditions. The status of the device The importance of making optimal maintenance decisions.

[0093] S23: Using attention weights as coefficients, perform a weighted summation of the value vectors of all devices to generate a global context vector. The global context vector includes the health of all devices, the work-in-progress loading of all buffers, the distribution of maintenance activities of all devices, and the distribution of downtime of all devices.

[0094] Specifically, the calculated attention weights are used as coefficients to perform a weighted summation of the value vectors of all devices, resulting in a weighted sum for each device. Generate a customized global context vector rich in collaborative information. .

[0095]

[0096] S3: Input the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions. These local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action. This step specifically involves:

[0097] S31: Concatenate the current agent's own state vector with the global context vector and input them into the current agent to obtain the action probability distribution.

[0098] Specifically, each device agent will concatenate the vectors. The input is fed into its local decision network, which outputs an action probability distribution and selects the final action through sampling or a greedy strategy.

[0099] S32: A greedy strategy is adopted to select the final action as the local predictive action decision from the action probability distribution. Each local decision network outputs the action probability, and the control system uses the greedy strategy to select the action with the highest probability to generate the final maintenance action decision.

[0100] S4: Transform local predictive action decisions into equipment maintenance actions for execution. The updated local predictive action decisions are then translated into specific equipment maintenance action selections and executed. The system collects the new equipment status and adjacent buffer parameters in real time after the action is executed. The collected sequence data is stored in the experience replay pool.

[0101] In an optional embodiment of this application, the predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning further includes: training the deep reinforcement learning model, wherein the deep reinforcement learning model further includes a centralized critic network;

[0102] The training methods for the deep reinforcement learning model include:

[0103] S110: Randomly sample from the stored device operation data to obtain the state data of all devices and the corresponding local predictive action decisions output by the intelligent agent, wherein the state data includes state vectors.

[0104] S120: Input the state data of all devices into a centralized critic network to obtain the optimal action policy with the goal of maximizing the global reward.

[0105] S130: Based on the optimal action policy and local predictive action decisions, the parameters of the local decision network in the agent and the parameters of the attention network in the global coordinator are updated by minimizing the loss function formed by the temporal difference error and using the gradient descent method.

[0106] Specifically, the collected sequence data is stored in an experience replay pool. During training, batch data is randomly sampled, and a centralized critic network calculates a global Q-value estimate with the global reward as the optimization objective. The policies of each agent are updated by maximizing this global value. The parameters of the local decision networks in the agents and the parameters of the attention network in the global coordinator are updated by minimizing the loss function composed of temporal difference errors. The neural network parameters in the attention network used to generate query, key, and value vectors are optimized based on the gradient of the global Q-value during backpropagation as part of the model. After training, the updated local decision network parameters and attention network parameters are deployed to the agents of each device on the production line, and the system then enters the next round of decision-making loop.

[0107] Each intermediate buffer has a work-in-process loading level. , indicating the first An intermediate buffer in The work-in-process loading rate at any given time, its range being: The initial work-in-process inventory in each buffer is 0. When the work-in-process load in the upstream buffer of a device on the assembly line is 0, the device enters a starved state and idles and stops. When the work-in-process load in its upstream buffer is greater than 0, it resumes operation. The positive and negative electrode cutting and winding devices on the assembly line do not have upstream buffers, so they do not enter a starved state. When the work-in-process load in the downstream buffer of a device on the assembly line is full, the device enters a blocked state and idles and stops. When the work-in-process load in its downstream buffer is not full, it resumes operation. The end-of-line top cover welding device on the assembly line does not have a downstream buffer, so it does not enter a blocked state. The work-in-process load is the degree to which work-in-process is loaded in the buffer. When there is no work-in-process in the buffer, the load is 0; when the buffer is full of work-in-process, the load is 1. Therefore, the range of the work-in-process load is... .

[0108] In the lithium battery production assembly line model, since the positive and negative electrode slitting machines operate synchronously, when one machine stops, the other must also be idle and shut down. Therefore, the maintenance and downtime costs of both positive and negative electrode slitting machines must be considered. For the tab pre-welding equipment, connector welding equipment, cover plate welding equipment, and top cover full welding equipment in the assembly line, the maintenance costs incurred from performing maintenance operations on these machines must be considered. In the assembly line model, a corresponding reward is awarded for each product produced.

[0109] For intermediate devices Its state vector is:

[0110]

[0111] For equipment and Since there is no upstream buffer, its state vector is:

[0112]

[0113] For end devices Since there is no downstream buffer, its state vector is:

[0114]

[0115] in, Indicates device At any moment Health status at that time ; Represents a buffer exist Work-in-process loading rate in the buffer zone at that time. . It is a buffer Maximum work-in-process loading ; Indicates the first The operating status of the equipment Indicates device exist It is constantly in a shutdown state. Indicates device exist It is always running. Indicates device exist It is constantly under maintenance. Indicates the first The remaining duration of the maintenance task currently in progress on the equipment at time t is used to determine when the equipment can be put back into use.

[0116] The state transition of the device is recorded as:

[0117]

[0118] in, Indicates time Time Machine The environment; Indicates in Machines in an environment Local predictive action decision-making; Represents local predictive action decision-making The reward received; Indicates the execution of an action After machine The environment reached. The state space is: The local predictive action decision space is: The local predictive action decision model is as follows:

[0119]

[0120] In the process of predictive maintenance of the assembly line, the state space, maintenance action space, and predictive action decisions are stored in the experience replay pool.

[0121] For positive and negative electrode cutting and winding equipment, in time Rewards received It is time arrive The incentive equation for the maintenance costs and equipment downtime costs of the production line is as follows:

[0122]

[0123] in, The variable can be either 0 or 1, representing the device, respectively. exist Whether preventative or corrective maintenance was performed at any given time. This represents the production loss cost incurred per unit of equipment downtime.

[0124] For the remaining four pieces of equipment on the assembly line, in time Rewards received It is time arrive The reward equation for the maintenance costs incurred during equipment maintenance is as follows:

[0125]

[0126] For the entire production assembly line, the entire system in time The reward received is time. arrive The global reward is calculated by subtracting the maintenance costs incurred from all maintenance actions performed on equipment within the production line from the revenue generated from products produced on the production line.

[0127]

[0128] in, Indicates revenue per unit of product; This indicates the output of products produced by the lithium battery assembly line. Indicates the first Maintenance costs incurred from corrective maintenance of the equipment; Indicates the first Maintenance costs incurred from preventative maintenance of equipment; Indicator variables that are either 0 or 1. Indicates the first The equipment is in Whether to perform corrective maintenance at all times Indicates the first The equipment is in Whether preventative maintenance is performed at all times; This represents the production loss cost incurred per unit of equipment downtime; Indicates the first The operating status of the equipment Indicates that the first device is in It is constantly in a shutdown state; This indicates that the second device is in It is constantly in a shutdown state.

[0129] The optimal action strategy is:

[0130]

[0131] in, Represents the strategy set for a lithium battery assembly line; Indicating in strategy Next deadline Total revenue generated by the production line.

[0132] In summary, compared with existing technologies, it has the following beneficial effects:

[0133] 1. This application not only considers the health of the equipment, but also comprehensively takes into account various dynamic factors such as the maintenance duration, the work-in-process load in the buffer zone, maintenance activities, and production losses due to equipment downtime, making the maintenance plan more adaptable. This adaptability helps reduce unnecessary downtime caused by fixed maintenance plans, further improving the responsiveness and stability of the production system.

[0134] 2. This application utilizes multi-agent reinforcement learning technology, employing an attention-based global coordinator, to more accurately identify dynamic changes in the production line and subsequently implement targeted maintenance measures. This method can adjust maintenance priorities in a timely manner, ensuring the entire production line is in optimal operating condition and preventing production interruptions due to equipment failure.

[0135] like Figure 3 As shown in the figure, an embodiment of this application provides a predictive maintenance system for lithium battery assembly lines based on deep reinforcement learning, comprising:

[0136] The state vector construction module 10 is used to construct the state vector of the equipment based on the current health of the equipment on the lithium battery assembly line, the work-in-process loading degree of the upstream buffer, the work-in-process loading degree of the downstream buffer, the operating status, and the remaining duration of the maintenance task. The lithium battery assembly line includes multiple equipment and multiple buffers.

[0137] The coordination module 20 is used to input the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent.

[0138] The prediction module 30 is used to input the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions, wherein the local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action.

[0139] The conversion execution module 40 is used to convert local predictive action decisions into equipment maintenance actions for execution.

[0140] In this embodiment, the beneficial effects of the predictive maintenance system for lithium battery assembly lines based on deep reinforcement learning are similar to those of the predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning described above, and will not be repeated here.

[0141] The deployment of deep reinforcement learning models is divided into two stages: offline training and online deployment. The offline training stage is conducted in a simulation environment, aiming to train a mature and usable multi-agent predictive maintenance policy model.

[0142] Offline training phase

[0143] S101: System Modeling and Parameter Initialization. Construct a virtual simulation environment in a computer. Initialize the parameters required for the deep reinforcement learning model, including:

[0144] Equipment health parameters: critical deviation for faults in positive and negative electrode coil cutting equipment, welding head wear sensitivity coefficient, laser welding power instability penalty coefficient, and health decay rate constant. Maintenance parameters: cost and time of preventative maintenance, cost and time of corrective maintenance, and preventative maintenance efficiency coefficient for each piece of equipment. Production and cost parameters: equipment downtime loss cost and maximum work-in-process loading in each buffer zone. Fault parameters: maximum random fault probability and fault probability shape parameters for each piece of equipment.

[0145] S102: Model Construction and Initialization. Construct a multi-agent deep reinforcement learning model with six agents. Each agent is equipped with a local decision network containing three fully connected layers. Simultaneously, construct a global coordinator based on an attention mechanism and a centralized critic network. Initialize the parameters of all neural networks.

[0146] S103: Simulated Interaction and Experience Collection. In a simulated environment, a group of intelligent agents interacts with the environment based on the current policy.

[0147] At each simulation time step, each agent outputs a local predictive action decision based on its local observations. The simulation environment calculates the new system state and global reward based on these actions, the dynamic model of device health, random failure probabilities, and other internal logic. The interaction data is then processed. As an experience tuple, it is stored in the experience replay pool.

[0148] S104: Model Training and Optimization. Once the amount of data in the experience replay pool reaches a set threshold, model training begins.

[0149] A small batch of empirical data is randomly sampled from the pool. A centralized critic network calculates a global Q-value estimate based on this batch of data and updates the network parameters by minimizing the temporal difference error loss. Using the policy gradient method, guided by the critic network, the local decision network parameters of all agents, as well as the neural network parameters in the attention mechanism that generate query, key, and value vectors, are updated.

[0150] Repeat steps S103 and S104 until the model policy converges, i.e., the long-run average return. It remains stable at a relatively high level.

[0151] The online deployment and execution phase involves deploying the offline-trained model to the actual production line.

[0152] S201: Model Consolidation and Deployment. The local decision network parameters and attention network parameters of each agent, after training convergence, are exported and embedded into the agents and global coordinator of the production line's central control system. It is important to note that during online operation, the system only needs to run the consolidated local decision networks for forward computation, eliminating the need for centralized training and complete backpropagation. This results in high computational efficiency and meets real-time requirements.

[0153] S202: Real-time data acquisition. The following data is acquired in real time through a sensor network deployed across various devices and buffer zones:

[0154] Equipment health data: winding alignment deviation of positive and negative electrode cutting equipment, welding head wear, output power of welding equipment, and continuous operating time of each piece of equipment. Buffer zone data: work-in-process loading level of each buffer zone. Equipment status data: operating status and remaining maintenance time of each piece of equipment.

[0155] S203: Collaborative Maintenance Decision. The central control system executes the following procedures:

[0156] The real-time data collected in step S202 is used to construct the state vectors of each device as defined in this invention. The input is fed into the corresponding pre-defined agent. Each agent outputs the probability of an action, and the control system uses a greedy strategy to select the action with the highest probability, generating the final local predictive action decision.

[0157] S204: Instruction Execution and Loop. The control system issues maintenance instructions to the production line execution layer. If the instruction is to perform predictive maintenance, the system automatically schedules the equipment to enter the maintenance process after completing the current work-in-process processing. The system then returns to step S202 to enter the next round of data acquisition and decision-making loop, forming a closed-loop control.

[0158] Furthermore, to ensure the predictive maintenance system of this application possesses long-term applicability and continuous optimization capabilities, a complete update and iteration mechanism is designed. This mechanism includes update strategies at the following different levels:

[0159] Online fine-tuning of model parameters is a continuous process after online deployment, serving as a fundamental means to ensure model adaptability.

[0160] S301: During normal operation, the system will continuously store the real-time collected status, action, reward and new status sequence data into a dedicated online experience replay pool.

[0161] S302: The system starts a low-priority background computing task to periodically sample data from the online experience replay pool.

[0162] S303: Using the sampled data, incrementally train the parameters of the local decision-making network and attention mechanism network of each deployed agent at a small learning rate.

[0163] S304: After training is complete, the system will verify the performance of the fine-tuned model on the validation set. If the performance is stable or improved, it will seamlessly switch to the new parameters, achieving silent model updates. This process aims to adapt the model to the slow aging of equipment performance, minor changes in production rhythm, etc.

[0164] This stage involves iterative upgrades of the model structure, performed when significant changes occur in the production system or when the model's performance degrades significantly.

[0165] S401: Triggering Condition: The system continuously monitors key performance indicators (KPIs), including but not limited to average revenue per unit time, frequency of unplanned equipment downtime, and the conflict rate between model decisions and expert experience. When any KPI continuously deviates from the preset threshold, the system automatically generates an alarm.

[0166] S402: Shadow Mode Operation: The system constructs a simulation environment that runs in parallel with the online environment and places the latest copy of the current online model into it for operation; this is shadow mode. In this mode, newly collected real-time data is simultaneously input into both the online model and the shadow model, but the decisions made by the shadow model do not affect the actual equipment and are only used for evaluation.

[0167] S403: Offline Retraining: During shadow mode operation, a large-scale new dataset is collected. Based on this dataset, the model is fully retrained in an offline environment. This process may involve adjusting the network structure depth and width or introducing new attention calculation methods.

[0168] S404: Validation and Deployment: The retrained model needs to be fully validated on historical data and a new test set. After successful validation, the old online model is smoothly replaced using hot-swapping technology, completing the iterative upgrade of the model version.

[0169] Based on rapid cross-line adaptation using transfer learning, this solution is enabled when it is necessary to deploy this invention to a new lithium-ion battery assembly line with a similar topology.

[0170] S501: Model reuse: Using a model that has been trained and matured on the original production line as a pre-trained model for the new production line.

[0171] S502: Parameter Freezing and Fine-tuning: Freeze the parameters of the attention mechanism and the front-end layer of the local decision network in the pre-trained model, and fine-tune the final output layer of the local decision network using only the initial running data of the new production line.

[0172] S503: Rapid Deployment: Once the fine-tuning process converges, the adapted model can be deployed to the new production line. This method significantly reduces the data requirements for the new production line and shortens the deployment cycle.

[0173] Co-evolution and multi-objective optimization: To meet the ever-changing optimization goals in long-term operation, the system supports evolution towards multi-objective optimization.

[0174] S601: Objective Expansion: During model iterative training, the reward function is expanded from a single long-term total benefit to multiple objectives that simultaneously consider production throughput, total maintenance cost, and equipment health balance.

[0175] S602: Policy Learning: A multi-objective reinforcement learning algorithm is used to train a model that can output a Pareto optimal solution set. This solution set contains a series of maintenance policies that make different trade-offs among different objectives.

[0176] S603: Strategy Selection: System administrators can flexibly select and activate corresponding maintenance strategies from the solution set according to actual production needs (such as "guarantee delivery", "reduce costs" or "extend lifespan" modes), making system decision-making more strategically flexible.

[0177] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning, characterized in that, include: Determine whether the current health status of the device is less than the health status threshold or whether the random failure probability of the device is greater than the failure probability threshold. If the current health status is less than the health status threshold or the random failure probability is greater than the failure probability threshold, then corrective maintenance is performed; If the current health status is greater than or equal to the health status threshold or the random failure probability is less than or equal to the failure probability threshold, then predict whether to perform predictive maintenance. Based on the current health of the equipment on the lithium battery assembly line, the work-in-process loading rate of the upstream buffer, the work-in-process loading rate of the downstream buffer, the operating status, and the remaining duration of the maintenance task, a state vector of the equipment is constructed. The lithium battery assembly line includes multiple equipment and multiple buffers. The state vectors of all devices are input into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent. The state vector and global context vector of each device are input into the corresponding agent of the device to obtain local predictive action decisions. The local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action. Transform local predictive action decisions into equipment maintenance actions for execution; The lithium battery assembly line includes positive electrode cutting and rolling equipment, negative electrode cutting and rolling equipment, tab pre-welding equipment, connecting piece welding equipment, cover plate welding equipment, and top cover full welding equipment. The health models for the positive electrode winding equipment and the negative electrode winding equipment are as follows: in, Indicates the first The equipment is in Health status at all times; The attenuation rate constant; In order to be in The winding alignment deviation of the positive and negative electrode cutting and winding equipment is measured at all times. The initial winding alignment deviation of the equipment is 0. This represents the critical deviation from the fault. The continuous operating time of the equipment refers to the time from the last maintenance until... The duration of continuous operation; The health model of the electrode pre-welding equipment and the connecting piece welding equipment is as follows: in, Indicates that the device is in The wear degree of the welding head at any given time; the initial wear degree of the welding head is 0. This is the wear sensitivity coefficient; The health model for the cover plate welding equipment and the top cover full welding equipment is as follows: in, This is the power instability penalty coefficient; For the first The equipment is in The actual output power at any given time; For the first The set power value of the device.

2. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 1, characterized in that, The process of inputting the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model yields a global context vector, including: The state vector of each device is input into the corresponding agent, and a query vector, key vector, and value vector are generated through neural network mapping. For the current agent, the similarity between the query vector and the key vectors output by the agents of all devices is calculated, and then normalized by a function to obtain a set of attention weights; The attention weights are used as coefficients to perform a weighted summation of the value vectors of all devices to generate a global context vector. The global context vector includes the health of all devices, the work-in-progress loading of all buffers, the distribution of maintenance activities of all devices, and the distribution of downtime of all devices.

3. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 1, characterized in that, The step of inputting the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions includes: The current agent's own state vector and the global context vector are concatenated and input into the current agent to obtain the action probability distribution; A greedy strategy is adopted to select the final action from the action probability distribution as the local predictive action decision.

4. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 1, characterized in that, The random failure probability of the device is: in, Indicates the first The probability of random failure of a device; This represents the highest probability of random failure throughout the entire lifespan of the equipment. Indicates the first The equipment is in Health status at all times; For both positive and negative electrode slitting equipment, the failure probability shape parameter is... For electrode pre-welding equipment and connecting piece welding equipment, the failure probability shape parameter For cover plate welding equipment and top cover full welding equipment, the failure probability shape parameter .

5. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 1, characterized in that, It also includes: training the deep reinforcement learning model, which further includes a centralized critic network; Training methods for deep reinforcement learning models include: Randomly sample from the stored device operation data to obtain the state data of all devices and the corresponding local predictive action decisions output by the intelligent agent; The optimal action strategy is obtained by inputting the status data of all devices into a centralized critic network with the goal of maximizing the global reward. Based on the optimal action policy and local predictive action decisions, the parameters of the local decision network in the agent and the parameters of the attention network in the global coordinator are updated by minimizing the loss function composed of temporal difference errors and using gradient descent.

6. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 5, characterized in that, The global reward is in, Indicates revenue per unit of product; This indicates the output of products produced by the lithium battery assembly line. Indicates the first Maintenance costs incurred from corrective maintenance of the equipment; Indicates the first Maintenance costs incurred from preventative maintenance of equipment; Indicator variables that are either 0 or 1. Indicates the first The equipment is in Whether to perform corrective maintenance at all times Indicates the first The equipment is in Whether preventative maintenance is performed at all times; This represents the production loss cost incurred per unit of equipment downtime; Indicates the first The operating status of the equipment Indicates that the first device is in It is constantly in a shutdown state; This indicates that the second device is in It is constantly in a shutdown state.

7. The predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in claim 6, characterized in that, The optimal action strategy is: in, Represents the strategy set for a lithium battery assembly line; Indicating in strategy Next deadline Total revenue generated by the production line.

8. A predictive maintenance system for lithium battery assembly lines based on deep reinforcement learning, characterized in that, Implementing the predictive maintenance method for lithium battery assembly lines based on deep reinforcement learning as described in any one of claims 1-7, comprising: The state vector construction module is used to construct the state vector of the equipment based on the current health of the equipment on the lithium battery assembly line, the work-in-process loading degree of the upstream buffer, the work-in-process loading degree of the downstream buffer, the operating status, and the remaining duration of the maintenance task. The lithium battery assembly line includes multiple equipment and multiple buffers. The coordination module is used to input the state vectors of all devices into the global coordinator of the trained deep reinforcement learning model to obtain the global context vector. The deep reinforcement learning model includes multiple agents equipped on the devices and a global coordinator. Each device on the lithium battery assembly line is equipped with one agent. The prediction module is used to input the state vector and global context vector of each device into the corresponding agent of the device to obtain local predictive action decisions. The local predictive action decisions include 0 and 1, where 0 indicates no maintenance action and 1 indicates predictive maintenance action. The conversion and execution module is used to convert local predictive action decisions into equipment maintenance actions for execution.

Citation Information

Patent Citations

  • Hybrid sensing-based intelligent positioning and diagnosis method for optical fiber composite fault of power distribution network

    CN119881542A