A method for continuous short-circuit detection of mass-produced lithium batteries

By introducing three detection states and a reinforcement learning model into batch lithium battery testing, the detection state and duration are adaptively adjusted. Combined with a short-circuit identification model and accuracy updates, the problems of detection accuracy and stability in batch testing are solved, and the ability to identify short circuits is improved.

CN121703678BActive Publication Date: 2026-04-21SHENZHEN ZHIJIANENG AUTOMATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN ZHIJIANENG AUTOMATION CO LTD
Filing Date
2026-02-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to cover different short-circuit paths, identify developing short circuits, and maintain consistent detection accuracy in batch lithium battery testing under limited throughput conditions. In particular, they are prone to misjudgment under the influence of switching constraints and sampling time.

Method used

A switchable detection space is constructed using three detection states (positive and negative sampling, positive shell sampling, and floating state). The detection state and duration are adaptively determined by a reinforcement learning model. Multiple detection data are judged through a short-circuit identification model, and the model is updated using the detection accuracy.

Benefits of technology

It improves the accuracy and stability of batch lithium battery testing, can identify internal short circuits and casing short circuits and their development stages, reduces the false judgment rate, and adapts to the testing needs of different short circuit morphologies and stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121703678B_ABST
    Figure CN121703678B_ABST
Patent Text Reader

Abstract

This invention discloses a method for continuous short-circuit detection of mass-produced lithium batteries, relating to the field of battery short-circuit detection. The method includes: for the current lithium battery among multiple lithium batteries to be continuously tested, each test uses one of three detection states: positive and negative electrode sampling state, positive casing sampling state, and floating state, and collects the detected voltage or current data to form corresponding detection data; after completing one test, based on the current lithium battery's detection data, a reinforcement learning model determines the detection state and corresponding detection duration for the next test; the detection process is repeated multiple times to form an observation dataset, which is then input into a short-circuit identification model to output a short-circuit determination and a short-circuit type determination; the lithium battery is replaced and the detection process is repeated. This invention solves the problems of difficulty in adaptively configuring the detection state and detection duration, and unstable short-circuit type identification in mass testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of battery short-circuit detection, and more specifically, to a method for continuous short-circuit detection of mass-produced lithium batteries. Background Technology

[0002] Before production, sorting, assembly, and shipping, lithium batteries typically undergo short-circuit related safety testing to reduce the risk of thermal runaway during subsequent storage, transportation, or use. Short circuits can manifest not only as abnormal conduction between the positive and negative electrodes but also as abnormal conduction paths between the electrodes and the battery casing. In the early stages, short circuits may also exhibit intermittent, conditionally triggered abnormal conduction characteristics, leading to inconsistent electrical responses of the same cell at different testing times or under different connection states.

[0003] Among existing solutions, the technology with publication number CN110187225A, entitled "A Method and System for Detecting Abnormal Voltage and Current in Internal Short Circuits of Lithium Batteries," mainly targets the charging and discharging processes such as formation and capacity testing. It detects anomalies by synchronously collecting voltage and current data and using rules such as threshold values, decreasing trends, and decreasing slopes. This triggers warnings or shutdowns when anomalies occur. Its core lies in the rule-based identification of voltage and current anomalies during the charging and discharging process. The technology with publication number CN112946522A, entitled "Online Monitoring Method for Internal Short Circuit Faults in Battery Energy Storage Systems Caused by Low Temperature Conditions," focuses on low-temperature energy storage system scenarios. It simulates the transient voltage response by establishing an equivalent model of internal short circuits and extracts curve features by combining inter-cell correlation coefficients, moving window filtering, and active noise reduction to achieve online monitoring and safety status assessment of internal short circuits at low temperatures. While the two approaches mentioned above can monitor or warn of internal short-circuit anomalies under specific operating conditions, they typically assume that the detection connection relationships and sampling processes are relatively fixed. They focus more on signal trend identification under charging / discharging conditions or low-temperature operating conditions, without directly addressing the switching constraints and cycle time constraints commonly found in continuous batch testing on production lines. They also fail to provide adaptive switching order and sampling duration determination mechanisms to address the issue that rapid switching between different detection connection states can alter measurement results. Furthermore, these two comparative documents do not provide suppression and discrimination paths strongly correlated with batch polling for shell-related false anomalies caused by shell-side potential residues or drift introduced by batch clamping and repeated measurements. Therefore, it remains difficult to simultaneously cover different short-circuit paths, identify developing short circuits, and maintain stable and consistent detection accuracy under throughput-limited conditions.

[0004] In batch continuous testing, to meet cycle time requirements, testing devices often need to rapidly switch between different measurement connections using relays and other switching devices, acquiring voltage or current signals within a short time window. The testing results are easily affected by the switching sequence, switching interval, and sampling duration. Simultaneously, batch clamping and repeated measurements may introduce residual potential or drift on the casing side, causing casing-related measurements to exhibit abnormal behavior not intrinsic to the battery cell in the initial stage, further increasing the probability of misjudgment. For production lines, it is desirable to cover different short-circuit paths and different development stages while maintaining stable accuracy and consistency under limited throughput conditions. Therefore, a continuous testing method is needed that can adaptively determine the testing state and duration in batch scenarios and reliably classify short-circuit states. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for continuous short-circuit detection of mass-produced lithium batteries, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for continuous short-circuit detection of mass-produced lithium batteries includes the following steps:

[0008] S1: For the current lithium battery among multiple lithium batteries to be continuously tested, each time the current lithium battery is tested using one of the following detection states: positive and negative electrode sampling state, positive shell sampling state, and floating state, and the detected voltage data or current data is collected to form corresponding detection data.

[0009] S2: After completing one detection, based on the current lithium battery detection data, the detection state and corresponding detection duration for the next detection are determined through a reinforcement learning model, and the detection state and corresponding detection duration are switched to in the next detection.

[0010] S3: Repeat the detection process of S1 and S2 multiple times to form the current observation dataset of the lithium battery;

[0011] S4: Input the observation dataset into the short circuit identification model, and the short circuit identification model outputs a determination of whether the current lithium battery has a short circuit, and a determination of the corresponding short circuit type when a short circuit exists;

[0012] S5: Replace the current lithium battery and repeat the detection process of S1~S4;

[0013] During the training of the reinforcement learning model, the detection accuracy obtained from multiple lithium batteries is used as a feedback signal to update the reinforcement learning model.

[0014] Preferably, the positive and negative electrode sampling state is the voltage or current detection state performed after connecting a resistor between the positive and negative electrodes of the current lithium battery.

[0015] The positive casing sampling state refers to the voltage or current detection state performed after connecting a resistor between the positive electrode of the current lithium battery and the battery casing.

[0016] The floating state is the detection state in contrast to the current lithium battery disconnected state.

[0017] Preferably, the short circuit type includes an internal short circuit state and a shell short circuit state. Each state is further divided into a completed short circuit state and a developing short circuit state. The completed short circuit state is a stable abnormal conduction state, and the developing short circuit state is a state with intermittent or conditionally triggered abnormal conduction characteristics.

[0018] Preferably, the training process of the short circuit identification model includes: acquiring a historical observation dataset, using the re-inspection results or downstream process confirmation results corresponding to the historical observation dataset as label data, wherein the label data includes a determination of whether a short circuit exists and a determination of the type of short circuit; and using the observation dataset and the label data as samples for supervised learning training.

[0019] Preferably, obtaining the detection accuracy includes: comparing the short-circuit determination result output by the short-circuit identification model with the re-inspection result of the corresponding lithium battery or the confirmation result of the downstream process; when the short-circuit determination result is consistent with the re-inspection result or the confirmation result of the downstream process, it is recorded as a correct detection, and when they are inconsistent, it is recorded as an incorrect detection; after a preset number of lithium batteries have been tested, the detection accuracy is determined based on the ratio between the number of correct detections and the total number of correct and incorrect detections.

[0020] Preferably, the reinforcement learning model is a reinforcement learning model based on a value function or a policy function. Its state input includes the current detection data of the lithium battery, the historical detection state sequence, and the corresponding detection duration information. Its action output is used to indicate the detection state and the corresponding detection duration for the next detection.

[0021] Preferably, the reinforcement learning model employs a deep Q-network model, a policy gradient model, or a model based on the Actor-Critic architecture.

[0022] Preferably, the detection accuracy is also used to update the short-circuit identification model, and a phased update mechanism is adopted. The phased update mechanism includes: freezing the parameters of the short-circuit identification model in the first stage, and using the detection accuracy as a feedback signal to update the policy of the reinforcement learning model.

[0023] In the second stage, the parameters of the reinforcement learning model are frozen, and the short-circuit identification model is trained and updated based on the detection accuracy.

[0024] Preferably, the detection accuracy is updated using a dual-timescale update mechanism for the short-circuit identification model and the reinforcement learning model. The dual-timescale update mechanism includes: the short-circuit identification model performing batch parameter updates according to a first preset training period, and the reinforcement learning model performing policy updates according to a second preset update period different from the first preset training period.

[0025] Preferably, the detection device used in the method includes: a detection interface for establishing an electrical connection with the positive electrode, negative electrode, and battery casing of the lithium battery; a relay switch module electrically connected to the detection interface for switching different detection states; a sampling module for collecting electrical signals of the lithium battery under different detection states; a control module for controlling the relay switch module to switch and receiving data collected by the sampling module; a communication module electrically connected to the control module; and a display module connected to the communication module through a communication interface for displaying detection results or abnormal alarms.

[0026] The advantage of this invention over the prior art is that it constructs a switchable detection space around the positive and negative electrode sampling state, the positive shell sampling state, and the floating state during batch continuous detection. After each detection, a reinforcement learning model is used to adaptively determine the detection state and corresponding detection duration for the next detection. This allows the detection process to dynamically adjust the measurement sequence and time allocation according to the response characteristics of the current battery cell, avoiding the problem of insufficient coverage or redundant sampling in different battery cells and different short-circuit stages due to fixed strategies.

[0027] This invention generates an observation dataset through repeated detections and inputs it into a short-circuit identification model to output whether a short circuit occurs and the type of short circuit. It can simultaneously identify internal short circuits and short circuits to the casing, and further distinguish between completed stable abnormal conduction states and developing intermittent or conditionally triggered abnormal conduction states, thereby improving the ability to detect early-stage risky cells. The floating state provides open-circuit measurement conditions, helping to reduce the interference of potential residues on the casing side during batch testing, making casing-related signals closer to the intrinsic response of the cell, and improving the stability of casing-to-casing short-circuit identification. The short-circuit identification model is trained using historical observation datasets and re-inspection or downstream confirmation results through supervised learning. This allows it to map multi-state, multi-time-series detection data into short-circuit judgments and type outputs, reducing errors caused by relying on a single threshold rule. Detection accuracy is obtained by statistically analyzing the consistency between the identification output and the re-inspection or downstream confirmation results, serving as a feedback signal to update the reinforcement learning model, aligning the strategy optimization objective with the actual quality closed loop of the production line. Furthermore, the detection accuracy can also be used to update the short-circuit identification model. A phased freeze update or dual-timescale update mechanism can decouple the strategy layer from the identification layer, reducing the risk of variable coupling caused by simultaneous updates of both models and improving training convergence and online operational controllability. The detection device, through modular configuration of detection interfaces, relay switches, sampling, control, and communication displays, facilitates rapid multi-state switching, data acquisition, and alarm output in production line environments, balancing detection cycle time with engineering feasibility. Attached Figure Description

[0028] Figure 1 This is the overall flowchart of the present invention;

[0029] Figure 2 This is a schematic diagram of the device of the present invention;

[0030] Figure 3 This is a schematic diagram of the three switching states of the present invention. Detailed Implementation

[0031] The specific embodiments of the present invention will now be described with reference to the accompanying drawings.

[0032] This invention relates to continuous short-circuit detection of lithium batteries in production line or laboratory settings. Short circuits in lithium batteries can occur between the positive and negative electrodes within the cell, or they can manifest as abnormal conductive paths between the electrodes and the battery casing. During batch testing, frequent battery changes and tight testing cycles require the testing device to sample rapidly under different connection relationships. If a fixed testing sequence and sampling duration are used, two types of problems can easily arise. One is the difference in sensitivity to different short-circuit types; for example, internal short circuits are usually more directly reflected in the terminal voltage, while casing short circuits are more dependent on the casing-related sampling connection relationships. The other is the residual effect brought about by batch testing. In particular, casing-related measurements may drift within a short period due to residual potential from previous batteries, fixture contact interfaces, or measurement circuits, making it easier to misjudge or miss detections under certain sequences.

[0033] To address this, the present invention dynamically switches between three detection states and introduces a reinforcement learning model to adaptively determine the next detection state and duration. Simultaneously, a short-circuit identification model is used to determine short-circuit events and types on the observation dataset formed from multiple detections, thereby improving detection accuracy and stability under throughput constraints. The overall process is as follows: Figure 1 As shown.

[0034] In one embodiment, the detection device used in this invention is as follows: Figure 2 As shown, the system includes a detection interface, a relay switch module, a sampling module (the embodiment shown is a voltage sampling module), a control module (MCU module), a communication module, and a display module. The detection interface is used to establish electrical connections with the positive and negative terminals of the lithium battery and the battery casing, respectively. The detection interface can be implemented using a spring-loaded pin, a clamp contact, or a flexible conductive sheet. Considering that contact consistency has a significant impact on measurement stability during batch testing, the detection interface preferably adopts a spring-loaded pin structure with elastic stroke, and a guide limit is set on the clamp to reduce the positional error of each clamping.

[0035] The relay switch module is used to switch between different detection states and can be implemented using electromagnetic relays, reed relays, or solid-state relays. Electromagnetic relays are low-cost and have high durability, reed relays have better contact stability in small signal measurements, and solid-state relays have higher switching speeds.

[0036] The sampling module is used to acquire voltage or current signals. The sampling module may include an analog-to-digital converter and a front-end sampling circuit. Voltage sampling may use a differential sampling structure to suppress common-mode interference, while current sampling may be achieved using a shunt resistor or Hall effect sampling.

[0037] The control module controls the switching of the relay switch module and receives data from the sampling module. It can be implemented using a microcontroller or embedded processor. The control module also runs reinforcement learning model inference and short-circuit identification model inference, or uploads data to a host computer or server for inference via the communication module. The communication module can be implemented using serial port, Ethernet, or wireless methods. The display module displays detection results or abnormal alarms; it can be a local display screen or a host computer interface.

[0038] In one embodiment, the three detection states are as follows: Figure 3 The diagram shows the sampling states for positive and negative electrodes, positive casing sampling, and floating state, respectively. The positive and negative electrode sampling state involves connecting a resistor between the positive and negative electrodes of the current lithium battery to detect voltage or current. This resistor is used to construct a controllable measurement circuit and current-limiting path, allowing the sampling to reflect both the terminal voltage state and the degree of abnormal conduction through the current response. The resistor value can range from 10Ω to 200kΩ, selected according to battery specifications, target sensitivity, and safety limitations; a common range is 100Ω to 20kΩ. The positive casing sampling state involves connecting a resistor between the positive electrode of the current lithium battery and the battery casing to detect voltage or current. This resistor value can range from 10Ω to 200kΩ, with a common range being 500Ω to 50kΩ.

[0039] The purpose of the positive casing sampling state is to incorporate the casing into the measurement circuit, thereby enabling more sensitive observation of any abnormal conduction between the electrodes and the casing. The floating state is the detection state relative to the current open-circuit state of the lithium battery; that is, the relay switch module disconnects the measurement circuit from the lithium battery electrodes and casing or switches to a high-resistance input, retaining only the high-resistance measurement channel at the sampling front end or completely disconnecting it. The introduction of the floating state not only provides control data but also suppresses interference from residual casing potentials in batch testing on subsequent positive casing sampling, allowing the casing potential to return to a drift trajectory closer to its intrinsic value, thus reducing the probability of misjudgment.

[0040] The batch continuous testing of this invention is performed according to the following process: The control module acquires multiple lithium batteries to be continuously tested and clamps or connects them to the testing interface in sequence. The currently clamped lithium battery is recorded as the current lithium battery. For the current lithium battery, one of three testing states is selected for each test. The control relay switch module switches to the testing state, and the sampling module collects voltage data or current data to form testing data within the corresponding testing duration. Voltage data may include terminal voltage sampling value, positive casing sampling voltage value, or their timing sequence. Current data may be the instantaneous value or timing sequence of the loop current. To balance cycle time and stability, the single testing duration can be from 20ms to 500ms, with a commonly used range of 50ms to 200ms. Too short a testing duration may result in the sampling not yet being stable, while too long a duration will reduce throughput. Since the loop impedance, contact conditions, and electrical response time constant are different under different testing states, the testing duration of the positive and negative electrode sampling states and the positive casing sampling states may be different. The testing duration of the floating state is usually shorter, used to provide a residual attenuation observation window or a switching buffer window.

[0041] After completing a detection, the control module combines the detection data with historical detection state sequences and historical detection duration information to form a state input, which is then input into the reinforcement learning model. The reinforcement learning model outputs the detection state and corresponding detection duration for the next detection, and switches to the corresponding detection state for the corresponding duration in the next detection. The purpose of this design is to enable the detection strategy to adaptively adjust according to the current response characteristics of the lithium battery, avoiding mismatches that occur under different short-circuit types or different development stages due to a fixed sequence.

[0042] For example, if the current lithium battery has a significant internal short circuit, the terminal voltage in the positive and negative electrode sampling states will be significantly lower than the normal range. The reinforcement learning model may tend to shorten the duration of subsequent floating or positive casing sampling to quickly confirm and improve throughput. If there is a short circuit to the casing, the positive and negative electrode sampling may show abnormal terminal voltages, while the positive casing sampling will show more significant casing conduction characteristics. The reinforcement learning model may choose the positive casing sampling state earlier and extend its detection duration to improve the stability of casing evidence. If there is a developing short circuit, a single sampling may only show anomalies under specific connection relationships or specific durations. The reinforcement learning model will gradually explore and strengthen the combinations of states and durations that can trigger anomalies through multiple rounds of detection to improve the capture probability. The reason why different detection orders affect the results is that batch detection has residual and switching transients, especially casing-related measurements are more sensitive to initial potential and contact interfaces. When the detection order is positive casing sampling first and then positive and negative electrode sampling, the casing may carry the residual potential of the previous battery, causing transient anomalies in the positive casing sampling. If a conclusion is immediately drawn based on this anomaly, it may lead to misjudgment. Inserting the device into a floating state first, or performing positive and negative sampling before entering positive shell sampling, allows the shell potential to exhibit a decay trajectory within the floating window, or allows the terminal voltage to provide a stable reference first, thereby reducing false positives. The detection order and detection duration directly determine the distribution and signal-to-noise ratio of the observed data, making adaptive selection of the order and duration necessary for the reinforcement learning model.

[0043] In one embodiment, the reinforcement learning model is a value function- or policy function-based model, whose action output includes detection state selection and detection duration selection. The action space for detection state selection consists of three discrete actions, corresponding to positive and negative electrode sampling states, positive shell sampling states, and floating states. Detection duration selection can be discretized, dividing the duration into several levels, for example, dividing 20ms to 500ms into 8 to 32 levels, or it can use continuous action output, with the policy network directly outputting the duration and then pruning it to meet range constraints. The state input includes the current lithium battery detection data, historical detection state sequences, and corresponding detection duration information. Detection data can be input in a time-series manner, with the voltage or current sampling sequence as a one-dimensional sequence input to the network, or features can be extracted first and then input, including features such as mean, extreme values, slope, variance, and stable plateau duration. The historical detection state sequences and duration information are used to describe the recently occurred switching and sampling conditions, facilitating the model's determination of whether certain anomalies are caused by transient or residual events. Reinforcement learning models can be implemented using deep Q-networks, suitable for discrete action scenarios. The network outputs the Q-value of each action and selects the action corresponding to the largest Q-value. Reinforcement learning models can also be implemented using policy gradients or an Actor-Critic architecture. The Actor outputs the detection state and detection duration, while the Critic evaluates the reward. The reward signal is derived from the detection accuracy, the calculation method of which is given later.

[0044] To avoid frequent switching that could reduce relay lifespan or increase transient interference, the control module can add constraints during action selection. For example, it can limit the number of detections per lithium battery to 2 to 10 times, limit the number of relay switching to 2 to 12 times, and limit the total detection time to a preset clock cycle threshold. This threshold can be set to 0.2 to 5 seconds, with a commonly used range of 0.5 to 2 seconds. If the reinforcement learning model outputs actions or durations exceeding these constraints, the control module can prune or revert to a preset safety strategy.

[0045] After repeated testing, an observation dataset is generated for the current lithium battery. This dataset can be stored as sequential samples according to the testing order, or grouped by testing state and accompanied by timestamps and testing durations. This dataset is then input into a short-circuit identification model, which outputs a determination of whether a short circuit exists and the type of short circuit. Short circuit types include internal short circuit states and shell-side short circuit states. Each state is further divided into completed and developing short circuit states. Completed short circuit states correspond to stable abnormal conduction, characterized by the repeated appearance of abnormal features in multiple tests without depending on specific triggering conditions. Developing short circuit states correspond to intermittent or conditionally triggered abnormal conduction, characterized by the abnormality appearing only in certain testing states or combinations of testing durations, or by an increasing trend in abnormal features as the testing sequence progresses. The output of the short-circuit identification model can be a four-category output, corresponding to completed internal short circuit, developing internal short circuit, completed shell-side short circuit, and developing shell-side short circuit. Alternatively, a two-stage output can be used: the first stage outputs whether a short circuit exists, and the second stage outputs the type and development stage under short-circuit conditions.

[0046] In one embodiment, the short-circuit identification model is trained using supervised learning. Training data comes from historical observation datasets, with the corresponding re-inspection results or downstream process confirmation results serving as label data. The label data includes a determination of whether a short circuit exists and the type of short circuit. Re-inspection results can come from higher-precision offline testing equipment, and downstream process confirmation results can come from determinations made during processes such as capacitance testing, internal resistance testing, X-ray examination, or disassembly analysis.

[0047] Two methods can be used to construct training samples from the observation dataset and label data. The first method directly uses time-series data from multiple detections, concatenating the voltage or current sequences from each detection over time and adding encoded information about the detection state and duration. This encoding can employ one-hot encoding or numerical encoding, enabling the model to understand the response differences of the same battery cell under different states. The second method extracts features from each detection data point to form feature vectors, then concatenates or pools these feature vectors sequentially to obtain a fixed-dimensional input. The short-circuit identification model can use gradient boosting trees to classify the feature vectors or lightweight neural networks to classify the time-series sequences. The neural network structure can include one-dimensional convolutional layers or gated recurrent units to extract time-series patterns, followed by fully connected layers to output the classification results. The training process uses cross-entropy loss or weighted cross-entropy loss, with weights used to balance the imbalance of samples from different short-circuit types; the weight range can be 1 to 20. The batch size for model training can range from 32 to 2048, the learning rate can range from 1e-5 to 1e-2, and the number of training epochs can range from 5 to 200, depending on the data size and convergence performance. After training, the model parameters are obtained and deployed on the control module or host computer.

[0048] After completing the current lithium battery inspection, replace the current lithium battery and repeat the above inspection process. To enable the reinforcement learning model to be progressively optimized under real production line feedback, the detection accuracy is used to update the reinforcement learning model. The detection accuracy is obtained by comparing the short-circuit judgment result output by the short-circuit identification model with the re-inspection result of the corresponding lithium battery or the confirmation result of the downstream process. If they match, it is recorded as a correct detection; if they do not match, it is recorded as an incorrect detection. After a preset number of lithium batteries have been inspected, the detection accuracy is determined by the ratio of the number of correct detections to the total number of correct and incorrect detections. The preset number can be 10 to 5000, with a commonly used range of 50 to 500. Using window statistics can reduce the impact of individual sample noise on strategy updates and match the production line cycle time. The reinforcement learning model can construct a reward signal based on this accuracy, for example, setting the reward as the accuracy itself or as the accuracy improvement of adjacent windows, and taking into account the constraints of the number of detections and the total detection time in the strategy update, so that the strategy can improve accuracy while ensuring throughput.

[0049] In another embodiment, in batch continuous detection scenarios, the training feedback of the reinforcement learning model can, in addition to detection accuracy, further incorporate detection efficiency, which is directly related to the production line cycle time, as part of the reward signal. This allows the strategy to improve accuracy while considering throughput and resource consumption. Detection efficiency can be characterized by a weighted average of one or more data points, including the total detection time for a single lithium battery, the number of detections required to complete one detection, the number of relay switching operations, and the number of batteries detected per unit time. This efficiency, along with detection accuracy, constitutes a comprehensive reward, enabling the reinforcement learning model to reduce unnecessary repeated detections and excessively long sampling when abnormal features are clear, and to reasonably increase the number of detections or extend the detection time of critical states when abnormal features are unstable or suspected of developing short circuits. This improves the overall detection cycle time and production line operating efficiency while keeping the false positive rate under control. When both detection accuracy and detection efficiency reward signals are introduced simultaneously, to avoid conflicts in policy learning objectives or instability in the training process, a phased or weighted fusion approach is typically used to apply both types of rewards. A common approach is to pre-train the reinforcement learning model using detection accuracy as the primary reward. This allows the model to learn to select more reliable detection states and durations under different short-circuit conditions. Once the accuracy reaches a preset stable level, a detection efficiency reward is introduced while maintaining the accuracy reward as the primary factor. By adjusting the weights, the model gradually reduces the total detection time, number of detections, or number of relay switching without significantly sacrificing accuracy. Another common approach is to apply linear weighting or hierarchical constraints to the two types of rewards. For example, detection accuracy can be used as a hard constraint or threshold condition, allowing the detection efficiency reward to take effect only when the accuracy reaches a preset threshold. Alternatively, when accuracy is insufficient, only policy parameters related to accuracy can be updated, while policy parameters related to efficiency optimization can be frozen, thus preventing the policy from introducing systematic misjudgments in pursuit of speed. An alternating update mechanism can also be used. In several training cycles, the efficiency weight is fixed to optimize only accuracy, and in subsequent training cycles, the accuracy weight is fixed to optimize only efficiency. This decouples the two objectives over time, reducing oscillations caused by mutual interference. The above approach enables both types of rewards to work together on policy updates while keeping the training process controllable and better meeting the engineering requirements of batch testing, which requires both accuracy and rhythm.

[0050] In a further embodiment, the detection accuracy is also used to update the short-circuit identification model. To avoid interference between policy changes and identification model changes, a phased update mechanism is adopted. In the first phase, the short-circuit identification model parameters are frozen, and the reinforcement learning model is updated only using the detection accuracy as feedback, allowing the policy to converge in a stable identifier environment. In the second phase, the reinforcement learning model parameters are frozen, making the sampling policy relatively stable. Then, the short-circuit identification model is trained and updated based on the detection accuracy, allowing the identification model to adapt to the data distribution collected by the current policy. The purpose of the phased approach is to control variables and avoid training oscillations caused by simultaneous changes in both models. Freezing can be achieved by stopping gradient updates or fixing the model version. The phase switching condition can be set to the accuracy changing less than a threshold within N consecutive windows, where N can be 3 to 30, and the threshold can be 0.001 to 0.05.

[0051] In another embodiment, a dual-timescale update mechanism is used to update the short-circuit identification model and the reinforcement learning model. The short-circuit identification model updates its parameters in batches according to a first preset training period, which can be 1 to 30 days or counted based on the cumulative sample size, for example, updating once every 1,000 to 500,000 samples. The reinforcement learning model updates its policy according to a second preset update period, which is different from the first preset training period. The second preset update period can be 1 minute to 2 hours or counted based on the cumulative sample size, for example, updating once every 10 to 5,000 samples. The slower update of the short-circuit identification model helps maintain the stability of the discrimination boundary, while the faster update of the reinforcement learning model helps to adapt to production line fluctuations and residual changes in a timely manner. Through the difference in timescales, policy optimization and identification optimization are decoupled in engineering, improving system controllability.

[0052] To further illustrate the relationship between different short-circuit conditions and different detection orders, several policy examples can be provided to help understand the possible detection orders that the reinforcement learning model may choose.

[0053] If the battery cell has completed an internal short circuit, the terminal voltage will remain below the normal range under the positive and negative sampling state, and this anomaly is relatively stable under different detection durations. The reinforcement learning model may choose to directly enter the short circuit identification after short-term positive and negative sampling, thereby reducing the dependence on positive shell sampling and improving the cycle time.

[0054] If the cell has completed the short circuit to the casing, the terminal voltage may also be abnormal. However, a more significant and repeatable casing conduction signal will appear under the positive casing sampling state. The reinforcement learning model may insert positive casing sampling as soon as possible after the first positive and negative sampling and extend its detection time to improve the stability of the casing evidence. Then, the terminal voltage consistency is verified by the second positive and negative sampling.

[0055] If the battery cell is in a short-circuit state during development, anomalies may only appear within a specific time window after a switchover, such as in the initial stage of positive casing sampling after floating or at the end of a longer positive and negative electrode sampling period. The reinforcement learning model may gradually increase the detection time from a short interval through multiple trials, or insert floating between positive casing sampling and positive and negative electrode sampling to reduce residual interference, thus making it easier to capture intermittent anomalies. If there is significant casing residue in batch detection, directly entering positive casing sampling may cause drift and false anomalies in the initial stage. The reinforcement learning model will tend to choose the floating state first or perform positive and negative electrode sampling first, waiting for the casing potential to decay before entering positive casing sampling, thereby improving accuracy.

[0056] The above cases demonstrate that changes in detection order and detection duration can alter the stability and separability of observed data. Reinforcement learning models, during training with detection accuracy as feedback, will gradually favor sequences that are more conducive to distinguishing between real short circuits and residual interference.

[0057] As can be seen from the above embodiments, the present invention provides three switchable detection states and a batch detection interface at the device level, and adaptively determines the detection state and detection duration through a reinforcement learning model at the method level. It also uses a short-circuit identification model to determine short circuits and their types from multiple detection data. At the same time, it introduces a closed-loop detection accuracy mechanism and a phased or dual-time-scale update mechanism, which enables throughput, stability and interpretability to be balanced in batch continuous detection scenarios, and improves the ability to identify internal short circuits, shell short circuits and their development stages.

[0058] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for continuous short-circuit detection of mass-produced lithium batteries, characterized in that, Includes the following steps: S1: For the current lithium battery among multiple lithium batteries to be continuously tested, each time the current lithium battery is tested using one of the following detection states: positive and negative electrode sampling state, positive shell sampling state, and floating state, and the detected voltage data or current data is collected to form corresponding detection data. S2: After completing one detection, based on the current lithium battery detection data, the detection state and corresponding detection duration for the next detection are determined through a reinforcement learning model, and the detection state and corresponding detection duration are switched to in the next detection. S3: Repeat the detection process of S1 and S2 multiple times to form the current observation dataset of the lithium battery; S4: Input the observation dataset into the short circuit identification model, and the short circuit identification model outputs a determination of whether the current lithium battery has a short circuit, and a determination of the corresponding short circuit type when a short circuit exists; the short circuit type includes internal short circuit state and shell short circuit state; The training process of the short circuit identification model includes: acquiring a historical observation dataset, using the re-inspection results or downstream process confirmation results corresponding to the historical observation dataset as label data, the label data including the determination of whether there is a short circuit and the determination of the short circuit type; and using the observation dataset and the label data as samples for supervised learning training. S5: Replace the current lithium battery and repeat the detection process of S1~S4; During the training of the reinforcement learning model, the detection accuracy obtained from multiple lithium batteries is used as a feedback signal to update the reinforcement learning model.

2. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, The positive and negative electrode sampling status refers to the voltage or current detection status after connecting a resistor between the positive and negative electrodes of the current lithium battery. The positive casing sampling state refers to the voltage or current detection state performed after connecting a resistor between the positive electrode of the current lithium battery and the battery casing. The floating state is the detection state in contrast to the current lithium battery disconnected state.

3. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, Each internal short circuit state and shell short circuit state is further divided into completed short circuit state and developing short circuit state. The completed short circuit state is a stable abnormal conduction state, and the developing short circuit state is a state with intermittent or conditionally triggered abnormal conduction characteristics.

4. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, The acquisition of the detection accuracy includes: comparing the short-circuit determination result output by the short-circuit identification model with the re-inspection result of the corresponding lithium battery or the confirmation result of the downstream process; when the short-circuit determination result is consistent with the re-inspection result or the confirmation result of the downstream process, it is recorded as a correct detection, and when they are inconsistent, it is recorded as an incorrect detection; after a preset number of lithium batteries have been tested, the detection accuracy is determined based on the ratio between the number of correct detections and the total number of correct and incorrect detections.

5. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1 or 4, characterized in that, The reinforcement learning model is a value function or policy function-based reinforcement learning model. Its state input includes the current lithium battery detection data, historical detection state sequence and corresponding detection duration information. Its action output is used to indicate the detection state and corresponding detection duration for the next detection.

6. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 5, characterized in that, The reinforcement learning model employs a deep Q-network model, a policy gradient model, or a model based on the Actor-Critic architecture.

7. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, The detection accuracy is also used to update the short-circuit identification model, and a phased update mechanism is adopted. The phased update mechanism includes: freezing the parameters of the short-circuit identification model in the first stage, and using the detection accuracy as a feedback signal to update the policy of the reinforcement learning model. In the second stage, the parameters of the reinforcement learning model are frozen, and the short-circuit identification model is trained and updated based on the detection accuracy.

8. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, The detection accuracy is updated using a dual-timescale update mechanism for the short-circuit identification model and the reinforcement learning model. The dual-timescale update mechanism includes: the short-circuit identification model performing batch parameter updates according to a first preset training period, and the reinforcement learning model performing policy updates according to a second preset update period different from the first preset training period.

9. The short-circuit continuous detection method for mass-produced lithium batteries according to claim 1, characterized in that, The detection device used in the method includes: a detection interface for establishing an electrical connection with the positive electrode, negative electrode, and battery casing of the lithium battery; a relay switch module electrically connected to the detection interface for switching different detection states; a sampling module for collecting electrical signals of the lithium battery under different detection states; a control module for controlling the relay switch module to switch and receiving data collected by the sampling module; a communication module electrically connected to the control module; and a display module connected to the communication module through a communication interface for displaying detection results or abnormal alarms.

Citation Information

Patent Citations

  • Method and system for detecting abnormality of short-circuit voltage and current in lithium battery

    CN110187225A

  • Online monitoring method for short-circuit fault in battery energy storage system caused by low-temperature working condition

    CN112946522A

  • Short circuit identification method and system for vehicle lithium battery, electronic equipment and storage medium

    CN117890797A

  • Method and apparatus with battery short detection

    US20230400518A1