Soft package battery tab cutting and welding process parameter optimization method and system

By collecting and processing cutting and welding data, and combining transfer learning and reinforcement learning to optimize process parameters, the shortcomings of traditional methods that rely on experience are solved, and precise control and efficient production of tab cutting and welding for pouch batteries are achieved.

CN121035544APending Publication Date: 2025-11-28CHANGZHOU ZHIKE AUTOMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511083229.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional pouch cell battery tab cutting and welding process parameters adjustment rely on experience and lack scientific monitoring and optimization methods, making it difficult to meet the requirements of high energy density and high safety.

Method used

Data from the cutting and welding processes are collected using a force gauge and displacement sensor. Signal processing is performed to extract feature parameters. The welding status is monitored using a thermal imager and a current acquisition device. Transfer learning and reinforcement learning methods are then employed to optimize the process parameters.

Benefits of technology

It achieves precise characterization of the cutting process, reduces burrs and deformation defects, improves the quality and consistency of the electrode tabs, ensures the strength and stability of the welded joints, reduces the time cost of optimizing process parameters, and improves production efficiency and product quality stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121035544A_ABST
    Figure CN121035544A_ABST
Patent Text Reader

Abstract

The invention provides a soft package battery tab cutting and welding process parameter optimization method and system, and relates to the technical field, the method comprises the following steps: extracting tab cutting cutting force data and cutter displacement track characteristics, and obtaining cutting characteristic parameters; temperature gradient and current characteristics of a welding point are collected, and welding characteristic parameters are obtained; carrying out similarity calculation and importance weighting on the source domain and target domain features by utilizing transfer learning; and iteratively optimizing the process parameters based on reinforcement learning. According to the invention, self-adaptive optimization of tab cutting and welding process parameters is realized, and the stability of battery production quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field, and particularly relates to a soft package battery tab cutting and welding process parameter optimization method and system. BACKGROUND

[0002] The soft package battery is an important energy storage device and is widely used in electric vehicles, portable electronic devices and energy storage systems. In the manufacturing process of the soft package battery, tab cutting and welding is a key process, and the quality directly affects the performance and safety of the battery. The tab is the positive and negative electrode lead-out end of the soft package battery, and bears the function of leading out the internal current of the battery. The quality of the tab cutting and welding process directly affects the internal resistance, cycle performance and safety and reliability of the battery.

[0003] The traditional soft package battery tab cutting and welding process parameter adjustment mainly depends on the experience and trial-and-error method of process personnel, and lacks scientific monitoring and optimization methods. With the development of battery products towards high energy density and high safety, the precision and consistency of the tab cutting and welding process are required more and more. SUMMARY

[0004] The embodiment of the present application provides a soft package battery tab cutting and welding process parameter optimization method and system, which can solve the problems in the prior art.

[0005] In a first aspect, the embodiment of the present application provides a soft package battery tab cutting and welding process parameter optimization method, comprising: The cutting force data of the tab cutting process is collected by the force measuring instrument, the cutting force data is subjected to segmented Hilbert transform, each order inherent modal function is obtained, and the instantaneous frequency and amplitude characteristics are extracted. The displacement trajectory of the tab cutting tool is collected by the displacement sensor, the displacement trajectory is subjected to spatial decomposition, and the trajectory curvature value of the displacement trajectory is extracted. The instantaneous frequency, amplitude characteristics and Gaussian curvature are weighted and fused to obtain the cutting characteristic parameters; The temperature gradient characteristics are extracted by the thermal imager, the current waveform is obtained by the current collector, the amplitude, rise time and duration of the current waveform are calculated as the current characteristics, the temperature gradient characteristics and the current characteristics are weighted and fused to obtain the welding characteristic parameters, and the cutting characteristic parameters and the welding characteristic parameters are stored in the process state characteristic parameter library. The feature parameters corresponding to the historical optimal process state are extracted from the process state feature parameter library as a source domain; the current processing process feature parameters are taken as a target domain, the similarity relationship of the source domain and the target domain features is calculated, and the importance of the source domain features is weighted according to the similarity relationship; the reweighted source domain features and the target domain features are combined to form training samples, the current process feature parameters are taken as state inputs, the process feature parameter adjustment amount is taken as action output, and the quality detection result is taken as a reward signal, and the optimal parameter adjustment strategy is obtained through iterative optimization by the gradient ascent method; when the quality detection results of the continuous multiple products are stable and meet the target requirements, the current process feature parameters are combined and recorded as the optimal process parameters.

[0006] The cutting force data of the tab cutting process is collected by a force measuring instrument, the cutting force data is subjected to piecewise Hilbert transform, each order inherent modal function is obtained, and instantaneous frequency and amplitude characteristics are extracted, including: The cutting force data is convoluted with a Hilbert transform kernel function to obtain a virtual component of the data; the cutting force data is taken as a real component, and the real component and the virtual component are combined to form an analytical signal; the amplitude function and the phase function of the analytical signal are extracted; All local extreme points of the analytical signal are extracted, a cubic spline interpolation with an adaptive tension coefficient is used to generate upper and lower envelope lines, and the tension coefficient is dynamically adjusted according to the distribution density of the extreme points; the mean values of the upper and lower envelope lines are calculated, and a nonlinear correction term is designed to correct the mean values in combination with the local curvature characteristics of the analytical signal; the original cutting force data is subtracted from the corrected mean values to obtain inherent modal components; For each inherent modal component, the phase function of the analytical signal is numerically differentiated to obtain the rate of change of the phase with time; the phase rate of change is normalized to a frequency value to obtain the instantaneous frequency; Based on the analytical signal of each inherent modal component, the local maximum points of the amplitude function are extracted to generate a signal envelope line, and the mean value, fluctuation range and modulation depth of the envelope line are calculated as amplitude characteristics.

[0007] The displacement trajectory of the tab cutting tool is collected by a displacement sensor, the displacement trajectory is subjected to spatial decomposition, and the trajectory curvature value of the displacement trajectory is extracted, including: The displacement trajectory of the cutting tool is projected into a parameterized space, the local curvature of each point of the displacement trajectory is calculated, and the curvature change rate is calculated according to the local curvatures of adjacent sampling points; when the curvature change rate is greater than a curvature fluctuation threshold, the current sampling interval is multiplied by an exponential decay factor for dynamic updating, the displacement trajectory is resampled according to the updated sampling interval to obtain a resampled trajectory; A unit tangent vector is calculated along the motion direction of the resampled trajectory, a unit normal vector is obtained by normalizing the derivative of the unit tangent vector, and the resampled trajectory is projected into a tangent plane determined by the unit tangent vector and the unit normal vector. In the local coordinate system of the tangent plane, the first fundamental form coefficients are calculated by using the first derivative of the resampled trajectory, and the second fundamental form coefficients are calculated by using the inner product of the second derivative of the resampled trajectory and the unit normal vector; and the trajectory curvature value is obtained by substituting the combination of the first fundamental form coefficients and the second fundamental form coefficients into the Gaussian curvature equation.

[0008] Calculate the feature similarity of the source domain and the target domain; and perform importance weighting on the source domain features according to the feature similarity, including: Reconstruct the time-delay trajectory of the feature parameters of the source domain in the phase space, select a time point corresponding to a minimum mutual information value as a time delay, expand the time-delay trajectory in the phase space, count the number of point pairs in the time-delay trajectory with a distance less than a reference distance threshold, divide the number of point pairs by the total number of points in the trajectory to obtain a correlation integral, and calculate the optimal dimensionality according to the correlation integral; Calculate the parameter sequence difference of the feature parameters of the source domain and the feature parameters of the target domain in each dimension, linearly fit the logarithm of the parameter sequence difference and the time sequence; and calculate the feature similarity of each dimension parameter according to the slope of the fitting straight line and the parameter sequence difference; Calculate the fluctuation amplitude of the feature parameters of the target domain in a plurality of continuous time windows, obtain a fluctuation stability indicator by calculating the standard deviation of the fluctuation amplitude, multiply the similarity value and the fluctuation stability indicator to obtain the importance weight of each dimension parameter, and multiply each dimension feature parameter of the source domain by the corresponding importance weight to obtain the reweighted source domain features.

[0009] The reweighted source domain features and the target domain features form a training sample, the current process feature parameters are taken as state inputs, the process feature parameter adjustment amount is taken as action output, and the quality detection result is taken as a reward signal, and an optimal parameter adjustment strategy is obtained by iterative optimization through gradient ascent method, including: A state vector is formed by acquiring the current process feature parameters and the historical quality detection results at a plurality of recent time points, a reasonable adjustment interval of each process parameter is formed by setting the process experience to form an action vector, the action vector is sampled to obtain an action sample set, the probability distribution of each action in the action sample set is counted, and a conditional entropy value is calculated based on the probability distribution as a regularization constraint term; Calculate the difference between the target domain feature parameters and the reweighted source domain feature parameters as a feature consistency loss; input the target domain features into a discrimination network to obtain a discrimination probability, and determine the authenticity of feature conversion based on the discrimination probability to obtain an adversarial loss, weight the adversarial loss, the feature consistency loss, and the conditional entropy regularization constraint term according to a preset weight to obtain a feature generation total loss, and the feature generation total loss guides the optimization direction of the feature generation network. The measured values of the quality indicators of the product are obtained by quality detection, and the deviation degree of the measured values is calculated according to the target values of the quality indicators to obtain a target reward value; the advantage values of each action in the current state are calculated based on the target reward value, the training samples are weighted and summed based on the advantage values, and the optimal parameter adjustment strategy is obtained by iterative optimization through the gradient ascent method.

[0010] The deviation degree of the measured values is calculated according to the target values of the quality indicators to obtain a target reward value; the advantage values of each action in the current state are calculated based on the target reward value, including: The measured values and target values of the plurality of quality indicators are obtained, and the absolute value of the difference between the measured value and the target value of each quality indicator is divided by the fluctuation threshold of the quality indicator to obtain a quality deviation; The cross-correlation coefficient and phase delay between the quality indicators are calculated based on historical data, the indicator pair with a cross-correlation coefficient higher than a correlation threshold is marked as a strongly coupled indicator, the causal order between the indicators is determined by extracting the phase delay, and the quality indicators are divided into a source indicator group and a response indicator group according to the causal order; The quality deviation of the strongly coupled indicator is input into a sliding time window, the mean square deviation of each indicator in the window is calculated, and the indicator group evaluation score is calculated based on the indicator with the maximum mean square deviation; the evaluation scores of each indicator group are weighted and fused to obtain a target reward value; The reward sequence in a plurality of evaluation periods is counted, the target reward mean value and fluctuation trend characteristics are extracted, the difference between the current target reward value and the target reward mean value is taken as a short-term return, the fluctuation trend characteristics are taken as a long-term return, and the action advantage value is determined according to the weighted sum of the short-term return and the long-term return.

[0011] In a second aspect of the embodiment of the present application, a soft package battery tab cutting and welding process parameter optimization system is provided, comprising: A first unit is configured to collect cutting force data of the tab cutting process by a force gauge, perform segmented Hilbert transform on the cutting force data, obtain each order inherent modal function, extract instantaneous frequency and amplitude characteristics, collect displacement trajectory of a tab cutting tool by a displacement sensor, perform spatial decomposition on the displacement trajectory, and extract trajectory curvature value of the displacement trajectory; and the instantaneous frequency, amplitude characteristics and Gaussian curvature are weighted and fused to obtain cutting characteristic parameters; A second unit is configured to obtain temperature gradient characteristics by a thermal imager, obtain welding current waveform by a current collector, calculate amplitude, rise time and duration time of the current waveform as current characteristics, and weightedly fuse the temperature gradient characteristics and the current characteristics to obtain welding characteristic parameters; and the cutting characteristic parameters and the welding characteristic parameters are stored in a process state characteristic parameter library; The third unit is configured to extract the feature parameters corresponding to the historical optimal process state from the process state feature parameter library as a source domain, take the current processing process feature parameters as a target domain, calculate the similarity relationship of the source domain and the target domain features, and weight the importance of the source domain features according to the similarity relationship; the re-weighted source domain features and the target domain features are combined to form a training sample, the current process feature parameters are taken as state input, the process feature parameter adjustment amount is taken as action output, and the quality detection result is taken as a reward signal, and the optimal parameter adjustment strategy is obtained through iterative optimization by the gradient ascent method; when the quality detection results of the continuous multiple products are stable and meet the target requirements, the current process feature parameters are recorded as the optimal process parameters.

[0012] A third aspect of the embodiments of the application, An electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method described above.

[0013] A fourth aspect of the embodiments of the application, A computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.

[0014] The beneficial effects of the present application are as follows: The present application obtains the cutting force data and the tool displacement trajectory in the tab cutting process through the force measuring instrument and the displacement sensor, and extracts the feature parameters by using an advanced signal processing method, so that the tab cutting process is accurately characterized, the burr and deformation defects in the cutting process are effectively reduced, and the tab cutting quality and consistency are improved.

[0015] The present application monitors the temperature distribution and the current waveform change in the welding process by using the thermal imager and the current collector, constructs a complete welding feature parameter system, accurately characterizes the welding state through feature fusion, effectively avoids problems such as poor welding and insufficient welding strength, and improves the strength and stability of the welding joint.

[0016] The present application adopts an intelligent optimization method combining transfer learning and reinforcement learning, calculates the similarity relationship of the source domain and the target domain, realizes efficient transfer of process knowledge and self-adaptive adjustment of parameters, greatly reduces the time cost of process parameter optimization, improves the production efficiency, and at the same time ensures the stability and consistency of the product quality, and is suitable for the process parameter rapid optimization demand in large-scale production of soft package batteries. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1This is a flowchart illustrating the method for optimizing the cutting and welding process parameters of the tabs in a soft-pack battery according to an embodiment of the present invention. Figure 2 A schematic diagram of the process for weighting the importance of source domain features. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0020] Figure 1 This is a flowchart illustrating the method for optimizing the cutting and welding process parameters of the tabs in a soft-pack battery according to an embodiment of the present invention. Figure 1 As shown, the method includes: Cutting force data during the tab cutting process is collected using a force gauge. A piecewise Hilbert transform is performed on the cutting force data to obtain the intrinsic mode functions of each order, and instantaneous frequency and amplitude features are extracted. The displacement trajectory of the tab cutting tool is collected using a displacement sensor, and the spatial domain is decomposed to extract the trajectory curvature value. The instantaneous frequency, amplitude features, and Gaussian curvature are then weighted and fused to obtain the cutting feature parameters. Temperature data of the welding point is acquired by a thermal imager, and temperature gradient features are extracted; welding current waveform is acquired by a current acquisition device, and the amplitude, rise time and duration of the current waveform are calculated as current features; the temperature gradient features and the current features are weighted and fused to obtain welding feature parameters; the cutting feature parameters and the welding feature parameters are stored in the process state feature parameter library. The source domain is obtained by extracting the feature parameters corresponding to the historical best process state from the process state feature parameter library; the target domain is obtained by taking the process feature parameters of the current process; the similarity relationship between the features of the source domain and the target domain is calculated, and the source domain features are weighted according to the similarity relationship; the reweighted source domain features and the target domain features are combined to form training samples; the current process feature parameters are used as the state input, the process feature parameter adjustment amount is used as the action output, and the quality inspection result is used as the reward signal; the optimal parameter adjustment strategy is obtained by iterative optimization through the gradient ascent method; when the quality inspection results of multiple consecutive products stably meet the target requirements, the current process feature parameter combination is recorded as the optimal process parameters.

[0021] In one optional implementation, cutting force data during the tab cutting process is collected using a force gauge. A piecewise Hilbert transform is performed on the cutting force data to obtain the intrinsic mode functions of each order, and instantaneous frequency and amplitude features are extracted, including: The cutting force data is convolved with the Hilbert transform kernel function to obtain the imaginary component of the data; the cutting force data is used as the real component and combined with the imaginary component to form an analytic signal; the amplitude function and phase function of the analytic signal are extracted. All local extreme points of the analytical signal are extracted, and cubic spline interpolation with an adaptive tension coefficient is introduced to generate upper and lower envelopes. The tension coefficient is dynamically adjusted according to the distribution density of extreme points. The mean of the upper and lower envelopes is calculated and combined with the local curvature characteristics of the analytical signal. A nonlinear correction term is designed to correct the mean. The original cutting force data is subtracted from the corrected mean to obtain the intrinsic modal components. For each intrinsic mode component, the phase function of its analytical signal is numerically differentiated to obtain the rate of phase change over time; the rate of phase change is normalized to a frequency value to obtain the instantaneous frequency. Based on the analytical signal of each intrinsic mode component, the local maxima of the amplitude function are extracted, the signal envelope is generated, and the mean, fluctuation range and modulation depth of the envelope are calculated as amplitude features.

[0022] In this embodiment, a high-precision piezoelectric force gauge is connected to the tab cutting device. The sampling frequency of the force gauge is set to 10kHz to ensure that high-frequency cutting force changes can be captured. The force gauge simultaneously collects cutting force data in the X, Y, and Z directions, where the X direction corresponds to the main cutting force in the cutting direction, the Y direction corresponds to the normal force perpendicular to the cutting surface, and the Z direction corresponds to the lateral force parallel to the cutting plane. The collected raw cutting force data is stored as a time series, and the force component in each direction is stored separately as a data vector.

[0023] Piecewise Hilbert transform was performed on the acquired cutting force data. The entire cutting process was divided into multiple time windows, with a window length of 512 data points and an overlap rate of 50% between adjacent windows. For the data within each window, the Hilbert transform was applied to extract the envelope features of the force signal. The Hilbert transform was implemented through convolution, using the original cutting force data as input and convolving it with a Hilbert transform kernel function. The kernel function was defined as an odd function in the time domain, and the convolution calculation was accelerated using a Fast Fourier Transform (FFT). Through the convolution operation, a sequence of imaginary components with the same length as the original cutting force data was obtained. In actual operation, the peak value of the X-direction cutting force data acquired in a certain experiment within one time window was 24.5 N; after the Hilbert transform, the corresponding imaginary component value at the same time was 18.3 N.

[0024] The original cutting force data is used as the real component and combined with the calculated imaginary component to form an analytic signal in complex form. The analytic signal has significant physical meaning; its amplitude represents the instantaneous magnitude of the signal, and its phase represents the instantaneous phase. During data processing, the amplitude function is extracted from the analytic signal by calculating the square root of the sum of the squares of the real and imaginary parts; simultaneously, the phase function is extracted by calculating the arctangent of the imaginary and real parts. For the data points in the aforementioned example, the amplitude of the analytic signal is 30.6 N, and the phase is 0.64 radians.

[0025] All local extrema, including maxima and minima, are extracted from the raw cutting force data. Considering the typical cutting force fluctuations during tab cutting, approximately 12-15 local extrema are found every 50ms of data. Based on these extrema, an upper and lower envelope are generated using a cubic spline interpolation method with adaptive tension coefficients. The tension coefficient is dynamically adjusted according to the distribution density of the extrema, using a smaller coefficient (e.g., 0.3) in densely populated areas and a larger coefficient (e.g., 0.7) in sparsely populated areas to ensure the envelope accurately fits the overall trend of the signal without overfitting or underfitting.

[0026] The mean of the upper and lower envelopes is calculated as the local trend line of the signal. A nonlinear correction term is designed to correct the mean, based on the local curvature characteristics of the analytic signal. The curvature characteristics are obtained by calculating the second derivative of the phase function of the analytic signal. In regions of high signal curvature, the correction coefficient is increased accordingly (up to a maximum of 1.5 times), while in regions of low curvature, the correction coefficient is decreased accordingly (down to a minimum of 0.7 times). The corrected mean is subtracted from the original cutting force data to obtain the first intrinsic mode component. The above process is repeated to further decompose the first intrinsic mode component until a predetermined number of intrinsic mode components are obtained or the remaining signal meets the termination condition. In practical applications, extracting 3-5 intrinsic mode components is usually sufficient to cover the main frequency characteristics during the tab cutting process.

[0027] For each extracted intrinsic mode component, the phase function of its analytical signal is numerically differentiated, and the rate of phase change over time is calculated using the central difference method. For two adjacent sampling points, the phase difference is calculated and divided by the time interval to obtain the instantaneous angular frequency at that time point. The angular frequency is then divided by 2π to convert it to an instantaneous frequency in Hertz units. In actual data processing, the instantaneous frequencies of the first intrinsic mode component are mainly distributed in the range of 800-1200Hz, corresponding to the basic cutting characteristics of the tab material; the instantaneous frequencies of the second intrinsic mode component are mainly distributed in the range of 1500-2500Hz, corresponding to the frictional vibration characteristics between the tool and the material; and the instantaneous frequencies of the third intrinsic mode component are mainly distributed in the range of 3000-4500Hz, corresponding to the high-frequency vibration characteristics of the equipment system.

[0028] Based on the analytical signal of each intrinsic mode component, local maxima of its amplitude function are extracted. These points typically represent local peaks in signal energy. The amplitude envelope of the signal is generated by connecting these maxima. The mean of the envelope is calculated as a signal strength feature, the difference between the maximum and minimum values ​​of the envelope is calculated as a fluctuation range feature, and the ratio of the fluctuation range to the mean is calculated as a modulation depth feature. In a certain electrode cutting experiment, the mean amplitude of the first intrinsic mode component was 5.6N, the fluctuation range was 3.2N, and the modulation depth was 0.57; the mean amplitude of the second intrinsic mode component was 3.1N, the fluctuation range was 2.4N, and the modulation depth was 0.77; the mean amplitude of the third intrinsic mode component was 1.8N, the fluctuation range was 1.5N, and the modulation depth was 0.83.

[0029] The instantaneous frequency and amplitude features extracted through the above process constitute a time-frequency feature set of the cutting force signal in the tab cutting process. These features can effectively characterize various dynamic characteristics in the cutting process, providing a data basis for subsequent condition monitoring and fault diagnosis.

[0030] In one optional implementation, the displacement trajectory of the tab cutting tool is acquired by a displacement sensor, and the displacement trajectory is spatially decomposed to extract the trajectory curvature value, including: The displacement trajectory of the cutting tool is projected onto the parameterized space, and the local curvature of each point on the displacement trajectory is calculated. The curvature change rate is calculated based on the local curvature of adjacent sampling points. When the curvature change rate is greater than the curvature fluctuation threshold, the current sampling interval is multiplied by an exponential decay factor for dynamic updating. The displacement trajectory is resampled based on the updated sampling interval to obtain the resampled trajectory. Calculate the unit tangent vector along the motion direction of the resampled trajectory, normalize the derivative of the unit tangent vector to obtain the unit normal vector, and project the resampled trajectory onto the tangent plane determined by the unit tangent vector and the unit normal vector; In the local coordinate system of the tangent plane, the first fundamental form coefficients are calculated using the first derivative of the resampled trajectory, and the second fundamental form coefficients are calculated using the inner product of the second derivative of the resampled trajectory and the unit normal vector; the combination of the first and second fundamental form coefficients is substituted into the Gaussian curvature equation to obtain the trajectory curvature value.

[0031] This embodiment provides a technical implementation method for acquiring the displacement trajectory of a tab cutting tool and calculating the trajectory curvature. The method utilizes a displacement sensor to acquire the displacement trajectory data of the tab cutting tool. The displacement sensor used can be a linear displacement sensor or a rotary encoder, with a sampling frequency set to 1000Hz to ensure the capture of subtle changes in the cutting tool's movement. The acquired raw data includes a timestamp and the corresponding position coordinates in the x, y, and z directions.

[0032] Projecting the displacement trajectory of the cutting tool into a parameterized space means representing the trajectory points in three-dimensional space as a function of parameter t, P(t)=(x(t),y(t),z(t)), where t can be understood as a normalized time parameter with a value range of [0,1]. For N discrete points collected, the parameter value of the i-th point can be set to t. i =i / (N-1).

[0033] When calculating the local curvature at each point on the displacement trajectory, for a parameterized curve P(t), its local curvature κ can be obtained through three adjacent points P(t). i -1), P(t) i ) and P(t i The positional relationship of +1) is determined. Specifically, the curvature can be estimated by calculating the reciprocal of the radius of the circle determined by these three points. For example, for a sequence of acquired trajectory points, the local curvature calculation of the 100th point involves the coordinate values ​​of the 99th, 100th, and 101st points. Assuming the coordinates of these three points are (10.2, 15.3, 5.1), (10.5, 15.5, 5.2), and (10.9, 15.8, 5.3), the local curvature value of this point can be calculated to be approximately 0.052.

[0034] The rate of curvature change is calculated by dividing the curvature difference between two adjacent sampling points by the parameter interval between them. When the rate of curvature change exceeds a preset curvature fluctuation threshold (e.g., set to 0.01), it indicates that the trajectory is changing drastically in that region, requiring more refined sampling. In this case, the current sampling interval is dynamically updated by multiplying it by an exponential decay factor (e.g., a value of 0.75). For example, if the original sampling interval is 0.01, in regions with significant curvature changes, the updated sampling interval will become 0.0075.

[0035] The displacement trajectory is resampled based on the updated dynamic sampling interval using cubic spline interpolation. A larger sampling interval is used in regions with gentle curvature changes, while a smaller interval is used in regions with dramatic curvature changes. This results in a resampled trajectory that more accurately reflects the characteristics of the tool motion. In practical applications, the original trajectory contains 500 sampling points, and after resampling, 650 points are obtained, with approximately 30% of the points concentrated in regions of significant curvature changes.

[0036] Calculating the unit tangent vector along the direction of motion of the resampled trajectory means that for each point on the resampled trajectory, the unit tangent vector T is obtained by calculating the first derivative of the trajectory at that point and normalizing it. For example, for a point on the trajectory, if its parametric derivative vector is (0.3, 0.4, 0.5), then the normalized unit tangent vector is (0.424, 0.566, 0.707).

[0037] The unit normal vector N is calculated by normalizing the derivative of the unit tangent vector T. If the derivative of the tangent vector is (0.02, -0.03, 0.01), the normalized unit normal vector is (0.535, -0.802, 0.267). The tangent plane is determined by the unit tangent vector T and the unit normal vector N at point P, and the resampled trajectory is projected onto this tangent plane.

[0038] In the local coordinate system of the tangent plane, the first fundamental form coefficients are parameters describing the metric properties on the tangent plane. The first fundamental form coefficients E, F, and G can be calculated using the first derivative of the resampled trajectory. Taking a specific point as an example, if the first parametric derivative components at that point are dx / dt = 0.8, dy / dt = 0.6, and dz / dt = 0.5, then E = dx / dt. 2 +dy / dt 2 +dz / dt 2 =1.25.

[0039] The second fundamental form coefficients are parameters describing the curvature of the surface, calculated by the dot product of the second derivative of the resampled trajectory and the unit normal vector. If the second derivative at a point is (0.05, -0.03, 0.02) and the unit normal vector is (0.535, -0.802, 0.267), then the dot product of the second derivative and the normal vector is 0.065. This value is used to calculate the second fundamental form coefficients L, M, and N.

[0040] Substituting the first and second fundamental form coefficients into the Gaussian curvature calculation formula yields the trajectory curvature value. For example, a Gaussian curvature value of 0.008 at a certain point indicates the curvature characteristics of the cutting tool's motion trajectory at that point. By analyzing the curvature of the entire trajectory, the motion state and working performance of the cutting tool can be evaluated, providing a basis for optimizing the tab cutting process.

[0041] By using this spatial decomposition and curvature analysis method of displacement trajectory, precise monitoring of the tab cutting process can be achieved, abnormal states can be identified early, and the quality and production efficiency of battery tab cutting can be improved.

[0042] Figure 2A schematic diagram illustrating the process of weighting the importance of source domain features. In one optional implementation, the feature similarity between the source and target domains is calculated; the source domain features are then weighted according to their similarity, including: Reconstruct the time delay trajectory of the feature parameters of the source domain in the phase space, and select the time point corresponding to the minimum mutual information value as the time delay; expand the time delay trajectory in the phase space, count the number of point pairs in the time delay trajectory whose distance is less than the reference distance threshold, divide the number of point pairs by the total number of points in the trajectory to obtain the correlation integral, and calculate the optimal dimension based on the correlation integral; Calculate the parameter sequence differences of each dimension of the feature parameters of the source domain and the feature parameters of the target domain, and perform linear fitting between the logarithm of the parameter sequence differences and the logarithm of the time series; calculate the feature similarity of each dimension of the parameters based on the slope of the fitted line and the parameter sequence differences. The fluctuation amplitude of the feature parameters of the target domain within multiple consecutive time windows is statistically analyzed, and the standard deviation of the fluctuation amplitude is calculated to obtain the fluctuation stability index. The similarity value is multiplied by the fluctuation stability index and normalized to obtain the importance weight of each dimension parameter. The feature parameters of each dimension of the source domain are multiplied by their corresponding importance weights to obtain the reweighted source domain features.

[0043] In this embodiment, a method is provided to calculate the feature similarity between the source domain and the target domain, and to weight the source domain features based on the feature similarity.

[0044] The first step in this method for calculating the feature similarity between the source and target domains is to reconstruct the time-delay trajectory of the source domain feature parameters in phase space. Specifically, feature parameter sequences are extracted from the source domain dataset, such as collecting parameters like temperature, pressure, and flow rate from industrial equipment to form time-series data. For each feature parameter, the optimal time delay is determined by calculating the mutual information value at different time lags. For example, for the temperature parameter, its mutual information value can be calculated at time lags of 1 second, 2 seconds, 3 seconds... up to 30 seconds. If the mutual information value is found to be the smallest at a time lag of 7 seconds, then 7 seconds is selected as the optimal time delay for that feature parameter. Using the determined time delay, the trajectory of the feature parameter is reconstructed in phase space. Taking a time delay of 7 seconds as an example, if the original temperature sequence is [T1, T2, T3, ..., T7], then the reconstructed trajectory points are [T1, T8, T...]. 15 [T2, T9, T], ...] 16 , ...], ..., [T7, T 14 , T 21 , ...]wait.

[0045] After determining the time delay, the optimal embedding dimension of the trajectory needs to be calculated. The time-delayed trajectory is unfolded in phase space, and a reference distance threshold is set, for example, 5% of the feature parameter value range. The number of point pairs in the trajectory whose distance is less than the reference distance threshold is counted. For example, if there are 1000 points for the temperature parameter in phase space, and 150 pairs are less than the threshold, then the correlation integral is 150 / 1000 = 0.15. By gradually increasing the embedding dimension (e.g., from 2 to 10 dimensions), the correlation integral at each dimension is calculated. When the correlation integral tends to stabilize with increasing embedding dimension, the optimal dimension is determined. For example, when the embedding dimension increases from 5 to 6, the correlation integral changes from 0.15 to 0.149, which is a small change, so the optimal dimension can be determined to be 5.

[0046] When calculating the similarity of feature parameters between the source and target domains, the differences between the parameter sequences of each dimension of the two domains are first calculated. For example, if the source domain temperature sequence is [25℃, 26℃, 27℃, 28℃, 29℃] and the target domain temperature sequence is [35℃, 37℃, 39℃, 41℃, 43℃], then the difference sequence is [10℃, 11℃, 12℃, 13℃, 14℃]. The logarithms of the difference sequence and the corresponding time series are taken respectively, resulting in the logarithmic difference sequence [2.30, 2.40, 2.48, 2.56, 2.64] and the logarithmic time series [0, 0.69, 1.10, 1.39, 1.61]. Linear fitting using the least squares method yields a fitted line with a slope of 0.2. Feature similarity is calculated based on the slope and the difference between parameter sequences. The similarity value is 1 / (1+0.2*average difference) = 1 / (1+0.2*12) = 0.29. This process is repeated for all feature parameters to obtain the feature similarity for each dimension.

[0047] Assessing the stability of the fluctuations in the target domain's characteristic parameters is also crucial. The target domain data is divided into multiple consecutive time windows, for example, every 10 minutes. Within each window, the difference between the maximum and minimum values ​​of the characteristic parameter is calculated to obtain the fluctuation amplitude. For example, if the fluctuation amplitudes of the target domain temperature parameter within five time windows are [2℃, 3℃, 2.5℃, 2.8℃, 2.7℃], the standard deviation of these fluctuation amplitudes is calculated, yielding a fluctuation stability index of 0.36. The smaller the fluctuation stability index, the more stable the parameter fluctuations and the higher the reliability.

[0048] The importance weights of each dimension parameter are calculated by combining feature similarity with the fluctuation stability index. Specifically, the similarity value is multiplied by the inverse of the fluctuation stability index and normalized to obtain the importance weight. For example, for the temperature parameter, the importance weight is calculated as 0.29*(1 / 0.36)=0.81. Assuming the system has three feature parameters: temperature, pressure, and flow rate, their importance weights are 0.81, 0.63, and 0.45, respectively, and after normalization, they are 0.43, 0.33, and 0.24. The importance weights of each dimension feature parameter of the source domain are multiplied by the corresponding importance weight to obtain the reweighted source domain features. For example, if the source domain temperature parameter is 25℃, the reweighted value is 25℃*0.43=10.75℃.

[0049] The above method achieves source domain feature importance weighting based on feature similarity, improving the adaptability of transfer learning across different domains. Experiments show that this method improves accuracy by more than 15% compared to traditional methods in applications such as predictive maintenance of industrial equipment and cross-domain recommendation systems, reducing the performance degradation caused by domain differences.

[0050] In one optional implementation, the reweighted source domain features and target domain features are used to form training samples. The current process feature parameters are used as state input, the process feature parameter adjustment amount is used as action output, and the quality detection result is used as a reward signal. The optimal parameter adjustment strategy is obtained through iterative optimization using the gradient ascent method, including: A state vector is formed by acquiring the current process feature parameters and the historical quality detection results from multiple recent moments. An action vector is formed by setting a reasonable adjustment range for each process parameter based on process experience. The action vectors are sampled to obtain an action sample set. The probability distribution of each action in the action sample set is statistically analyzed. The conditional entropy value is calculated based on the probability distribution as a regularization constraint. The difference between the target domain feature parameters and the reweighted source domain feature parameters is calculated as the feature consistency loss. The target domain features are input into a discriminant network to obtain the discrimination probability. Based on the discrimination probability, the authenticity of the feature transformation is determined to obtain the adversarial loss. The adversarial loss, feature consistency loss, and conditional entropy regularization constraint are weighted according to preset weights to obtain the total feature generation loss. The total feature generation loss guides the optimization direction of the feature generation network. The product undergoes quality testing to obtain measured values ​​of multiple quality indicators. Based on the target values ​​of each quality indicator, the deviation of the measured values ​​is calculated to obtain the target reward value. Based on the target reward value, the advantage value of each action in the current state is calculated. Based on the advantage value, the training samples are weighted and summed. The optimal parameter adjustment strategy is obtained through iterative optimization using the gradient ascent method.

[0051] This invention provides an automatic optimization method for process parameters based on domain adaptation reinforcement learning. The method uses reweighted source domain features and target domain features to form training samples, takes the current process feature parameters as the state input, takes the process feature parameter adjustment amount as the action output, takes the quality detection result as the reward signal, and obtains the optimal parameter adjustment strategy through gradient ascent method iterative optimization.

[0052] In practical applications, the process parameter optimization system collects process characteristic data from both the source and target domains. Source domain data includes historical production process parameters and corresponding quality indicators. For example, in an aluminum alloy extrusion process, 5000 sets of historical data were collected, each set containing 10 process parameters such as extrusion temperature, extrusion speed, and cooling rate, as well as 5 quality indicators such as product strength and hardness. Target domain data consists of real-time process parameter data collected from the current production line, such as 200 sets of data from currently ongoing new product development.

[0053] When reweighting the source domain data, the system constructs a feature generation network and a discriminator network. The feature generation network adopts a three-layer fully connected neural network structure, with the number of nodes in the input layer being 10 times the feature dimension of the source domain, the number of nodes in the intermediate layers being 20 and 15 respectively, and the number of nodes in the output layer being 10 times the feature dimension of the target domain. The discriminator network also adopts a three-layer fully connected structure, with the number of nodes in the input layer being 10 times the feature dimension, the number of nodes in the intermediate layers being 15, and the number of nodes in the output layer being 1, used to determine whether the input features come from the source domain or the target domain.

[0054] The feature generation network maps source domain features to a space similar to the target domain features, achieving source-to-target domain transfer. For example, if the extrusion temperature range in the source domain is 450-500℃, while the target domain is 470-520℃, the feature generation network can map the source domain temperature parameters to the distribution range of the target domain. The discriminant network attempts to distinguish between the transformed source domain features and the true target domain features; the two networks achieve feature distribution alignment through adversarial training.

[0055] When acquiring current process parameters and historical quality results, the system collects the process parameter values ​​and corresponding quality inspection results from the last five moments to form a state vector. For example, if the current extrusion temperature is 485℃, and the strengths of the last five quality inspections are 245MPa, 248MPa, 250MPa, 247MPa, and 249MPa, these data, along with other process parameters, constitute the state vector.

[0056] When setting parameter adjustment ranges based on process experience, the system limits the adjustment range of each process parameter to a reasonable range. For example, the adjustment range for extrusion temperature is [-5℃, +5℃], the adjustment range for extrusion speed is [-2mm / s, +2mm / s], and the adjustment range for cooling rate is [-1℃ / s, +1℃ / s]. The system samples and generates a set of action samples within these ranges. For example, for the temperature parameter, it samples five adjustment values: -5℃, -2.5℃, 0℃, +2.5℃, and +5℃. Other parameters are processed similarly and combined to form a set of action samples.

[0057] The system statistically analyzes the probability distribution of each action in the action sample set and calculates the conditional entropy value as a regularization constraint. For example, if the probability of the temperature adjustment value of +2.5℃ is 0.25, the probability distribution of other parameter adjustment values ​​is calculated in a similar way. The conditional entropy value is calculated using these probability distributions and serves as a measure of the degree of exploration.

[0058] When calculating the difference between the target domain features and the reweighted source domain features, the system calculates the mean squared error for each feature dimension. For example, if the mean temperature parameter of the transformed source domain is 482℃ and the mean temperature parameter of the target domain is 485℃, the difference is 3℃. The differences across all feature dimensions constitute the feature consistency loss. The target domain features are input into the discriminant network to obtain the discrimination probability. For example, a discriminant network output of 0.65 indicates a 65% probability of identifying the features as target domain features. Based on this, the adversarial loss is calculated.

[0059] The system weights the adversarial loss, feature consistency loss, and conditional entropy regularization constraint term with weights of 0.4, 0.4, and 0.2 respectively to obtain the total feature generation loss. This loss value is used to guide the optimization direction of the feature generation network and updates the network parameters through backpropagation.

[0060] When conducting quality inspections on products, the system measures multiple quality indicators, such as a target strength value of 250 MPa and a measured value of 245 MPa, with a deviation of -2%; and a target hardness value of 95 HB and a measured value of 92 HB, with a deviation of -3.2%. Based on these deviations, the system calculates a comprehensive target reward value, for example, by inverting the weighted average of the deviations of each indicator as the reward value.

[0061] Based on the target reward value, the system calculates the advantage value of each action in the current state. For example, in the current state, the advantage value of the action that increases the temperature by 2.5℃ is 0.85, meaning that this action is 0.85 units better than the average level. The system uses these advantage values ​​to perform a weighted summation on the training samples and iteratively optimizes the policy network parameters using the gradient ascent method.

[0062] After 1000 rounds of iterative training, the system obtains the optimal parameter adjustment strategy. In practical applications, after inputting the current process state, the system outputs optimal parameter adjustment suggestions, such as temperature +2.5℃, speed -1mm / s, and cooling rate +0.5℃ / s. Experimental verification shows that the optimized process parameters using this method increase the product quality pass rate from 92% to 97.5%, and the optimization process requires an average of only 5 iterations, far fewer than the 15-20 adjustments required by traditional methods, significantly improving production efficiency.

[0063] In one optional implementation, a target reward value is obtained by calculating the deviation of the measured values ​​based on the target values ​​of each quality indicator; the advantage value of each action in the current state is calculated based on the target reward value, including: Obtain the measured values ​​and target values ​​of multiple quality indicators, and divide the absolute value of the difference between the measured value and the target value of each quality indicator by the fluctuation threshold of that quality indicator to obtain the quality deviation. Based on historical data, the cross-correlation coefficient and phase delay between quality indicators are calculated. Indicator pairs with cross-correlation coefficients higher than the correlation threshold are marked as strongly coupled indicators. The phase delay is extracted to determine the causal order between indicators. Based on the causal order, the quality indicators are divided into source indicator group and response indicator group. The quality deviation of the strongly coupled indicators is input into the sliding time window, and the mean square deviation of each indicator within the window is calculated. The indicator with the largest mean square deviation is used as the benchmark to calculate the evaluation score of the indicator group. The evaluation scores of each indicator group are weighted and fused to obtain the target reward value. By statistically analyzing the reward sequences over multiple evaluation periods, the mean target reward and fluctuation trend characteristics are extracted. The difference between the current target reward value and the mean target reward is taken as the short-term gain, and the fluctuation trend characteristics are taken as the long-term gain. The action advantage value is determined based on the weighted sum of the short-term and long-term gains.

[0064] This invention provides a method for calculating target reward values ​​and determining action advantage values ​​based on quality indicators. In industrial production processes, multiple quality indicators often have complex interrelationships. By analyzing these relationships, the production status can be more accurately assessed and subsequent control decisions can be guided.

[0065] The system acquires measured and target values ​​for multiple quality indicators. Taking the steel production process as an example, it can monitor indicators such as temperature, pressure, and component content. Assume a production line monitors five quality indicators: temperature (target value 1200℃, measured value 1220℃, fluctuation threshold set at 30℃); pressure (target value 5.5MPa, measured value 5.2MPa, fluctuation threshold 0.5MPa); carbon content (target value 0.45%, measured value 0.48%, fluctuation threshold 0.05%); sulfur content (target value 0.03%, measured value 0.035%, fluctuation threshold 0.01%); and hardness (target value 45HRC, measured value 43HRC, fluctuation threshold 3HRC). The system divides the absolute value of the difference between the measured and target values ​​for each quality indicator by the fluctuation threshold for that indicator to obtain the normalized quality deviation. For example, the mass deviation of temperature is |1220-1200| / 30=0.67, the mass deviation of pressure is |5.2-5.5| / 0.5=0.6, and so on to obtain the deviation values ​​of other indicators.

[0066] The system calculates the cross-correlation coefficients and phase delays between quality indicators based on historical data. By analyzing 30 consecutive days of production data, the system calculates the cross-correlation coefficients for each indicator pairwise. Assuming the calculation results show that the cross-correlation coefficient between temperature and carbon content is 0.85 with a phase delay of 2 time units; the cross-correlation coefficient between temperature and hardness is 0.72 with a phase delay of 3 time units; and the cross-correlation coefficient between pressure and sulfur content is 0.65 with a phase delay of 1 time unit, and setting the correlation threshold to 0.7, temperature and carbon content, and temperature and hardness are marked as strongly coupled indicator pairs. According to phase delay analysis, temperature changes precede changes in carbon content and hardness; therefore, temperature is classified into the source indicator group, while carbon content and hardness are classified into the response indicator group.

[0067] The system uses a sliding time window to handle quality deviations of strongly coupled indicators. Taking temperature, carbon content, and hardness as examples, a 12-hour sliding window is set, with data sampled hourly. For each indicator, the mean square deviation (MSD) of the 12 sample points within the window is calculated. Assuming the calculation results show that the MSD for temperature is 0.55, for carbon content it is 0.42, and for hardness it is 0.38, the indicator with the largest MSD deviation (temperature) is used as the benchmark to calculate the evaluation score for the source indicator group. The evaluation score for the source indicator group is calculated by multiplying the temperature MSD deviation by a weighting factor (set to 1.2), resulting in 0.66. The evaluation score for the response indicator group is the weighted average of the MSD deviations for carbon content and hardness (weights of 0.6 and 0.4 respectively), resulting in 0.404. Considering the evaluation scores of both indicator groups, with the source indicator group weighted at 0.7 and the response indicator group weighted at 0.3, the target reward value is 0.66 × 0.7 + 0.404 × 0.3 = 0.583.

[0068] To evaluate the long-term effects of the control actions, the system statistically analyzes the reward sequence over the past 10 evaluation periods. Assuming the target reward value sequence for the past 10 periods is [0.583, 0.562, 0.601, 0.548, 0.572, 0.590, 0.610, 0.595, 0.568, 0.576], the calculated average target reward is 0.581. The difference of 0.002 between the current target reward value of 0.583 and the average of 0.581 is taken as the short-term gain. By analyzing the changing trend of the reward sequence, fluctuation trend characteristics can be extracted. Using a linear regression method, the slope of the reward sequence is calculated, yielding a slope value of 0.0015, indicating a slight upward trend in the reward value; this trend characteristic is taken as the long-term gain. If the weight of short-term returns is 0.4 and the weight of long-term returns is 0.6, then the overall advantage value is 0.002×0.4+0.0015×0.6=0.0018.

[0069] For multiple selectable control actions, the system calculates their respective advantage values ​​using the same method. Assume the system considers three adjustment actions: increasing heating power, decreasing feed rate, and adjusting cooling water flow rate. Simulation predictions show that increasing heating power has an advantage value of 0.0018, decreasing feed rate has an advantage value of 0.0012, and adjusting cooling water flow rate has an advantage value of 0.0025. The action with the highest advantage value (adjusting cooling water flow rate) will be selected as the optimal control strategy.

[0070] In practical applications, the system can adjust calculation parameters according to production characteristics. For example, for rapidly changing processes, the sliding window length can be shortened; for steady-state processes, the historical data analysis cycle can be increased. The system can also introduce a feedback mechanism to dynamically adjust the weights of each indicator based on the actual control effect, thereby making the reward calculation more accurately reflect the production goals.

[0071] The present invention provides a system for optimizing the cutting and welding process parameters of tabs for pouch batteries, comprising: The first unit is used to collect cutting force data during the tab cutting process using a force measuring instrument, perform piecewise Hilbert transform on the cutting force data to obtain the intrinsic mode functions of each order, and extract instantaneous frequency and amplitude features; collect the displacement trajectory of the tab cutting tool using a displacement sensor, perform spatial decomposition on the displacement trajectory, and extract the trajectory curvature value; and weight and fuse the instantaneous frequency, amplitude features, and Gaussian curvature to obtain cutting feature parameters. The second unit is used to acquire welding point temperature data through a thermal imager and extract temperature gradient features; acquire welding current waveform through a current acquisition device and calculate the amplitude, rise time and duration of the current waveform as current features; weight and fuse the temperature gradient features and the current features to obtain welding feature parameters; and store the cutting feature parameters and the welding feature parameters into a process state feature parameter library. The third unit is used to extract the feature parameters corresponding to the historical best process state from the process state feature parameter library as the source domain; take the current process feature parameters as the target domain, calculate the similarity relationship between the features of the source domain and the target domain, and weight the source domain features according to the similarity relationship; combine the reweighted source domain features and target domain features to form training samples, take the current process feature parameters as the state input, take the process feature parameter adjustment amount as the action output, take the quality inspection result as the reward signal, and obtain the optimal parameter adjustment strategy through gradient ascent method iterative optimization; when the quality inspection results of multiple consecutive products stably meet the target requirements, the current process feature parameter combination is recorded as the optimal process parameters.

[0072] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0073] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0074] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing the cutting and welding process parameters of tabs in soft-pack batteries, characterized in that, include: Cutting force data during the tab cutting process is collected by a force measuring instrument. Piecewise Hilbert transform is performed on the cutting force data to obtain the natural mode functions of each order and extract the instantaneous frequency and amplitude features. The displacement trajectory of the tab cutting tool is acquired by a displacement sensor, the displacement trajectory is decomposed in the spatial domain, and the trajectory curvature value is extracted; the instantaneous frequency, amplitude features and Gaussian curvature are weighted and fused to obtain the cutting feature parameters. Temperature data of the welding point is acquired by a thermal imager, and temperature gradient features are extracted; welding current waveform is acquired by a current acquisition device, and the amplitude, rise time and duration of the current waveform are calculated as current features; the temperature gradient features and the current features are weighted and fused to obtain welding feature parameters; the cutting feature parameters and the welding feature parameters are stored in the process state feature parameter library. Extract the feature parameters corresponding to the historical best process state from the process state feature parameter library as the source domain; take the process feature parameters of the current process as the target domain, calculate the similarity relationship between the features of the source domain and the target domain, and weight the features of the source domain according to the similarity relationship; The reweighted source domain features and target domain features are combined to form training samples. The current process feature parameters are used as state input, the process feature parameter adjustment amount is used as action output, and the quality detection result is used as reward signal. The optimal parameter adjustment strategy is obtained by iterative optimization through gradient ascent method. When the quality inspection results of multiple consecutive products consistently meet the target requirements, the current combination of process characteristic parameters is recorded as the optimal process parameters.

2. The method according to claim 1, characterized in that, Cutting force data during the tab cutting process is collected using a force gauge. A piecewise Hilbert transform is performed on the cutting force data to obtain the natural mode functions of each order. Instantaneous frequency and amplitude features are extracted, including: The cutting force data is convolved with the Hilbert transform kernel function to obtain the imaginary component of the data; the cutting force data is used as the real component and combined with the imaginary component to form an analytic signal; the amplitude function and phase function of the analytic signal are extracted. All local extreme points of the analytical signal are extracted, and cubic spline interpolation with an adaptive tension coefficient is introduced to generate upper and lower envelopes. The tension coefficient is dynamically adjusted according to the distribution density of extreme points. The mean of the upper and lower envelopes is calculated and combined with the local curvature characteristics of the analytical signal. A nonlinear correction term is designed to correct the mean. The original cutting force data is subtracted from the corrected mean to obtain the intrinsic modal components. For each intrinsic mode component, the phase function of its analytical signal is numerically differentiated to obtain the rate of phase change over time; the rate of phase change is normalized to a frequency value to obtain the instantaneous frequency. Based on the analytical signal of each intrinsic mode component, the local maxima of the amplitude function are extracted, the signal envelope is generated, and the mean, fluctuation range and modulation depth of the envelope are calculated as amplitude features.

3. The method according to claim 1, characterized in that, The displacement trajectory of the tab cutting tool is acquired by a displacement sensor, and the spatial domain decomposition of the displacement trajectory is performed to extract the trajectory curvature value, including: The displacement trajectory of the cutting tool is projected onto the parameterized space, and the local curvature of each point on the displacement trajectory is calculated. The curvature change rate is calculated based on the local curvature of adjacent sampling points. When the curvature change rate is greater than the curvature fluctuation threshold, the current sampling interval is multiplied by an exponential decay factor for dynamic updating. The displacement trajectory is resampled based on the updated sampling interval to obtain the resampled trajectory. Calculate the unit tangent vector along the motion direction of the resampled trajectory, normalize the derivative of the unit tangent vector to obtain the unit normal vector, and project the resampled trajectory onto the tangent plane determined by the unit tangent vector and the unit normal vector; In the local coordinate system of the tangent plane, the first fundamental form coefficients are calculated using the first derivative of the resampled trajectory, and the second fundamental form coefficients are calculated using the inner product of the second derivative of the resampled trajectory and the unit normal vector; the combination of the first and second fundamental form coefficients is substituted into the Gaussian curvature equation to obtain the trajectory curvature value.

4. The method according to claim 1, characterized in that, Calculate the feature similarity between the source domain and the target domain; Source domain features are weighted based on feature similarity, including: Reconstruct the time delay trajectory of the feature parameters of the source domain in the phase space, and select the time point corresponding to the minimum mutual information value as the time delay; expand the time delay trajectory in the phase space, count the number of point pairs in the time delay trajectory whose distance is less than the reference distance threshold, divide the number of point pairs by the total number of points in the trajectory to obtain the correlation integral, and calculate the optimal dimension based on the correlation integral; Calculate the parameter sequence differences of each dimension of the feature parameters of the source domain and the feature parameters of the target domain, and perform linear fitting between the logarithm of the parameter sequence differences and the logarithm of the time series; calculate the feature similarity of each dimension of the parameters based on the slope of the fitted line and the parameter sequence differences. The fluctuation amplitude of the feature parameters of the target domain within multiple consecutive time windows is statistically analyzed, and the standard deviation of the fluctuation amplitude is calculated to obtain the fluctuation stability index. The similarity value is multiplied by the fluctuation stability index and normalized to obtain the importance weight of each dimension parameter. The feature parameters of each dimension of the source domain are multiplied by their corresponding importance weights to obtain the reweighted source domain features.

5. The method according to claim 1, characterized in that, The reweighted source domain features and target domain features are used to form training samples. The current process feature parameters are used as state input, the process feature parameter adjustment amount is used as action output, and the quality inspection result is used as reward signal. The optimal parameter adjustment strategy is obtained through iterative optimization using the gradient ascent method, including: A state vector is formed by acquiring the current process feature parameters and the historical quality detection results from multiple recent moments. An action vector is formed by setting a reasonable adjustment range for each process parameter based on process experience. The action vectors are sampled to obtain an action sample set. The probability distribution of each action in the action sample set is statistically analyzed. The conditional entropy value is calculated based on the probability distribution as a regularization constraint. The difference between the target domain feature parameters and the reweighted source domain feature parameters is calculated as the feature consistency loss. The target domain features are input into a discriminant network to obtain the discrimination probability. Based on the discrimination probability, the authenticity of the feature transformation is determined to obtain the adversarial loss. The adversarial loss, feature consistency loss, and conditional entropy regularization constraint are weighted according to preset weights to obtain the total feature generation loss. The total feature generation loss guides the optimization direction of the feature generation network. The product undergoes quality testing to obtain measured values ​​of multiple quality indicators. Based on the target values ​​of each quality indicator, the deviation of the measured values ​​is calculated to obtain the target reward value. Based on the target reward value, the advantage value of each action in the current state is calculated. Based on the advantage value, the training samples are weighted and summed. The optimal parameter adjustment strategy is obtained through iterative optimization using the gradient ascent method.

6. The method according to claim 5, characterized in that, The target reward value is obtained by calculating the deviation of the measured value from the target value of each quality indicator. Calculate the advantage value of each action in the current state based on the target reward value, including: Obtain the measured values ​​and target values ​​of multiple quality indicators, and divide the absolute value of the difference between the measured value and the target value of each quality indicator by the fluctuation threshold of that quality indicator to obtain the quality deviation. Based on historical data, the cross-correlation coefficient and phase delay between quality indicators are calculated. Indicator pairs with cross-correlation coefficients higher than the correlation threshold are marked as strongly coupled indicators. The phase delay is extracted to determine the causal order between indicators. Based on the causal order, the quality indicators are divided into source indicator group and response indicator group. The quality deviation of the strongly coupled indicators is input into the sliding time window, and the mean square deviation of each indicator within the window is calculated. The indicator with the largest mean square deviation is used as the benchmark to calculate the evaluation score of the indicator group. The evaluation scores of each indicator group are weighted and fused to obtain the target reward value. By statistically analyzing the reward sequences over multiple evaluation periods, the mean target reward and fluctuation trend characteristics are extracted. The difference between the current target reward value and the mean target reward is taken as the short-term gain, and the fluctuation trend characteristics are taken as the long-term gain. The action advantage value is determined based on the weighted sum of the short-term and long-term gains.

7. A system for optimizing the cutting and welding process parameters of tabs for soft-pack batteries, used to implement the method as described in any one of claims 1-6, characterized in that, include: The first unit is used to collect cutting force data during the tab cutting process using a force measuring instrument, perform piecewise Hilbert transform on the cutting force data to obtain the intrinsic mode functions of each order, and extract instantaneous frequency and amplitude features; collect the displacement trajectory of the tab cutting tool using a displacement sensor, perform spatial decomposition on the displacement trajectory, and extract the trajectory curvature value; and weight and fuse the instantaneous frequency, amplitude features, and Gaussian curvature to obtain cutting feature parameters. The second unit is used to acquire welding point temperature data through a thermal imager and extract temperature gradient features; acquire welding current waveform through a current acquisition device and calculate the amplitude, rise time and duration of the current waveform as current features; weight and fuse the temperature gradient features and the current features to obtain welding feature parameters; and store the cutting feature parameters and the welding feature parameters into a process state feature parameter library. The third unit is used to extract the feature parameters corresponding to the historical best process state from the process state feature parameter library as the source domain; take the process feature parameters of the current process as the target domain, calculate the similarity relationship between the features of the source domain and the target domain, and weight the features of the source domain according to the similarity relationship; The reweighted source domain features and target domain features are combined to form training samples. The current process feature parameters are used as state input, the process feature parameter adjustment amount is used as action output, and the quality detection result is used as reward signal. The optimal parameter adjustment strategy is obtained by iterative optimization through gradient ascent method. When the quality inspection results of multiple consecutive products consistently meet the target requirements, the current combination of process characteristic parameters is recorded as the optimal process parameters.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.