A stainless steel cathode plate self-adaptive repair method simulating manual operation
By using machine learning and reinforcement learning algorithms based on multimodal data, an adaptive stainless steel cathode plate repair strategy is generated, which solves the problem of difficulty in quantifying repair techniques in existing technologies and achieves efficient and personalized cathode plate repair results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHIFENG YUNTONG NON FERROUS METAL CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to transform the stainless steel cathode plate repair techniques, which rely on personal experience, into an adaptive automated process, especially when faced with complex nonlinear deformations, resulting in limited repair effectiveness.
An initial repair action sequence is generated using a machine learning model trained on multimodal expert data, and the repair strategy is optimized online using a reinforcement learning algorithm, which is dynamically adjusted using 3D shape data, tapping force, and sound data.
It achieves high-precision and personalized cathode plate repair, improves the flatness and consistency of the repaired plate, and has the ability to continuously learn and self-evolve, significantly improving repair efficiency and success rate.
Smart Images

Figure CN122114050A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent repair technology for metal plates, specifically relating to an adaptive repair method for stainless steel cathode plates that simulates manual operation. Background Technology
[0002] As a key component in the electrolytic refining process, the surface flatness of stainless steel cathode plates directly affects production efficiency and product quality. During use, cathode plates are prone to complex plastic deformations such as warping and bending, requiring regular, meticulous repairs to restore their function. Currently, the core challenge in achieving efficient and high-quality repairs lies in transforming non-standardized repair techniques that rely on personal experience into a stable, automated process capable of adaptive decision-making.
[0003] Currently, cathode plate repair mainly relies on experienced workers, whose skills are difficult to quantify and pass on. Existing automation technologies attempt to replace manual labor, but each has its limitations: some methods use fixed-program CNC stamping, lacking adaptability; others introduce real-time detection and feedback control. For example, prior art document CN117816749A discloses a system for adjusting rollers based on laser scanning and a preset evaluation model. Its control logic is based on relatively simple mathematical rules, and its effectiveness is limited when faced with complex nonlinear deformations. Similarly, prior art document CN1597166B involves monitoring the strip state through sensors and adjusting the correction load according to an algorithm, and its adjustment logic is also relatively fixed. Furthermore, in the broader field of intelligent manufacturing, prior art document US11676007B2 demonstrates a method for identifying and removing surface defects of objects using an image-to-image machine learning algorithm. Its innovation lies in distinguishing between natural deformations during manufacturing and defects that need to be removed through algorithm training. However, this technology focuses on eliminating "protrusion" type defects by removing material through milling and other methods. Its decision-making results are used to modify the surface model for directly driving the processing path. It does not address or solve the key problem of how to generate and optimize a series of dynamic physical repair actions that can actively release and balance internal stress in processes such as cathode plate leveling.
[0004] Therefore, a new method for repairing stainless steel cathode plates is needed that can deeply integrate expert experience and knowledge with real-time adaptive decision-making capabilities. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention provides an adaptive repair method for stainless steel cathode plates that simulates manual operation, comprising the following steps:
[0006] S1. Obtain the current three-dimensional topography data of the cathode plate to be repaired;
[0007] S2. Input the current three-dimensional topography data into the first machine learning model to generate an initial repair action sequence, wherein the first machine learning model is trained based on pre-collected expert repair operation data;
[0008] S3. Control the repair execution unit to execute the repair actions in the initial repair action sequence;
[0009] S4. During the repair process, the intermediate three-dimensional shape of the cathode plate is obtained online as an intermediate state. Based on the difference between the intermediate state and the preset target state, the subsequent repair actions to be executed are dynamically adjusted using a reinforcement learning algorithm.
[0010] Furthermore, the expert repair operation data is multimodal data, including first three-dimensional morphological data before the expert repairs the cathode plate, operation sequence data during the repair process, and second three-dimensional morphological data after repair; the operation sequence data includes three-dimensional position data, tapping force data, and sound data for each repair action.
[0011] Furthermore, the processed sound data is used to extract sound spectrum feature data, which includes at least one of the following: dominant frequency, peak energy spectral density, and bandwidth features.
[0012] Furthermore, the first machine learning model includes a convolutional neural network and a long short-term memory network; the convolutional neural network is used to extract spatial features from the three-dimensional topography data, and the long short-term memory network is used to generate the initial repair action sequence with temporal relationships based on the spatial features.
[0013] Further, in step S4, the step of dynamically adjusting the subsequent repair actions based on the difference between the intermediate state and the preset target state using a reinforcement learning algorithm specifically includes: converting the flatness difference between the intermediate state and the preset target state into a reward signal; using the reward signal to optimize the policy network of the reinforcement learning algorithm, and outputting the adjustment amount for the three-dimensional position and / or tapping force in the subsequent repair actions.
[0014] Furthermore, the reinforcement learning algorithm employs a proximal policy optimization algorithm.
[0015] Furthermore, the method also includes a learning evolution step S5, which specifically involves: adding successfully repaired case data to the training database, wherein the case data includes at least the three-dimensional morphological data before repair, the complete sequence of actual repair actions, and the three-dimensional morphological data after repair; and retraining the first machine learning model using the updated training database.
[0016] Furthermore, steps S3 and S4 are executed in a phased iterative manner: after each phase consisting of N repair actions is completed, step S4 is triggered to obtain the current intermediate state and adjust the repair actions of the next phase, where N is an integer greater than 1.
[0017] Furthermore, the repair process terminates when one of the following conditions is met: the overall flatness of the cathode plate is less than 0.5 mm; or, the improvement in flatness is less than 0.05 mm for M consecutive stages, where M is an integer greater than 1.
[0018] This invention also provides an adaptive repair system for stainless steel cathode plates that simulates manual operation, used to implement the aforementioned method, comprising: a data acquisition unit for acquiring current three-dimensional morphological data of the cathode plate to be repaired and intermediate three-dimensional morphological data during the repair process; a repair execution unit for executing repair actions; and a processing unit communicatively connected to the data acquisition unit and the repair execution unit; wherein the processing unit is configured to: call a first machine learning model to process the current three-dimensional morphological data to generate the initial repair action sequence; control the repair execution unit to execute according to the sequence; and during the repair process, based on the difference between the intermediate three-dimensional morphological data acquired by the data acquisition unit and the preset target state, drive a reinforcement learning algorithm module to run, so as to dynamically adjust the control instructions on the repair execution unit.
[0019] Furthermore, the data acquisition unit includes: a 3D scanner for acquiring 3D topographic data; a force sensor for acquiring the striking force data of the end tool of the repair execution unit; a positioning sensor for acquiring the 3D position data of the end tool; and a microphone for acquiring the sound data generated by the striking.
[0020] Furthermore, the system also includes a knowledge database; the processing unit is further configured to store successfully repaired case data into the knowledge database and periodically update the first machine learning model using data from the knowledge database.
[0021] Compared with existing technologies, the adaptive repair method and system for stainless steel cathode plates that simulates manual operation provided by the present invention has the following advantages:
[0022] (1) Achieving the digital inheritance and transcendence of expert restoration skills: Through the first machine learning model trained on multimodal expert data, the system can effectively model and solidify the non-standardized restoration decision-making logic that relies on personal experience, generating a reliable initial restoration plan and solving the problem of the difficulty in inheriting human skills. Subsequently, by introducing an online optimization mechanism based on state differences through reinforcement learning, the system can autonomously explore and dynamically adjust its strategy during execution, possessing the potential to discover restoration paths superior to initial expert experience, thus achieving a leap from "imitation" to "optimization".
[0023] (2) Achieving high-precision and highly adaptive personalized repair: This invention does not execute a fixed program, but generates a customized initial action sequence based on the unique three-dimensional morphology of each cathode plate. More importantly, during the repair process, online decision-making and optimization can be performed based on the intermediate state feedback in real time, thereby effectively dealing with complex and ever-changing warping deformation, realizing a "one plate, one policy" fine repair for each plate, and significantly improving the flatness, consistency and success rate after repair.
[0024] (3) Possesses the ability to continuously learn and self-evolve: By feeding back successful repair case data to the training database and using it for model retraining, the system constructs a closed-loop knowledge accumulation and performance optimization cycle. This enables the system's repair strategy to continuously iterate and evolve with the increase of practical data, thereby achieving long-term performance improvement of "becoming smarter with use" and enhancing the system's practical value and life cycle. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the adaptive repair method for stainless steel cathode plates according to the present invention. Detailed Implementation
[0026] The present invention will be further described in detail through the following embodiments. These embodiments are illustrative examples of specific implementations of the present invention and are not intended to limit the scope of protection of the present invention.
[0027] Example 1
[0028] This embodiment details the complete implementation process of an adaptive repair method for stainless steel cathode plates based on CNN-LSTM and PPO reinforcement learning.
[0029] First, expert experience data was collected and the model was trained. A large amount of expert restoration case data was collected using a high-precision 3D line laser scanner (e.g., Keyence LJ-X8000 series). For each cathode plate to be restored, scanning was performed before and after restoration to obtain high-resolution point cloud data, which was then processed into a 256x256 pixel depth map as the 3D topographic data before and after restoration. While the expert was performing the restoration operation using specialized tools (e.g., a pneumatic hammer), multimodal operation sequence data was simultaneously recorded: the 3D position coordinates (X, Y, Z) of the tool tip were recorded at each strike using a motion capture system (e.g., VICON); the striking force (F) was recorded using a six-dimensional force sensor (e.g., ATI Omega160) mounted on the tool; and the sound waveform data generated by the strikes was collected using microphones placed in the work area. These data together constitute a paired training sample of "pre-restoration topography – operation sequence – post-restoration topography".
[0030] Using the training samples described above, a "Digital Craftsman Model" was trained as the first machine learning model. This model was built using the PyTorch framework, and its structure included a ResNet-50 convolutional neural network (CNN) as the encoder to extract abstract spatial topography features from the input pre-repair depth map. This feature vector was then fed into a Long Short-Term Memory (LSTM) decoder consisting of two layers with 512 hidden units each. The LSTM decoder generated a predicted sequence of repair actions in an autoregressive manner based on the input feature sequence. Each action was represented by a vector {x, y, z, F}, corresponding to the suggested tapping point coordinates and force. The model was trained using the Adam optimizer, with the loss function being the mean squared error between the predicted action sequence and the actual expert's action sequence.
[0031] When it is necessary to repair a new cathode plate with a central bulge deformation, the following steps should be performed:
[0032] S1. Use the same 3D scanner to acquire the current three-dimensional topographic data of the cathode plate to be repaired and process it into a depth map.
[0033] S2, input this depth map into the pre-trained "digital artisan model". The model outputs an initial repair action sequence R0 containing approximately 60 repair actions.
[0034] S3 controls a KUKA KR 16 six-axis industrial robot (as a repair execution unit) to execute the initial sequence. The robot's end effector is a force-controlled pneumatic stamping tool.
[0035] S4, the repair process adopts a phased iterative approach. The initial sequence R0 is divided into 6 phases, each containing 10 consecutive actions.
[0036] After each stage is completed, an online 3D scan is immediately triggered to obtain the current intermediate 3D topography of the cathode plate, and its overall flatness index P_curr (defined as the average absolute deviation of all points on the plate surface from the best-fitting plane) is calculated.
[0037] A reinforcement learning module based on the Proximal Policy Optimization (PPO) algorithm then intervenes. The state of this module is jointly composed of the depth map of the current intermediate topography and the remaining unexecuted action sequence encoded in R0. Its action space is defined as the fine-tuning amount for the parameters of the next action to be executed, such as {Δx, Δy, ΔF}. The reward function is designed as: R = (P_prev - P_curr)×10, where P_prev is the flatness at the end of the previous stage. If the flatness is improved (P_curr < P_prev), a positive reward is obtained, otherwise a negative reward is obtained.
[0038] The PPO algorithm optimizes its policy network based on the current state and the reward history, and outputs specific adjustment instructions for the first action in the next stage of R0. The robot then executes this dynamically adjusted action. This "execution - evaluation - adjustment" cycle is repeated after each stage.
[0039] The termination condition of the repair process is set as: when the overall flatness P_curr is lower than 0.5 mm, or the improvement amount of flatness (P_prev - P_curr) is less than 0.05 mm for three consecutive stages, the repair is stopped.
[0040] In addition, the system includes a learning and evolution mechanism. After each successful repair, the system automatically adds the complete data of this case - including the depth map before repair, the complete action sequence actually executed (optimized and adjusted by PPO), and the depth map after repair - as a new sample to the training database. The system regularly (e.g., weekly) retrains or incrementally trains all training samples using the updated database, so as to continuously optimize and improve the performance of the first machine learning model.
[0041] Embodiment 2
[0042] This embodiment is optimized based on Embodiment 1. By integrating acoustic features to simulate the experience of experts in "identifying shapes by listening to sounds", the accuracy and efficiency of reinforcement learning decision-making are improved.
[0043] The main difference between this embodiment and Embodiment 1 lies in the way of constructing the state of the reinforcement learning module. During the repair execution process, when the force control tool completes a knocking action, the system not only obtains 3D topography data, but also collects the instantaneous sound signal generated by the knocking through a high-sensitivity microphone.
[0044] After preprocessing, the audio signal undergoes a Fast Fourier Transform (FFT) to convert it from the time domain to the frequency domain. Subsequently, three key features are extracted from the spectrum: Peak Frequency (the frequency component with the highest energy), Peak Spectral Density (Peak PSD), and -20dB bandwidth (the frequency range covered when the signal power drops by 20 dB). These three features are then normalized.
[0045] When constructing the state vector of the PPO algorithm, in addition to including the current plate depth map features and the remaining action sequence encoding as described in Example 1, the normalized acoustic feature vector is also concatenated with it to form a more informative composite state vector. This enhanced state vector can provide the reinforcement learning agent with supplementary information about the local material response and stress state at the impact point, because different internal stress conditions often lead to regular changes in the spectral characteristics of the impact sound.
[0046] Experiments show that, with this state enhancement approach, the PPO agent converges faster when learning repair strategies. When repairing cathode plates with similar initial conditions to those in Example 1, the total number of taps required to achieve the same or better final flatness (e.g., below 0.5 mm) is reduced by an average of approximately 15%, significantly improving repair efficiency. This demonstrates the effectiveness of multimodal information fusion in enhancing the performance of adaptive repair systems.
[0047] Example 3
[0048] This embodiment demonstrates the applicability of the method of the present invention for repairing complex, atypical damage morphologies (such as severely warped edges).
[0049] One corner of the cathode plate to be repaired was severely warped, with a local warping height of up to 15mm. Following steps S1 to S3 as described in Example 1, the "digital craftsman model," based on its training experience, generated an initial repair action sequence R0 that focused on high-intensity tapping of the most warped corner area.
[0050] However, after performing the first few actions of the initial sequence, online scanning revealed that although the local warp height was reduced, new stress concentrations were induced in its neighboring areas, causing the overall flatness index P_curr to rebound, and the reinforcement learning module thus received a negative reward.
[0051] This negative feedback triggered the exploration mechanism of the PPO algorithm. In subsequent decisions, the algorithm no longer limited itself to minor modifications to the R0 sequence, but explored an action that deviated significantly from the original sequence: performing a moderate-force tap on a seemingly flat area about 300mm away from the severely warped region. This "abnormal" operation successfully released some of the macroscopic stress transmitted from the edge warping, resulting in a significant improvement in the morphology of the entire warped region during subsequent scans, thus yielding a substantial positive reward.
[0052] Subsequently, the reinforcement learning policy network incorporated this effective experience into its decision weights. The entire repair process evolved into an effective combination of initial expert experience guidance and online exploration and discovery through reinforcement learning. Ultimately, after approximately 72 taps, the flatness of the complex warped plate was successfully restored to 0.65 mm, meeting the usage requirements. This embodiment demonstrates that the "imitation learning + reinforcement learning" architecture possesses a powerful adaptive capability in handling unseen cases and complex nonlinear problems.
[0053] The above description is a preferred embodiment of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for adaptive repair of stainless steel cathode plates simulating manual operation, characterized in that, Includes the following steps: S1. Obtain the current three-dimensional topography data of the cathode plate to be repaired; S2. Input the current three-dimensional topography data into the first machine learning model to generate an initial repair action sequence, wherein the first machine learning model is trained based on pre-collected expert repair operation data; S3. Control the repair execution unit to execute the repair actions in the initial repair action sequence; S4. During the repair process, the intermediate three-dimensional shape of the cathode plate is obtained online as an intermediate state. Based on the difference between the intermediate state and the preset target state, the subsequent repair actions to be executed are dynamically adjusted using a reinforcement learning algorithm.
2. The method according to claim 1, characterized in that, The expert repair operation data is multimodal data, including the first three-dimensional morphological data before the expert repairs the cathode plate, the operation sequence data during the repair process, and the second three-dimensional morphological data after the repair; the operation sequence data includes the three-dimensional position data, tapping force data, and sound data of each repair action.
3. The method according to claim 2, characterized in that, The processed sound data is used to extract sound spectrum feature data, which includes at least one of the following: dominant frequency, peak energy spectral density, and bandwidth features.
4. The method according to claim 1, characterized in that, The first machine learning model includes a convolutional neural network and a long short-term memory network; the convolutional neural network is used to extract spatial features from the three-dimensional topography data, and the long short-term memory network is used to generate the initial repair action sequence with temporal relationship based on the spatial features.
5. The method according to claim 1, characterized in that, In step S4, the step of dynamically adjusting the subsequent repair actions based on the difference between the intermediate state and the preset target state using a reinforcement learning algorithm specifically includes: converting the flatness difference between the intermediate state and the preset target state into a reward signal; using the reward signal to optimize the policy network of the reinforcement learning algorithm, and outputting the adjustment amount for the three-dimensional position and / or tapping force in the subsequent repair actions.
6. The method according to claim 5, characterized in that, The reinforcement learning algorithm employs a proximal policy optimization algorithm.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes a learning evolution step S5, which specifically involves: adding successfully repaired case data to the training database, wherein the case data includes at least the three-dimensional morphological data before repair, the complete sequence of actual repair actions, and the three-dimensional morphological data after repair; and retraining the first machine learning model using the updated training database.
8. The method according to any one of claims 1 to 6, characterized in that, Steps S3 and S4 are executed in a phased iterative manner: after each phase consisting of N repair actions is completed, step S4 is triggered to obtain the current intermediate state and adjust the repair actions of the next phase, where N is an integer greater than 1.
9. The method according to claim 8, characterized in that, The repair process terminates when one of the following conditions is met: (a) The overall flatness of the cathode plate is less than 0.5 mm; or, (b) Within M consecutive stages, the improvement in flatness is less than 0.05 mm, where M is an integer greater than 1.
10. An adaptive repair system for stainless steel cathode plates that simulates manual operation, used to implement the method according to any one of claims 1 to 9, characterized in that, include: The data acquisition unit is used to acquire the current three-dimensional morphological data of the cathode plate to be repaired, as well as the intermediate three-dimensional morphological data during the repair process. The repair execution unit is used to perform repair actions; The processing unit is communicatively connected to the data acquisition unit and the repair execution unit; The processing unit is configured to: invoke the first machine learning model to process the current 3D shape data and generate the initial repair action sequence; and control the repair execution unit to execute the sequence. During the repair process, based on the difference between the intermediate three-dimensional topography data acquired by the data acquisition unit and the preset target state, the reinforcement learning algorithm module is driven to run, so as to dynamically adjust the control instructions on the repair execution unit.
11. The system according to claim 10, characterized in that, The data acquisition unit includes: A 3D scanner is used to acquire 3D topographic data. A force sensor is used to collect data on the striking force of the end tool of the repair execution unit; A positioning sensor is used to collect three-dimensional position data of the end effector. A microphone is used to collect sound data produced by striking.
12. The system according to claim 10 or 11, characterized in that, The system also includes a knowledge database; the processing unit is further configured to store successfully repaired case data in the knowledge database and periodically update the first machine learning model using data from the knowledge database.
Citation Information
Patent Citations
Control system of stainless steel band curved surface flatness adjusting device
CN117816749A
Methods and apparatus for monitoring and conditioning strip material
CN1597166B
Defect removal from manufactured objects having morphed surfaces
US11676007B2