Robot welding seam grinding control method and system based on DDPG reinforcement learning algorithm and storage medium
Through the robot weld grinding control method based on DDPG reinforcement learning algorithm, the PointNet++ network and vision sensor adapt to workpieces of different specifications are solved, and the dust hazards, difficulty in ensuring accuracy and low efficiency of traditional manual grinding are achieved, and efficient and accurate automatic grinding is achieved.
Patent Information
- Application Number
- CN202510665452.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The wheel diameter of the wheels of traditional manual grinding rollers has problems such as dust hazard, difficulty in guaranteeing accuracy, low efficiency and inability to adapt to workpieces of different specifications.
The robot weld grinding control method based on DDPG reinforcement learning algorithm is adopted, and weld features are extracted through the PointNet++ network and compared with the process database. The grinding parameters are directly called or generated. The quality is evaluated in real time with vision sensors, and the pre-trained DDPG model is used to adapt to new specifications of workpieces.
It realizes efficient and precise grinding of workpieces of different specifications, reduces manual intervention, improves work efficiency and ensures consistency of grinding quality.
Smart Images

Figure CN120480905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to robot weld grinding, and in particular to a robot weld grinding control method, system and storage medium based on a DDPG reinforcement learning algorithm. Background Art
[0002] For the grinding of straight welds on the outer rims of roller wheels with a diameter of 1100mm to 1800mm, the traditional grinding method mainly relies on manual operation. The grinding site is prone to dust particles and noise, which leads to many disadvantages. These include serious harm to the human body caused by long-term exposure to dust; the processing of curved workpieces is complex, and manual operation is difficult to maintain accuracy; workers are easily disturbed by environmental factors, making it difficult to maintain high-precision work quality for a long time, resulting in poor efficiency and difficulty in ensuring consistency in grinding quality.
[0003] Industrial robots are superior to manual labor in terms of operating accuracy and anti-interference ability. Using robots to replace manual grinding has achieved certain results, but there are still many shortcomings, mainly reflected in the lack of real-time quality feedback and parameter iteration mechanism, reliance on manual input of grinding parameters, and inability to autonomously adapt to new specifications of workpieces. Because weld characteristics are complex and diverse, fixed parameters are difficult to adapt to the grinding of workpieces of multiple specifications. Summary of the Invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a robot weld grinding control method, system and storage medium based on the DDPG reinforcement learning algorithm that can adaptively generate corresponding grinding parameters for welds of different specifications.
[0005] Technical solution: The robot weld grinding control method based on the DDPG reinforcement learning algorithm described in the present invention includes the following steps:
[0006] S1. Use the RGV transfer vehicle to transport the workpiece to be polished to the support platform and fix it;
[0007] S2. Use the visual sensor at the front end of the robotic arm to scan and confirm the positions of the arc-starting plates at both ends of the weld of the workpiece to be ground and cut them. After cutting, the cutting tool is replaced with a grinding tool;
[0008] S3, scanning the weld of the workpiece to be polished by a visual sensor to obtain point cloud data and perform denoising processing;
[0009] S4. Use the PointNet++ network to extract multiple weld features from the denoised point cloud data. Compare the extracted features with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters.
[0010] S5. The robot arm grinds the weld based on the grinding parameters. After grinding, the weld surface is scanned by the visual sensor. The robot arm determines whether to grind again based on the preset quality assessment criteria. If so, the robot arm returns to step S3; otherwise, the robot arm goes to step S6.
[0011] S6. Record the current workpiece parameters and corresponding grinding parameters into the process database;
[0012] S7. Release the workpiece fixation and transport the workpiece to the unloading rack via the RGV transfer vehicle.
[0013] By extracting features through the PointNet++ network, it can directly process disordered and unstructured weld point cloud data. Through its hierarchical feature learning and local-global feature fusion, it can accurately extract multi-dimensional geometric features such as weld width, depth, curvature, and residual height shape. Through its interpolation strategy, it can restore the local features of the missing area to ensure high discrimination and robustness of feature expression. Through its symmetric function, it can reduce the amount of calculation and improve the feature extraction speed. According to the extracted features, it is compared with the existing data in the process database. If there is weld grinding data with the same features, the corresponding grinding parameters can be directly called for rapid grinding. For weld types that do not exist in the process database, the corresponding grinding parameters can be generated by inputting the extracted features into the pre-trained DDPG reinforcement learning model. Grinding parameters. After the robot arm grinds according to the grinding parameters, it scans the weld surface through the visual sensor at its front end, and evaluates and feedbacks whether the grinding quality is qualified in real time. If it is unqualified, the above steps are repeated until the quality is qualified and the corresponding parameters of the entire grinding process are stored in the process database to provide a call basis for subsequent grinding. This method can not only store a large amount of grinding data through the process database, realize rapid call data for subsequent grinding, and improve work efficiency, but also generate new grinding parameters based on the pre-trained DDPG reinforcement learning model. It can adapt to workpieces of various specifications without manual input of response grinding parameters, has wide applicability, and also greatly improves work efficiency. Moreover, the visual sensor can provide real-time feedback on the grinding quality to ensure the final grinding effect and avoid the occurrence of unqualified grinding.
[0014] Preferably, the training process of the pre-trained DDPG reinforcement learning model in step S4 includes:
[0015] S101, collect weld point cloud data of workpieces of various specifications in various environments and perform denoising;
[0016] S102, using the PointNet++ network to extract multiple features of the weld from the denoised point cloud data as the weld status;
[0017] S103, initializing the Actor network and the Critic network, and the Actor network generates corresponding grinding actions according to the status of all welds;
[0018] S104. Execute all polishing actions in the simulation environment and calculate their rewards, and store all obtained interaction data in the experience replay buffer;
[0019] S105. Randomly sample multiple groups of interaction data from the experience replay buffer, update the critic network parameters through the backpropagation algorithm, use the Q value gradient output by the critic network to calculate the policy gradient to update the actor network parameters, and use the soft update strategy to synchronize the target network. Repeat this step until the model converges and the training is completed.
[0020] By collecting weld point cloud data of workpieces of various specifications in various environments, the applicability of the pre-trained DDPG reinforcement learning model can be improved. By using the PointNet++ network to extract features, and through its multi-level feature fusion capability, the local details of the weld are combined with the global state, providing the DDPG model with a feature vector containing the complete geometric and spatial information of the weld, enhancing the model's ability to learn the action strategies corresponding to grinding complex welds, thereby improving the generalization and decision-making accuracy of the pre-trained DDPG reinforcement learning model.
[0021] Preferably, in the step S5, the robot arm obtains the weld size change in real time through the visual sensor during the grinding process, and calculates the cumulative cutting amount based on the change. When the cumulative cutting amount reaches the threshold, the grinding tool is replaced and the cumulative cutting amount is recalculated. The calculation formula of the cumulative cutting amount is:
[0022] V=ΔW×ΔT×L
[0023] Among them, V is the cumulative cutting amount, ΔW is the change in the width of the weld, ΔT is the change in the thickness of the weld, and L is the ground length.
[0024] The cumulative cutting volume is monitored and calculated in real time through visual sensors. When the cumulative cutting volume reaches a threshold, it is determined that the tool wear is already serious. Continuing to use the tool for grinding will not only be inefficient but also have poor grinding effects. Replacing it accordingly can ensure the final grinding quality.
[0025] Preferably, the preset quality assessment standards in step S5 include a roughness standard and a concave depth standard. When both meet corresponding thresholds, re-polishing is unnecessary; otherwise, re-polishing is necessary.
[0026] Roughness and indentation depth reflect the flatness of the polished surface. Only when both meet the corresponding threshold requirements can the polished surface be considered sufficiently flat. Using these two parameters as preset quality assessment standards can not only reasonably evaluate the polishing quality, but these two parameters can also be easily obtained through visual sensors, reducing the complexity of the assessment and improving work efficiency.
[0027] The robot weld grinding control system based on the DDPG reinforcement learning algorithm described in the present invention includes:
[0028] Workpiece pre-processing module: used to transport the workpiece to be ground to the support platform via the RGV transfer vehicle and fix it; the visual sensor at the front end of the robotic arm scans and confirms the position of the arc-starting plates at both ends of the weld of the workpiece to be ground and cuts it. After cutting, the cutting tool is replaced with the grinding tool;
[0029] Data acquisition and preprocessing module: used to scan the weld seam of the workpiece to be polished through a visual sensor to obtain point cloud data and perform denoising;
[0030] Grinding parameter generation module: This module uses the PointNet++ network to extract multiple weld features from denoised point cloud data. The extracted features are then compared with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters.
[0031] Evaluation module: used to grind the weld using a robotic arm based on the grinding parameters. After grinding, the weld surface is scanned by a visual sensor and the system determines whether to grind again based on the preset quality evaluation criteria. If so, the system returns to the data acquisition and preprocessing module; otherwise, the system enters the recording module.
[0032] Recording module: used to record the current workpiece parameters and corresponding grinding parameters to the process database;
[0033] Unloading module: used to release the workpiece fixation and transport the workpiece to the unloading rack via the RGV transfer vehicle.
[0034] The computer-readable storage medium storing one or more programs according to the present invention includes one or more programs including instructions, which, when executed by a computing device, enable the computing device to perform any of the above methods.
[0035] Beneficial effects: Extracting weld features through the PointNet++ network can efficiently parse disordered point cloud data, adapt to complex weld structures, accurately capture key geometric properties of weld width, curvature, excess height, etc., generate highly discriminative and interpretable features, achieve end-to-end automation and meet real-time requirements, and input the extracted features into the pre-trained DDPG reinforcement learning model to generate the grinding parameters that can generate the weld to control the robotic arm for grinding, so that this method can adapt to workpieces of various specifications and has a wide applicability. There is no need to manually input grinding parameters, which improves work efficiency. In addition, all grinding data will be stored in the process database. After feature extraction, it can be compared with the process database first. If there are welds with the same features, the corresponding grinding parameters can be directly called. As the number of grinding increases, the process database will gradually expand and improve, which can further improve work efficiency. In addition, real-time monitoring and feedback of grinding quality through visual sensors can ensure the flatness of the final polished surface. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of the overall structure of the robot system of the present invention;
[0037] Figure 2 Schematic diagram of the three-dimensional structure of the roller frame of the present invention;
[0038] Figure 3 It is a structural schematic diagram of the pneumatic pressing device in the roller frame of the present invention;
[0039] Figure 4 Schematic diagram of the structure of the robotic arm of the present invention;
[0040] Figure 5 Schematic diagram of the replaceable tool and its bracket structure, where Figure 5 a is a schematic diagram of the grinding tool and bracket structure. Figure 5 b is a schematic diagram of the cutting tool and bracket structure;
[0041] Figure 6 It is a schematic diagram of the structure of the closed tool magazine 4;
[0042] Figure 7 This is a hierarchical structure diagram of the DDPG reinforcement learning model;
[0043] Figure 8 This is a schematic diagram of the pre-training process of the DDPG reinforcement learning model;
[0044] Figure 9 Schematic diagram of the execution flow of the present invention;
[0045] Figure 10 This is a visual recognition flow chart of the present invention;
[0046] Figure 11This is a flow chart of the process database data calling mechanism of the present invention. DETAILED DESCRIPTION
[0047] As shown in the figure, the robot weld grinding control method based on the DDPG reinforcement learning algorithm described in the present invention uses a robot system, which specifically includes a base, on which are provided an RGV transfer vehicle 1, a support platform 2, a robotic arm 3, a closed tool magazine 4, a robot control cabinet 5, an electrical control device 6 and a spiral air duct 7.
[0048] The supporting platform 2 includes a collecting box 2-1, a workbench 2-2, a supporting column 2-3, a fixing seat 2-4, a roller barrel 2-5, a worm gear reduction motor 2-6, a linear guide rail 2-7 and a pneumatic clamping device 2-8. The collecting box 2-1 is installed on the workbench 2-2, and the arc-starting plate falls into the collecting box 2-1 after being cut. The RGV transfer vehicle 1 is provided with a loading platform with a raised part for placing the workpiece; a retractable supporting column 2-3 is provided on the top of the RGV transfer vehicle 1, and a loading platform for carrying the workpiece is connected to the top of the supporting column 2-3, and the workpiece is loaded and unloaded by raising and lowering the loading platform; the fixing seat 2-4 is installed on both sides of the workbench 2-2, and a platform is provided on the top, a linear guide rail 2-7 is installed on the platform, and a roller barrel 2-5 is installed on the linear guide rail 2-7. Two roller barrels 2-5 are provided on the fixing seat 2-4 on each side, and the four roller barrels 2-5 on both sides of the workbench 2-2 cooperate to clamp Hold the workpiece; the roller drum 2-5 is moved along the linear guide rail 2-7 through the internal transmission device to achieve fine-tuning of the wheelbase (that is, adjustment of the spacing between the roller drums 2-5 on both sides). For large workpieces, the spacing is increased to prevent instability and shaking or slipping during grinding due to the high center of gravity and too close spacing; for small workpieces, the spacing is reduced and the height of the workpiece is increased to adapt to the grinding height range of the robot arm 3 (800~1850mm); the worm gear reduction motor 2-6 is connected to the roller wheel 2-5, and the rotation of the motor drives the roller drum 2-5 to rotate, and the rotation of the roller drum 2-5 can drive the straight weld of the workpiece to rotate to an upward position; the pneumatic clamping device 2-8 is installed in the middle position of the workbench 2-2. When the workpiece is positioned, it extends downward to support the inner surface of the workpiece, and the roller drum 2-5 contacts the outer surface of the workpiece to form a support. The roller drum 2-5 and the pneumatic clamping device 2-8 work together to fix the workpiece.
[0049] The grinding system includes a robotic arm 3 and a closed tool magazine 4. The robotic arm 3 is installed on the base. The robot is equipped with a visual sensor, which drives the visual sensor to scan the workpiece through joint movement. A quick-turn male joint is installed at the end position. The quick-turn male joint is connected to the quick-change female joint installed on the force-controlled sanding machine 8 (grinding tool) or the force-controlled cutting machine 9 (cutting tool) to realize fully automatic tool replacement. The force-controlled cutting machine 9 is used to cut the arc-starting plate, and the force-controlled sanding machine 8 removes burrs, grinds the end face and chamfers the cutting area, and then grinds the weld; the support frame 10 is installed near the robotic arm 3, and two support frames 10 can be set to place tools during the grinding operation, which is convenient for the robot to quickly pick up and improve the system operation efficiency; that is, when the robotic arm 3 replaces two tools, the idle tools can be placed on the support frame 10.
[0050] The closed tool magazine 4 includes a cutting disc replacement station 4-1 and an automatic sanding belt replacement station 4-2. The robot control cabinet 5 is provided with a constant force floating monitoring module and a sanding belt wear assessment module. The constant force floating monitoring module is communicated with the force position compensator installed on the robot force-controlled sanding belt machine. The constant force floating monitoring module is used to maintain a constant force state during the sanding process in real time. Under constant force conditions, the sanding belt wear assessment module uses a visual sensor to obtain the width dimension change ΔW, thickness dimension change ΔT and sanding length L of the workpiece during the sanding process, and approximately calculates the cutting amount V. The calculation formula is:
[0051] V=ΔW×ΔT×L
[0052] The system presets a cumulative cutting amount threshold V th , when V=V th When the maximum wear limit is reached, the system immediately generates a control command to suspend grinding and control the robot arm 3 to send the force-controlled sanding belt machine 8 to the automatic sanding belt replacement station 4-2 in the closed tool magazine 4 for belt replacement. The cutting volume calculation parameter is reset to zero, and the calculation is restarted when grinding is resumed until the cutting volume V reaches the threshold again. When the cutting blade of the force-controlled cutting machine 9 needs to be replaced, the robot arm 3 can also be controlled to send the force-controlled cutting machine 9 to the cutting blade replacement station 4-1 for replacement.
[0053] In addition, when grinding is not in progress, the double station can be used to store the tools, which is equipped with a sensor that can detect the approach of the robotic arm and send a signal to the induction switch to open the warehouse door.
[0054] The support platform 2 is suitable for supporting and placing workpieces of various specifications, with a diameter range of 1100-1800mm and a length range of 1200-2500mm. The electrical control device 6 establishes a communication connection with each device to automatically control and monitor the operation of the equipment. During system operation, the spiral air duct 7 collects and processes the chips, grinding chips, and dust generated during the operation of the equipment. The robot control cabinet 5 is equipped with a controller for controlling the operation of the various components of the robot system, including the movement of the various components mentioned above (the robot systems used in this invention are all existing devices, so their specific control principles are not described in detail), as well as collecting weld data and generating corresponding grinding data, which is the control method and core content of the present invention.
[0055] The control method of the present invention comprises the following steps:
[0056] S1, transport the workpiece to be polished to the support platform 2 by the RGV transfer vehicle 1 and fix it;
[0057] Specifically, the workpiece to be polished is transported to the support platform 2 by the RGV transfer vehicle 1, the support column 2-3 drives the loading platform to descend, and the workpiece falls on the roller drum 2-5. The roller drum 2-5 adjusts the wheelbase along the linear guide rail 2-7 according to the size of the workpiece under the drive of the internal transmission device, and drives the workpiece to rotate to the position where the weld is facing upward through its own rotation. The pneumatic clamping device 2-8 extends downward and presses on the inner surface of the workpiece to fix the workpiece.
[0058] S2, using the visual sensor at the front end of the robot arm 3 to scan and confirm the positions of the arc-starting plates at both ends of the weld of the workpiece to be ground and cut them, and after cutting, the cutting tool is replaced with a grinding tool;
[0059] Specifically, the grinding robot 3 scans the two ends of the workpiece weld through a visual sensor, obtains the position of the arc-starting plate, and uses the force-controlled cutting machine 9 to cut it. The arc-starting plate automatically falls into the collection box 2-1 for recycling. After the cutting of both ends is completed, the force-controlled sanding machine 8 is replaced to grind the end faces, chamfer and remove burrs.
[0060] S3, scanning the weld of the workpiece to be polished by a visual sensor to obtain point cloud data and perform denoising processing;
[0061] S4. Use the PointNet++ network to extract multiple weld features from the denoised point cloud data. Compare the extracted features with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters.
[0062] Specifically, after extracting the weld features, perform the following steps:
[0063] In step a, the system compares the acquired weld size information (including thickness, width, etc.) with the existing grinding data in the process database, first determines the interval in which the acquired weld size information is located, and then searches for the grinding data in the corresponding interval in the grinding process database;
[0064] Step b: If there is weld seam data with the same characteristics in the process database, that is, weld seams of workpieces with the same specifications have been polished before, the corresponding polishing parameters are directly retrieved from the database for this polishing operation; otherwise, proceed to step c.
[0065] Step c: If there is no weld-related data with the same characteristics in the database, that is, the weld of the workpiece of this specification has not been polished, the system calls the model, inputs the extracted weld characteristics, and generates polishing parameters suitable for this operation;
[0066] The pre-training of the DDPG reinforcement learning model includes the following steps:
[0067] S101, collecting weld point cloud data of workpieces of various specifications in various environments and performing statistical filtering and denoising;
[0068] S102. Use the PointNet++ network to extract multiple features of the weld from the denoised point cloud data as the state of the weld, the multiple features including the width, depth, height, curvature, reinforcement shape, edge slope, circumferential position, and radial position of the weld.
[0069] S103. Initialize the Actor network and the Critic network. The Actor network generates corresponding grinding actions according to the status of all welds. The grinding actions are grinding parameters, including grinding pressure, grinding time, grinding posture, spindle speed, and end TCP point.
[0070] S104. Execute all grinding actions in the simulation environment and calculate their rewards. The rewards are calculated based on the reward function designed according to the grinding effect. All interaction data obtained are stored in the experience replay buffer. The interaction data includes the current state of the weld, the grinding action, the reward corresponding to the grinding action, and the next state of the weld. The next state of the weld is derived from the current state and the grinding action.
[0071] S105. Randomly sample multiple groups of interaction data from the experience replay buffer, update the critic network parameters through the backpropagation algorithm, use the Q value gradient output by the critic network to calculate the policy gradient to update the actor network parameters, and use the soft update strategy to synchronize the target network. Repeat this step until the model converges and the training is completed.
[0072] Updating the Critic network parameters includes the following process:
[0073] The target Q value is calculated by the following formula
[0074] y=r+γ·Q target (s',μ target (s'))
[0075] Where y is the target Q value, r is the current reward, γ is the discount factor, s' is the next state, μ target is the polishing action generated by the target Actor network, Q target is the predicted value of the target Critic network;
[0076] Minimize the mean square error loss function L between the predicted Q value and the target Q value. The calculation formula of L is
[0077]
[0078] Among them, Q(s,a) is the Q value predicted by the Critic network, s is the current state of the weld, a is the grinding action generated by the Actor network based on s, and N is the number of groups of interaction data sampled from the experience replay buffer.
[0079] The policy gradient is calculated according to the following formula
[0080]
[0081] Where J is the objective function, i.e. the expected long-term cumulative reward; N is the number of groups of interaction data sampled from the experience replay buffer, The gradient of the Q value predicted by the Critic network for action a indicates the influence of a on the Q value under the current state s of the weld. a is the grinding action generated by the Actor network based on the current state s of the weld. The polishing action μ(s) generated by the Actor network for its own parameters θ μ The gradient of , which represents how adjusting the Actor parameters changes the generated action μ(s).
[0082] S5. Robot arm 3 grinds the weld based on the grinding parameters. After the grinding parameter format is converted, it is transmitted to the robot arm control system via Ethernet and TCP / IP protocol. After parsing, instructions are generated to drive each joint to move and perform grinding. During the process, the constant force floating monitoring module realizes constant force compensation, and the sanding belt wear assessment module monitors the sanding belt wear. If the sanding belt is excessively worn, the grinding is stopped and replaced.
[0083] After grinding, the weld surface is scanned by a visual sensor and judged whether to grind again according to the preset quality assessment standard. If yes, the process returns to step S3; otherwise, the process proceeds to step S6.
[0084] The preset quality assessment standards include roughness standard and concave depth standard. When both meet the corresponding thresholds, no re-grinding is required; otherwise, re-grinding is required.
[0085] The roughness is calculated by the arithmetic mean of the absolute values of the distances from each point on the surface profile to the reference line measured by the visual sensor;
[0086] The sink depth is detected by constructing a 3D model of the weld, identifying the sink area and calculating the maximum depth.
[0087] The roughness standard range can be set to 6.3μm≤R a ≤12.5μm, the concave depth standard can be set to no more than 0.4mm. Of course, other reasonable thresholds can be set for both standards based on actual conditions. If any preset quality assessment standard deviates far from its threshold and is deemed to have a serious defect, the system will immediately alarm and suspend polishing, waiting for manual intervention.
[0088] S6. Record the current workpiece parameters and corresponding grinding parameters to the process database; the workpiece parameters include relevant information such as workpiece specifications and weld characteristics.
[0089] S7. Release the workpiece fixation and transport the workpiece to the unloading rack via RGV transfer vehicle 1.
[0090] Specifically, the pneumatic clamping device 2-8 retracts to release the clamping, and the loading platform of the RGV transfer vehicle 1 is lifted, carrying the workpiece, moving along its walking guide rail, and automatically transporting the workpiece to the unloading rack. The workpiece is lifted away and the next cycle is carried out.
[0091] The robot weld grinding control system based on the DDPG reinforcement learning algorithm described in the present invention includes:
[0092] Workpiece pre-processing module: used to transport the workpiece to be ground to the support platform via the RGV transfer vehicle and fix it; the visual sensor at the front end of the robotic arm scans and confirms the position of the arc-starting plates at both ends of the weld of the workpiece to be ground and cuts it. After cutting, the cutting tool is replaced with the grinding tool;
[0093] Data acquisition and preprocessing module: used to scan the weld seam of the workpiece to be polished through a visual sensor to obtain point cloud data and perform denoising;
[0094] Grinding parameter generation module: This module uses the PointNet++ network to extract multiple weld features from denoised point cloud data. The extracted features are then compared with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters.
[0095] Evaluation module: used to grind the weld using a robotic arm based on the grinding parameters. After grinding, the weld surface is scanned by a visual sensor and the system determines whether to grind again based on the preset quality evaluation criteria. If so, the system returns to the data acquisition and preprocessing module; otherwise, the system enters the recording module.
[0096] Recording module: used to record the current workpiece parameters and corresponding grinding parameters to the process database;
[0097] Unloading module: used to release the workpiece fixation and transport the workpiece to the unloading rack via the RGV transfer vehicle.
[0098] The computer-readable storage medium storing one or more programs according to the present invention includes one or more programs including instructions, which, when executed by a computing device, enable the computing device to perform any of the above methods.
Claims
1. A robot weld grinding control method based on DDPG reinforcement learning algorithm, characterized in that: The following steps are involved: S1, transporting the workpiece to be polished to the support platform (2) by the RGV transfer vehicle (1) and fixing it; S2, using the visual sensor at the front end of the robotic arm (3) to scan and confirm the positions of the arc-starting plates at both ends of the weld of the workpiece to be ground and cutting, and replacing the cutting tool with a grinding tool after cutting; S3, scanning the weld of the workpiece to be polished by a visual sensor to obtain point cloud data and perform denoising processing; S4. Use the PointNet++ network to extract multiple weld features from the denoised point cloud data. Compare the extracted features with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters. S5, the robot arm (3) grinds the weld based on the grinding parameters, scans the weld surface with a visual sensor after grinding, and determines whether to grind again based on a preset quality assessment standard. If yes, the process returns to step S3, otherwise, the process proceeds to step S6; S6. Record the current workpiece parameters and corresponding grinding parameters into the process database; S7, release the workpiece fixation, and transport the workpiece to the unloading rack via the RGV transfer vehicle (1).
2. The method according to claim 1, wherein: The training process of the pre-trained DDPG reinforcement learning model in step S4 includes: S101, collect weld point cloud data of workpieces of various specifications in various environments and perform denoising; S102, using the PointNet++ network to extract multiple features of the weld from the denoised point cloud data as the weld status; S103, initializing the Actor network and the Critic network, and the Actor network generates corresponding grinding actions according to the status of all welds; S104. Execute all polishing actions in the simulation environment and calculate their rewards, and store all obtained interaction data in the experience replay buffer; S105. Randomly sample multiple groups of interaction data from the experience replay buffer, update the critic network parameters through the backpropagation algorithm, use the Q value gradient output by the critic network to calculate the policy gradient to update the actor network parameters, and use the soft update strategy to synchronize the target network. Repeat this step until the model converges and the training is completed.
3. The method according to claim 2, wherein: The multiple features of the weld extracted in step S102 include the width, depth, height, curvature, shape of the weld reinforcement, edge slope, circumferential position and radial position.
4. The method according to claim 2, wherein: The interactive data in step S104 includes the current state of the weld, the grinding action, the reward corresponding to the grinding action, and the next state of the weld; the next state of the weld is derived from the current state and the grinding action.
5. The method according to claim 2, wherein: Updating the critic network parameters in step S105 includes the following process: The target Q value is calculated by the following formula y=r+γ·Q target (s',μ target (s')) Where y is the target Q value, r is the current reward, γ is the discount factor, s' is the next state, μ target is the polishing action generated by the target Actor network, Q target is the predicted value of the target Critic network; Minimize the mean square error loss function L between the predicted Q value and the target Q value. The calculation formula of L is Among them, Q(s,a) is the Q value predicted by the Critic network, s is the current state of the weld, a is the grinding action generated by the Actor network based on s, and N is the number of groups of interaction data sampled from the experience replay buffer.
6. The method according to claim 2, wherein: In step S105, the policy gradient is calculated according to the following formula Where J is the objective function, N is the number of interaction data groups sampled from the experience replay buffer, is the gradient of the Q value predicted by the Critic network for action a, and a is the grinding action generated by the Actor network according to the current state s of the weld. The polishing action μ(s) generated by the Actor network for its own parameters θ μ gradient.
7. The method according to claim 1, wherein: In the step S5, the robot arm (3) obtains the weld size change in real time through the visual sensor during the grinding process, and calculates the cumulative cutting amount based on the change. When the cumulative cutting amount reaches a threshold, the grinding tool is replaced and the cumulative cutting amount is recalculated. The calculation formula of the cumulative cutting amount is: V=ΔW×ΔT×L Among them, V is the cumulative cutting amount, ΔW is the change in the width of the weld, ΔT is the change in the thickness of the weld, and L is the ground length.
8. The method according to claim 1, wherein: The preset quality assessment standards in step S5 include a roughness standard and a concave depth standard. When both meet corresponding thresholds, re-polishing is unnecessary; otherwise, re-polishing is required.
9. A robot weld grinding control system based on DDPG reinforcement learning algorithm, characterized in that: The system comprises: Workpiece pre-processing module: used to transport the workpiece to be ground to the support platform via the RGV transfer vehicle and fix it; the visual sensor at the front end of the robotic arm scans and confirms the position of the arc-starting plates at both ends of the weld of the workpiece to be ground and cuts it. After cutting, the cutting tool is replaced with the grinding tool; Data acquisition and preprocessing module: used to scan the weld seam of the workpiece to be polished through a visual sensor to obtain point cloud data and perform denoising; Grinding parameter generation module: This module uses the PointNet++ network to extract multiple weld features from denoised point cloud data. The extracted features are then compared with the data in the process database. If weld-related data with the same features exists, the corresponding grinding parameters in the process database are directly called. Otherwise, the extracted features are input into the pre-trained DDPG reinforcement learning model to generate the grinding parameters. Evaluation module: used to grind the weld using a robotic arm based on the grinding parameters. After grinding, the weld surface is scanned by a visual sensor and the system determines whether to grind again based on the preset quality evaluation criteria. If so, the system returns to the data acquisition and preprocessing module; otherwise, the system enters the recording module. Recording module: used to record the current workpiece parameters and corresponding grinding parameters to the process database; Unloading module: used to release the workpiece fixation and transport the workpiece to the unloading rack via the RGV transfer vehicle.
10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 8.
Citation Information
Patent Citations
Online visual detecting system for robot polishing
CN106584273A
Automatic grinding device and grinding method
CN111136532A
Weld joint identification method based on deep learning and 3D point cloud
CN115965960A
Robot grinding track planning method and device based on machine vision
CN118342505A
Grinding force online detection control device and control method
CN119839775A
Cited By
Reinforcement learning polishing control method and system for intelligent mechanical arm
CN121105058A
Intelligent cooperative control method and system for welding and polishing indexable robot
CN122125696A