Two-point incremental forming manufacturing method and device based on deep reinforcement learning

The two-point incremental forming method employs deep reinforcement learning to dynamically adjust the support policy, addressing the limitations of traditional methods by improving flexibility and accuracy in forming complex shapes.

JP7766209B2Active Publication Date: 2025-11-07HANKAISI INTELLIGENT TECH CO LTD GUIZHOU
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024561985
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-04-19
Filing Date
2023-03-22
Publication Date
2025-11-07
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

Traditional two-point incremental forming methods suffer from fixed support policies of the secondary tool head, low geometric accuracy, and limited forming range, hindering industrial application and optimization of forming control.

Method used

A two-point incremental forming method utilizing deep reinforcement learning to dynamically adjust the support policy of the slave robot by training a deep reinforcement learning model with a digital simulation environment, optimizing the support path based on the deviation between the forming surface and the target surface.

Benefits of technology

Enhances flexibility and improves forming accuracy by dynamically adjusting the support policy, allowing for more precise and complex shapes to be formed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766209000001
    Figure 0007766209000001
  • Figure 0007766209000002
    Figure 0007766209000002
  • Figure 0007766209000003
    Figure 0007766209000003
Patent Text Reader

Abstract

The present invention provides a two-point incremental forming manufacturing method and device based on deep reinforcement learning, the method includes the steps of obtaining a three-dimensional model to be manufactured, dividing it into multiple layers, obtaining multiple main working paths and multiple candidate support paths, and selecting an initial current main working path and a current support path; respectively cyclically controlling the robot arms of the master robot and the slave robot to perform incremental forming in a real application environment according to the current main working path and the selected current support path, and obtaining a forming surface; the deviation value between the forming surface and the target surface is taken as a state vector, and applied to a pre-trained deep reinforcement learning model to perform reinforcement learning of the support policy, cyclically outputting the support path corresponding to the next main working path, and cyclically updating the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed. The present invention can adjust the support policy of the slave robot, which has high flexibility and high forming accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to additive manufacturing technology, and more particularly to a two-point incremental forming manufacturing method and apparatus based on deep reinforcement learning. [Background technology]

[0002] Incremental forming is a flexible manufacturing technology that can process a desired product without using a dedicated mold. In incremental forming, a single hemispherical tool attached to a robot arm or NC machine tool moves along a preprogrammed path to locally plastically deform a metal sheet into the desired shell shape. Two-point incremental forming uses two hemispherical forming tools to achieve local incremental deformation of the material to obtain the final formed product. The principle of two-point incremental forming is that the plate is processed on one side by a main tool head (forming pressure head), while the plate is supported on the other side by a secondary tool head (support pressure head), whose movement trajectory follows that of the main tool head. This two-point incremental forming method can further improve the formability of plates and effectively improve the dimensional accuracy of formed products. However, the traditional two-point incremental forming method has many drawbacks, such as the fixed and unchangeable support policy of the secondary tool head during incremental forming, the low geometric accuracy of the forming result, and the small forming range, which limits its industrial application.The traditional method for improving incremental forming accuracy mainly involves measuring and compensating for the material rebound, which makes it difficult to optimize forming control. Summary of the Invention

[0003] The present invention provides a two-point incremental forming manufacturing method and apparatus based on deep reinforcement learning to solve the traditional problem of the slave robot's fixed support policy and low forming accuracy.

[0004] Based on the above objectives, an embodiment of the present invention provides a two-point incremental forming manufacturing method based on deep reinforcement learning. The two-point incremental forming manufacturing method based on deep reinforcement learning includes the steps of obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into multiple layers to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, and selecting one main working path and one support path corresponding to the selected main working path as an initial current main working path and a current support path based on the forming direction; cyclically controlling the robot arms of the master robot and the slave robot to perform incremental forming in an actual application environment based on the current main working path and the selected current support path, thereby obtaining a forming surface corresponding to the current main working path; using the deviation value between the forming surface and the target surface as a state vector, applying it to a pre-trained deep reinforcement learning model to perform reinforcement learning of a support policy, and cyclically outputting a support path corresponding to the next main working path; and cyclically updating the current main working path and the current support path based on the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.

[0005] Preferably, the step of obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into a plurality of layers, and obtaining a plurality of main working paths and a plurality of candidate support paths corresponding to each of the main working paths includes the steps of obtaining a three-dimensional model to be manufactured, and dividing the three-dimensional model into a plurality of layers along the molding direction with a predetermined layer thickness by applying an offset function on the surface to obtain a first predetermined number of curved paths; distributing a second predetermined number of discrete points at a predetermined point interval for each of the curved paths, and generating a main working path corresponding to the curved path based on the discrete points; and obtaining, for each of the main working paths, a plurality of candidate support paths corresponding to the main working path based on a respective plurality of support policies, wherein the support policy is one of a global support policy, a local peripheral support policy, a local frontal support policy, and a following support policy.

[0006] Preferably, before the step of having a robot arm perform incremental forming in an actual application environment based on a current main operating path and a selected current support path, and obtaining a forming surface corresponding to the current main operating path, the method includes the steps of: constructing a digital simulation environment in Grasshopper that matches the actual application environment of the 3D model to be manufactured; simulating the 3D model in the digital simulation environment, training the deep reinforcement learning model based on the simulation results, and obtaining the pre-trained deep reinforcement learning model.

[0007] Preferably, the step of simulating the three-dimensional model in the digital simulation environment, training the deep reinforcement learning model based on the simulation result, and obtaining the pre-trained deep reinforcement learning model includes the steps of selecting an initial main action path as a current simulation main action path based on a forming direction, and randomly selecting one of a plurality of candidate support paths as an initial current simulation support path based on the current simulation main action path; and applying the current simulation main action path and the current simulation support path to the digital simulation environment to perform simulation forming, and generating a simulation forming surface and a simulation forming surface spring corresponding to the current simulation main action path. The method includes the steps of: acquiring a springback value; inputting the deviation value between the simulation forming surface and the target surface as a state vector into the deep reinforcement learning model, performing reinforcement learning of the support policy, and updating the simulation support path corresponding to the next simulation main operating path and the current reward value based on the springback value of the simulation forming surface; cyclically updating the current simulation main operating path and the current simulation support path based on the next simulation main operating path and the corresponding simulation support path, respectively; cyclically controlling the robot arm to perform incremental forming based on the updated current simulation main operating path and the current simulation support path, and cyclically updating the simulation forming surface.

[0008] Preferably, the step of applying the digital simulation environment to perform simulation forming based on the current simulation main motion path and the current simulation support path, and obtaining a simulation forming surface and a simulation forming surface springback value corresponding to the current simulation main motion path, includes the steps of converting the coordinates and directions of discrete points of the current simulation main motion path and the current simulation support path into robot motion commands according to robot grammar rules; and constructing a simulation model of plate deformation using simulation software, performing simulation forming based on the robot motion commands, and returning a simulation forming surface and a simulation forming surface springback value corresponding to the current simulation main motion path.

[0009] Preferably, the step of inputting the deviation value between the simulation forming surface and the target surface as a state vector into the deep reinforcement learning model, performing reinforcement learning of a support policy, and updating the simulation support path corresponding to the next simulation main operating path and the current reward value based on the simulation forming surface springback value includes the steps of obtaining each second reference point on the simulation forming surface corresponding to each first reference point on the target surface, calculating the error value between each second reference point and each corresponding first reference point, and constructing the state vector; inputting the state vector into the deep reinforcement learning model, performing reinforcement learning of a support policy, and outputting the simulation support path corresponding to the next simulation main operating path; and updating the current reward value based on the simulation forming surface and the simulation forming surface springback value.

[0010] Preferably, the step of updating the current reward value based on the simulated forming surface and the simulated forming surface springback value includes the steps of: an initial value of the current reward value is 0, and if the springback value of the simulated forming surface is equal to or greater than a reference value, controlling the current reward value to be decreased by a first preset value; if the springback value of the simulated forming surface is less than the reference value, controlling the current reward value to be increased by a first preset value; and if forming of the simulated forming surface fails, controlling the current reward value to be a second preset value.

[0011] Based on the same inventive idea, an embodiment of the present invention further provides a two-point incremental forming manufacturing device based on deep reinforcement learning. The two-point incremental molding manufacturing device based on deep reinforcement learning includes: a path acquisition unit that acquires a three-dimensional model to be manufactured, divides the three-dimensional model into multiple layers, acquires multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, and selects an initial current main working path and one support path corresponding to the initial main working path as an initial current main working path and a current support path based on the forming direction; an incremental forming unit that cyclically controls the robot arms of the master robot and the slave robot based on the current main working path and the current support path in an actual application environment to perform incremental forming, and acquires a forming surface corresponding to the current main working path; and a reinforcement learning unit that uses the deviation value between the forming surface and the target surface as a state vector, applies it to a pre-trained deep reinforcement learning model to perform reinforcement learning of a support policy, and cyclically outputs a support path corresponding to the next main working path, and cyclically updates the current main working path and the current support path based on the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.

[0012] Based on the same inventive idea, an embodiment of the present invention further provides an electronic device, the electronic device including: a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the deep reinforcement learning-based two-point incremental forming manufacturing method is realized when the program is executed by the processor.

[0013] Based on the same inventive idea, an embodiment of the present invention further provides a computer storage medium having at least one executable command stored therein, the executable command being executed by a processor to perform the deep reinforcement learning-based two-point incremental forming manufacturing method.

[0014] As can be seen from the above, in the two-point incremental forming manufacturing method and apparatus based on deep reinforcement learning provided by the embodiments of the present invention, the two-point incremental forming manufacturing method based on deep reinforcement learning includes the steps of obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into multiple layers, obtaining multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, selecting an initial main working path and one support path corresponding to the initial main working path as an initial current main working path and a current support path based on the forming direction, and moving the robot arms of the master robot and the slave robot in an actual application environment based on the current main working path and the selected current support path. The method includes the steps of performing incremental forming by cyclically controlling each of the three-dimensional motion paths to obtain a forming surface corresponding to the current main operating path, and using the deviation value between the forming surface and the target surface as a state vector and applying it to a pre-trained deep reinforcement learning model to perform reinforcement learning of the support policy, cyclically outputting a support path corresponding to the next main operating path, and cyclically updating the current main operating path and the current support path based on the next main operating path and the support path corresponding to the next main operating path until the incremental forming of the three-dimensional model is completed, thereby achieving the effects of the present invention in that the support policy of the slave robot can be adjusted, thereby increasing flexibility and improving forming accuracy. [Brief explanation of the drawings]

[0015] In order to more clearly describe the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings necessary for describing the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without any creative efforts.

[0016] [Figure 1] FIG. 1 is a flowchart of a two-point incremental forming manufacturing method based on deep reinforcement learning in an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram of a global support policy in an embodiment of the present invention. [Figure 3] FIG. 3 is a schematic diagram of a local frontal support policy in an embodiment of the present invention. [Figure 4] FIG. 4 is a schematic diagram of a follow-up support policy in an embodiment of the present invention. [Figure 5] FIG. 5 is a schematic diagram of a local neighborhood support policy in an embodiment of the present invention. [Figure 6] FIG. 6 is a schematic diagram of two-point incremental forming in an embodiment of the present invention. [Figure 7] FIG. 7 is a schematic diagram of molding in an embodiment of the present invention. [Figure 8] FIG. 8 is a schematic diagram of a two-point incremental forming manufacturing apparatus based on deep reinforcement learning in an embodiment of the present invention. [Figure 9] FIG. 9 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] To make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be described in more detail below with reference to specific embodiments and with reference to the drawings.

[0018] Unless otherwise defined, technical or scientific terms used in the embodiments of the present invention have the common meanings understood by those skilled in the art to which the present disclosure belongs. The terms "first," "second," and similar terms used in the embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish between different components. Words similar to "have" or "include" mean that the component or object appearing before the word covers the component or object exemplified after the word and its equivalents, and do not exclude other components or objects. Terms similar to "connect" or "couple" are not limited to physical or mechanical connections, but may also include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" indicate relative positional relationships. If the absolute position of the object being described is changed, the relative positional relationships will also be changed accordingly.

[0019] An embodiment of the present invention provides a two-point incremental forming manufacturing method based on deep reinforcement learning. As shown in Figure 1, the two-point incremental forming manufacturing method based on deep reinforcement learning includes the following steps:

[0020] In step S11, a three-dimensional model to be manufactured is obtained, and the three-dimensional model is divided into multiple layers to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths. Based on the molding direction, an initial main working path and one support path corresponding to the initial main working path are selected as an initial current main working path and a current support path.

[0021] In an embodiment of the present invention, a 3D model to be manufactured is loaded into a Grasshopper programming environment based on Rhino. In step S11, preferably, the 3D model to be manufactured is first obtained, and an offset on surface function is applied to divide the 3D model into multiple layers along the molding direction with a predetermined layer thickness to obtain a first predetermined number of curved paths. Here, the predetermined layer thickness is determined based on the dimensions of the forming tool head of the master robot and is typically within the range of 1.5-2 mm. The molding direction may be along the X-axis direction or another direction, and is not limited thereto. The first predetermined number is determined by the predetermined layer thickness and the size of the 3D model. The smaller the predetermined layer thickness, the larger the 3D model and the larger the first predetermined number. The curved path is a two-dimensional curved path.

[0022] Then, a second predetermined number of discrete points are allocated to each curved path at a predetermined point interval, and a main working path corresponding to the curved path is generated based on the discrete points. The predetermined point interval can be set as needed and is usually within the range of 1-3 mm. The second predetermined number is determined by the predetermined point interval and the size of the curved path. A series of planes are established along the tangent direction of the curved surface of each discrete point, and the center point of the plane is the movement position of the forming tool head, and the Z axis of the plane is the forming direction of the forming tool head. Each discrete point becomes a forming pass point, and the combination of the movement position and forming direction of the forming tool head corresponding to each discrete point on any curved path constitutes one main working path.

[0023] Finally, for each of the main working paths, obtain a plurality of candidate supporting paths corresponding to the main working path based on a respective plurality of supporting policies, where the supporting policy is one of a global supporting policy, a local peripheral supporting policy, a local frontal supporting policy, and a following supporting policy.

[0024] Based on the global support policy, a global support curve is generated by offsetting the outer contour of the forming area by a first preset distance, and the forming pass points are mapped to the global support curve to generate a global support path. The first preset distance can be set as needed, and is preferably equal to the radius of the forming tool head. As shown in Figure 2, the slave robot moves the support tool head along the part boundary.

[0025] Based on the local front support policy, the main working path of the forming area is mirrored with respect to the plane of the metal plate, and the direction of the forming pass point is reversed with respect to the plane to generate the local front support path, as shown in Figure 3, that is, the support tool head of the slave robot directly faces the front of the forming tool head of the master robot.

[0026] Based on the tracking support policy, the main working path of the forming section is mirrored relative to the plane of the metal sheet, the forming pass points are reversed in direction relative to the plane, and the forming pass point list is offset three terms forward to generate the tracking support path. As shown in Figure 4, in the case of local front support, the support tool head of the slave robot lags behind the forming tool head of the master robot by a second preset distance. The second preset distance can be set as needed and is preferably equal to the diameter of the forming tool head.

[0027] Based on the local peripheral support policy, the following support path is offset one layer backward to generate a local peripheral support path, as shown in Figure 5. That is, the support tool head of the slave robot directly follows the corresponding offset path of the forming tool head of the master robot, and one forming gap is formed between the forming tools.

[0028] Each of the four support policies has its own merits and demerits. Local support policies (including the local frontal support policy and the local peripheral support policy) improve the shaping of local details but result in a large overall error. Global support policies can unify the overall error but result in a large error in shaping details. Following support policies can improve the quality of the shaped surface. In the subsequent reinforcement learning environment, the slave robot's behavior space is to implement the four support policies and select one of the four candidate support paths mentioned above.

[0029] An initial current main working path is selected from the plurality of main working paths based on a forming direction, and an initial current support path is selected from four candidate support paths corresponding to the initial main working path.

[0030] In step S12, based on the current main working path and the selected current support path, the robot arms of the master robot and the slave robot are respectively circulated and controlled in the actual application environment to perform incremental forming, and a forming surface corresponding to the current main working path is obtained.

[0031] In an embodiment of the present invention, the action space of the deep reinforcement learning model is the support policy of the slave robot. The output of the deep reinforcement learning model includes four dimensions, each corresponding to one support path of the slave robot, which are a global support policy, a local front support policy, a local surrounding support policy, and a following support policy, respectively. For example, the deep reinforcement learning model includes one input layer, three hidden layers, and one output layer. The state vector input by the input layer includes multiple parameters, for example, 16 parameters S: {d1; d2; d3... d15; d16}. The three hidden layers have 64 neurons, 32 neurons, and 16 neurons, respectively. The output layer outputs four parameters A: {g; l1; l2; f}, where g indicates the probability that the support path corresponding to the next main working path is a global support path (global), l1 indicates the probability that the support path corresponding to the next main working path is a local frontal support path (local1), l2 indicates the probability that the support path corresponding to the next main working path is a local peripheral support path (local2), and f indicates the probability that the support path corresponding to the next main working path is a following support path (follow). Of the four output parameters, only one parameter is non-zero and the remaining parameters are zero, indicating that the support path corresponding to the next main working path is the support path corresponding to the non-zero parameter.

[0032] Before step S12, a digital simulation environment that matches the actual application environment of the 3D model to be manufactured is established in Grasshopper; the 3D model is simulated in the digital simulation environment, and the deep reinforcement learning model is trained based on the simulation results to obtain the pre-trained deep reinforcement learning model.

[0033] In a digital simulation environment that matches the actual application environment, the coordinates and directions of discrete points on each layer are converted into robot motion commands according to the robot grammar rules (KRL). The slave robot provides support in convex areas, while the master robot provides support in concave areas. A simulation model of sheet metal deformation is built using LS Dyna simulation software, and communication with the Grasshopper simulation environment is established via Socket to receive path data for the forming tool head and supporting tool head. The returned data is the formed curved surface and the formed curved surface springback value after forming.

[0034] When training the deep reinforcement learning model, first, an initial main working path is selected as the current simulation main working path based on the forming direction, and one of multiple candidate support paths is randomly selected as the initial current simulation support path based on the current simulation main working path. Then, the current simulation main working path and the current simulation support path are applied to the digital simulation environment to perform simulated forming, and a simulated forming surface and a simulated forming surface springback value corresponding to the current simulation main working path are obtained. Specifically, the coordinates and directions of discrete points of the current simulation main working path and the current simulation support path are converted into robot motion commands according to robot grammar rules, and a simulation model of sheet deformation is constructed using LS Dyna simulation software. Simulation forming is performed based on the robot motion commands, and a simulated forming surface and a simulated forming surface springback value corresponding to the current simulation main working path are returned.

[0035] Furthermore, the deviation value between the simulated forming surface and the target surface is input as a state vector to the deep reinforcement learning model, reinforcement learning of a support policy is performed, and a simulated support path corresponding to the next simulated main operating path and a current reward value are updated based on the simulated forming surface springback value. Preferably, second reference points on the simulated forming surface corresponding to each first reference point on the target surface are obtained, and error values ​​between each second reference point and each corresponding first reference point are calculated to construct the state vector. The state vector is then input to the deep reinforcement learning model to perform reinforcement learning of a support policy, and a simulated support path corresponding to the next simulated main operating path is output. The current reward value is then updated based on the simulated forming surface and the simulated forming surface springback value. The initial value of the current reward value is 0. If the springback value of the simulated forming surface is equal to or greater than a reference value, the current reward value is controlled to be decreased by a first preset value. If the springback value of the simulated forming surface is less than the reference value, the current reward value is controlled to be increased by a first preset value. If the forming of the simulated forming surface fails, the current reward value is controlled to be a second preset value. Here, the reference value, the first preset value, and the second preset value may be set as needed. Preferably, the reference value is 10%, the first preset value is 0.1, and the second preset value is -1.0.

[0036] Finally, the current simulated main motion path and the current simulated support path are cyclically updated based on the next simulated main motion path and the corresponding simulated support path, respectively. The robot arm is cyclically controlled based on the updated current simulated main motion path and the current simulated support path to perform incremental shaping and cyclically update the simulated shaping surface. The state vector is cyclically updated based on the updated simulated shaping surface and the target surface, and model parameters of the deep reinforcement learning model are adjusted based on the updated state vector and the reward value until a convergence condition for the deep reinforcement learning model is met. The convergence condition may be that a predetermined number of training attempts has been reached or that the error between the shaping surface and the target surface has reached or is smaller than a target value. A trained deep reinforcement learning model is then obtained and used for incremental shaping in a real environment.

[0037] After obtaining the pre-trained deep reinforcement learning model, in step S12, the robot arms of the master robot and the slave robot are respectively controlled to perform incremental forming in the actual application environment based on the current main operating path and the current support path, thereby obtaining a forming surface corresponding to the current main operating path. Specifically, as shown in FIG. 6, the metal sheet is fixed by a jig, and the robot arm of the master robot is controlled to operate along the current main operating path, and the robot arm of the slave robot is controlled to operate along the current support path, thereby performing incremental forming in the actual application environment to obtain a forming surface corresponding to the current main operating path. During incremental forming, the master-slave relationship between the master robot and the slave robot can be exchanged to form concave and convex molds within the same part. When the forming path coincides with the main operating path and the master robot performs the forming operation, the metal sheet is plastically deformed locally and stepwise along the direction of the main operating path of the forming tool head.

[0038] In step S13, the deviation value between the forming surface and the target surface is used as a state vector, and is applied to a pre-trained deep reinforcement learning model to perform reinforcement learning of the support policy, cyclically outputting a support path corresponding to the next main operating path, and cyclically updating the current main operating path and the current support path based on the next main operating path and the support path corresponding to the next main operating path until the incremental forming of the three-dimensional model is completed.

[0039] In an embodiment of the present invention, the deviation value between the forming surface and the target surface is input as a state vector into a deep reinforcement learning model, and reinforcement learning of the support policy is performed to output a support path corresponding to the next main operating path. Furthermore, in step S12, the forming surface springback value is also obtained. During reinforcement learning, the current reward value is updated based on the forming surface and the forming surface springback value. Preferably, a third preset number of first reference points are set on the target surface. After the forming surface is obtained, a third preset number of second reference points are set on the corresponding forming surface in one-to-one correspondence with each first reference point, and the error value between each second reference point and its corresponding first reference point is calculated to form a state vector. For example, as shown in FIG. 7, 16 second reference points are set on the forming surface. The 16 reference points are uniformly distributed on the forming surface and correspond one-to-one to the 16 first reference points on the target surface. The corresponding generated state vector is S:{d1;d2;d3···d15;d16}. The state vector is input to the deep reinforcement learning model to perform reinforcement learning of the support policy, and the support path corresponding to the next main working path is output, so that the next incremental shaping of the shaping surface can be subsequently performed.

[0040] After outputting a support path corresponding to the next main operating path through reinforcement learning of the support policy using the deep reinforcement learning model, the current main operating path and the current support path are updated based on the next main operating path and the support path corresponding to the next main operating path. That is, the current main operating path is updated to the next main operating path, and the current support path is updated to the support path corresponding to the next main operating path. Then, returning to step S12, the robot arm is cyclically controlled based on the updated current main operating path and current support path to perform incremental forming and generate a new forming surface. This cycle is continued until incremental forming of the 3D model is completed.

[0041] In an embodiment of the present invention, a trained deep reinforcement learning model is used to control a robot arm in a real environment to perform incremental molding, and then a 3D scanner is used to digitize the molding results, an error calculation is performed, and the model parameters of the deep reinforcement learning model are adjusted based on the calculated error.

[0042] The two-point incremental forming manufacturing method based on deep reinforcement learning according to an embodiment of the present invention uses two robots (e.g., KR-210) working together as hardware. The master robot performs the forming operation, plastically deforming the metal plate locally and incrementally along the tool path direction. The slave robot has four support policies, and at each step, a deep reinforcement learning model adjusts and selects a support path corresponding to one of the support policies. Each step is defined as the master robot's forming tool head completing one path. In the training phase, a digital simulation environment is first established, and a deep neural network is constructed. The deep neural network outputs a support policy for the slave robot based on the error value between the forming surface and the target surface in the simulation environment, and the parameters of the deep neural network are optimized based on the effect of the support policy. Finally, the task is completed when the robot completes the entire path and the error of the formed part meets the target requirement. The two-point incremental forming manufacturing method based on deep reinforcement learning according to an embodiment of the present invention has strong real-time capabilities and high flexibility, effectively improving the accuracy of the formed part and reducing experimental costs.

[0043] In accordance with an embodiment of the present invention, a two-point incremental forming manufacturing method based on deep reinforcement learning obtains a 3D model to be manufactured, divides the 3D model into multiple layers, obtains multiple main working paths and multiple candidate support paths corresponding to each main working path, and selects an initial current main working path and a support path corresponding to the initial main working path based on the forming direction as an initial current main working path and a current support path; based on the current main working path and the selected current support path, cyclically controls the robot arms of the master robot and the slave robot in an actual application environment to perform incremental forming and obtain a forming surface corresponding to the current main working path; uses the deviation value between the forming surface and the target surface as a state vector, and applies it to a pre-trained deep reinforcement learning model to perform reinforcement learning of the support policy, cyclically outputting a support path corresponding to the next main working path; and cyclically updating the current main working path and the current support path based on the next main working path and the support path corresponding to the next main working path until the incremental forming of the 3D model is completed, thereby adjusting the support policy of the slave robot, thereby improving flexibility and improving forming accuracy.

[0044] The foregoing describes specific embodiments of the present invention. In some cases, the actions or steps described in the embodiments of the present invention can be performed in a different order than the examples and still achieve desirable results. Also, processes depicted in the figures do not necessarily have to be performed in the particular order or sequential order shown to achieve desirable results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0045] Based on the same idea, an embodiment of the present invention further provides a two-point incremental forming manufacturing apparatus based on deep reinforcement learning. As shown in Fig. 8, the two-point incremental forming manufacturing apparatus based on deep reinforcement learning includes a path acquisition unit, an incremental forming unit, and a reinforcement learning unit.

[0046] The path obtaining unit obtains a three-dimensional model to be manufactured, divides the three-dimensional model into multiple layers, obtains multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, and selects an initial main working path and one support path corresponding to the initial main working path as an initial current main working path and a current support path based on the molding direction.

[0047] The incremental forming unit performs incremental forming by respectively circulating the robot arms of the master robot and the slave robot in an actual application environment based on the current main working path and the current support path, and obtains a forming surface corresponding to the current main working path.

[0048] The reinforcement learning unit uses the deviation value between the forming surface and the target surface as a state vector, applies it to a pre-trained deep reinforcement learning model to perform reinforcement learning of the support policy, cyclically outputs a support path corresponding to the next main working path, and cyclically updates the current main working path and the current support path based on the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.

[0049] For simplicity of explanation, the above apparatus has been described as being divided into various modules each having different functions. Of course, when implementing an embodiment of the present invention, the functions of each module may be implemented by the same or multiple pieces of software and / or hardware.

[0050] The apparatus according to the above-mentioned embodiment is applied to the method corresponding to the above-mentioned embodiment, and has the beneficial effects of the corresponding method embodiment, so the description thereof will be omitted here.

[0051] Based on the same inventive idea, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable by the processor, and when the processor executes the program, the method according to any of the above embodiments is realized.

[0052] An embodiment of the present invention provides a non-volatile computer storage medium having stored thereon at least one executable command, the command being executed by a computer to perform the method of any of the above embodiments.

[0053] 9 is a schematic diagram showing a specific hardware structure of an electronic device according to this embodiment. The device may include a processor 901, a memory 902, an input / output interface 903, a communication interface 904, and a bus 905. The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other via the bus 905 within the device.

[0054] The processor 901 may be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, and executes related programs to implement the technical solutions provided in the embodiments of the method according to the present invention.

[0055] The memory 902 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), an SRAM (Static Random Access Memory), a DRAM (Dynamic Random Access Memory), etc. The memory 902 stores an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the method according to the present invention through software or firmware, the relevant program codes are stored in the memory 902 and are read and executed by the processor 901.

[0056] The input / output interface 903 is connected to an input / output module to realize input and output of information. The input / output module may be disposed as a component in the device (not shown) or may be externally attached to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator lamp, etc.

[0057] The communication interface 904 is connected to a communication module (not shown) to realize communication interaction between the device and other devices, where the communication module may realize communication via a wired method (e.g., USB, network cable, etc.) or a wireless method (e.g., mobile network, WIFI, Bluetooth, etc.).

[0058] The bus 905 includes a path for transmitting information between each component of the device (eg, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904).

[0059] Although the above device only shows a processor 901, a memory 902, an input / output interface 903, a communication interface 904, and a bus 905, in a specific implementation, the device may further include other components necessary for normal operation. Those skilled in the art will understand that the above device may only include the components necessary to implement the solutions of the embodiments of the present invention, and does not necessarily include all the components shown in the drawings.

[0060] For those skilled in the art, the above-mentioned embodiments discussed are merely illustrative and do not imply that the scope of the present application is limited to these examples. Based on the concept of the present application, the technical features in the above-mentioned embodiments or different embodiments may be combined, steps may be implemented in any order, and there are many other variations of the above-mentioned different aspects of the present application, which are not described in detail for the sake of brevity.

[0061] The purpose of this application is to cover all such substitutions, modifications and variations that fall within the broad scope of the embodiments of the present invention, and therefore any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall all be included in the scope of protection of this application.

Claims

1. obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into a plurality of layers to obtain a plurality of main working paths and a plurality of candidate support paths corresponding to each of the main working paths; and selecting an initial main working path and one support path corresponding to the initial main working path as an initial current main working path and a current support path based on a molding direction; According to the current main working path and the selected current support path, in an actual application environment, respectively, circularly control the robot arms of the master robot and the slave robot to perform incremental forming, and obtain a forming surface corresponding to the current main working path; a step of cyclically updating the current main operating path and the current support path based on the next main operating path and the support path corresponding to the next main operating path until the incremental forming of the three-dimensional model is completed; A two-point incremental forming manufacturing method based on deep reinforcement learning, comprising:

2. The step of obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into a plurality of layers, and obtaining a plurality of main actuation paths and a plurality of candidate support paths corresponding to each of the main actuation paths, comprises: Obtaining a three-dimensional model to be manufactured, and applying an offset function on the surface to divide the three-dimensional model into a plurality of layers along a molding direction with a predetermined layer thickness to obtain a first predetermined number of curved paths; distributing a second predetermined number of discrete points at a predetermined point interval for each of the curved paths, and generating a main operating path corresponding to the curved paths based on the discrete points; for each of the primary working paths, obtaining a plurality of candidate supporting paths corresponding to the primary working path based on a respective plurality of supporting policies; the support policy is one of a global support policy, a local peripheral support policy, a local frontal support policy, and a following support policy; The two-point incremental forming manufacturing method based on deep reinforcement learning as claimed in claim 1.

3. Before the step of: a robot arm performs incremental forming in an actual application environment according to a current main working path and a selected current support path to obtain a forming surface corresponding to the current main working path; building a digital simulation environment in Grasshopper that matches the actual application environment of the three-dimensional model to be manufactured; Simulating the three-dimensional model in the digital simulation environment, and training the deep reinforcement learning model based on a simulation result to obtain a pre-trained deep reinforcement learning model; The two-point incremental forming manufacturing method based on deep reinforcement learning according to claim 1, comprising:

4. The step of simulating the three-dimensional model in the digital simulation environment and training the deep reinforcement learning model based on a simulation result to obtain the pre-trained deep reinforcement learning model includes: selecting an initial main working path as a current simulation main working path based on a forming direction, and randomly selecting one of a plurality of candidate support paths as an initial current simulation support path based on the current simulation main working path; Applying the digital simulation environment to perform simulation forming based on the current simulation main working path and the current simulation support path, and obtaining a simulation forming surface and a simulation forming surface springback value corresponding to the current simulation main working path; inputting the deviation value between the simulation forming surface and the target surface as a state vector into the deep reinforcement learning model, performing reinforcement learning of the support policy, and updating the simulation support path corresponding to the next simulation main working path and the current reward value according to the simulation forming surface springback value; cyclically updating the current simulation main motion path and the current simulation support path based on the next simulation main motion path and the corresponding simulation support path, respectively, and cyclically controlling a robot arm to perform incremental forming based on the updated current simulation main motion path and the updated current simulation support path, thereby cyclically updating the simulation forming surface; The two-point incremental forming manufacturing method based on deep reinforcement learning according to claim 3, comprising:

5. The step of applying the digital simulation environment to perform simulation forming based on the current simulation main working path and the current simulation support path, and obtaining a simulation forming surface and a simulation forming surface springback value corresponding to the current simulation main working path, converting the coordinates and directions of discrete points of the current simulated main action path and the current simulated support path into robot motion commands according to robot grammar rules; Using simulation software to construct a simulation model of sheet deformation, and performing simulation forming according to the robot motion command, and returning a simulation forming curved surface and a simulation forming curved surface springback value corresponding to the current simulation main motion path; The two-point incremental forming manufacturing method based on deep reinforcement learning according to claim 4, comprising:

6. The step of inputting the deviation value between the simulation forming surface and the target surface as a state vector into the deep reinforcement learning model, performing reinforcement learning of a support policy, and updating the simulation support path corresponding to the next simulation main working path and the current reward value according to the simulation forming surface springback value, acquiring second reference points on the simulation forming surface corresponding to the first reference points on the target surface, calculating an error value between each of the second reference points and the corresponding first reference points, and constructing the state vector; inputting the state vector into the deep reinforcement learning model, performing reinforcement learning of a support policy, and outputting a simulated support path corresponding to a next simulated main action path; updating the current reward value based on the simulated forming surface and the simulated forming surface springback value; The two-point incremental forming manufacturing method based on deep reinforcement learning according to claim 4, comprising:

7. updating the current reward value based on the simulated forming surface and the simulated forming surface springback value; an initial value of the current reward value is 0, and when the springback value of the simulated forming curved surface is equal to or greater than a reference value, the current reward value is controlled to be reduced by a first preset value; If the springback value of the simulated forming surface is less than a reference value, controlling the current reward value to increase by a first preset value; If the shaping of the simulation shaping surface fails, controlling the current reward value to be a second preset value; The two-point incremental forming manufacturing method based on deep reinforcement learning according to claim 6, comprising:

8. a path obtaining unit for obtaining a three-dimensional model to be manufactured, dividing the three-dimensional model into a plurality of layers, obtaining a plurality of main working paths and a plurality of candidate support paths corresponding to each of the main working paths, and selecting an initial main working path and one support path corresponding to the initial main working path as an initial current main working path and a current support path according to a molding direction; an incremental forming unit for controlling the robot arms of the master robot and the slave robot to perform incremental forming in an actual application environment according to a current main working path and a current support path, and obtaining a forming surface corresponding to the current main working path; a reinforcement learning unit that uses a deviation value between the forming surface and the target surface as a state vector, applies it to a pre-trained deep reinforcement learning model to perform reinforcement learning of a support policy, cyclically outputs a support path corresponding to a next main operating path, and cyclically updates the current main operating path and the current support path based on the next main operating path and the support path corresponding to the next main operating path until the incremental forming of the three-dimensional model is completed; A two-point incremental forming manufacturing device based on deep reinforcement learning, characterized by:

9. a memory; a processor; and a computer program stored in the memory and executable by the processor; When the processor executes the program, the two-point incremental forming manufacturing method based on deep reinforcement learning according to any one of claims 1 to 7 is realized. An electronic device characterized by:

10. At least one executable command is stored, and the processor executes the executable command to perform the deep reinforcement learning-based two-point incremental forming manufacturing method according to any one of claims 1 to 7. A computer storage medium comprising:

Citation Information

Patent Citations

  • Incremental forming method based on deep learning

    CN111633111A

  • Nc program preparation method for successive molding, successive molding method and record medium

    JP1999327619A

  • Systems and methods for compensating for spring back of structures formed through incremental sheet forming

    JP2022001381A

  • System and method for accumulative double sided incremental forming

    US20130103177A1