Double-point incremental forming manufacturing method and device based on deep reinforcement learning
Through the two-point progressive forming manufacturing method based on deep reinforcement learning, the support strategy of the secondary tool head is dynamically adjusted, and the problem of low forming accuracy in the existing technology is solved, achieving higher forming accuracy and flexibility.
Patent Information
- Application Number
- CN202210410301.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing two-point progressive forming method has the problem of fixed support strategies of secondary tool heads and low forming accuracy, which limits its widespread use in industrial applications.
The two-point progressive forming manufacturing method based on deep reinforcement learning is adopted. By obtaining the three-dimensional model to be manufactured and obtaining the main working path and support path layered, the pre-trained deep reinforcement learning model is used for reinforcement learning of support strategies, and the support strategy of the secondary tool head is dynamically adjusted.
It improves forming accuracy and flexibility, can effectively optimize forming control, and improves the dimensional accuracy and forming range of molded parts.
Smart Images

Figure CN114757102B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of additive manufacturing, and specifically relates to a double-point incremental forming manufacturing method and device based on deep reinforcement learning. Background Art
[0002] Incremental forming is a flexible manufacturing technology that can produce target parts without the need for special molds. Incremental forming uses a hemispherical tool attached to a robotic arm or CNC machine tool. The tool moves along a pre-programmed path to produce local plastic deformation on the metal sheet, so that it reaches the desired shell shape. Double-point incremental forming uses two hemispherical forming tools to achieve local incremental deformation of the material to obtain the final formed part. The principle of double-point incremental forming is that while the main tool head (forming ram) is processing the sheet on one side, the auxiliary tool head (support ram) supports the sheet on the other side, and the running trajectory of the auxiliary tool head is subordinate to the main tool head. This double-point incremental forming method can further improve the formability of the sheet and effectively improve the dimensional accuracy of the formed part. However, the existing double-point incremental forming method has many disadvantages. The support strategy of the auxiliary tool head is fixed during the incremental forming process, the geometric accuracy of the forming result is low and the forming range is small, which limits its wide industrial application. At present, the improvement method of the incremental forming accuracy problem is mainly to measure the rebound amount of the material for compensation, which is difficult to optimize in terms of forming control. Summary of the invention
[0003] The present invention provides a double-point incremental forming manufacturing method and device based on deep reinforcement learning to solve the problem that the existing support strategy of the slave robot is fixed and the forming accuracy is low.
[0004] Based on the above purpose, an embodiment of the present invention provides a double-point incremental forming manufacturing method based on deep reinforcement learning, including: obtaining a three-dimensional model to be manufactured and layering the three-dimensional model to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, selecting a main working path and a support path corresponding to the selected main working path as the initial current main working path and current support path; according to the current main working path and the selected current support path, cyclically controlling the mechanical arms of the master and slave robots in an actual application environment to perform incremental forming, and obtaining a forming surface corresponding to the current main working path; using the deviation value between the forming surface and the target surface as a state vector, applying a pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically outputting the support path corresponding to the next main working path, and cyclically updating the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path, until the incremental forming of the three-dimensional model is completed.
[0005] Optionally, the method of obtaining a three-dimensional model to be manufactured and layering the three-dimensional model to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths includes: obtaining the three-dimensional model to be manufactured, applying an along-surface offset function to layer the three-dimensional model along a forming direction with a preset layer thickness, and obtaining a first preset number of curved paths; for each of the curved paths, dividing a second preset number of discrete points with a preset point spacing, and generating a main working path corresponding to the curved path based on the discrete points; for each of the main working paths, obtaining multiple candidate support paths corresponding to the main working path according to multiple support strategies, wherein the support strategy is one of a global support strategy, a local peripheral support strategy, a local frontal support strategy, and a follow-up support strategy.
[0006] Optionally, before controlling the robotic arm to perform incremental forming in a real environment according to the current main working path and the selected support path to obtain a forming surface corresponding to the current main working path, it includes: constructing a digital simulation environment in Grasshopper that is consistent with the actual application environment of the three-dimensional model to be manufactured; simulating the three-dimensional model in the digital simulation environment, and training the deep reinforcement learning model in combination with the simulation results to obtain the pre-trained deep reinforcement learning model.
[0007] Optionally, the three-dimensional model is simulated in the digital simulation environment, and the deep reinforcement learning model is trained in combination with the simulation results to obtain the pre-trained deep reinforcement learning model, including: selecting an initial main working path as the current simulation main working path according to the forming direction, and randomly selecting one of the multiple candidate support paths as the initial current simulation support path from the candidate support paths according to the current simulation main working path; applying the digital simulation environment to perform simulated forming according to the current simulation main working path and the current simulation support path, and obtaining a simulated forming surface and a simulated forming surface springback value corresponding to the current simulation main working path; using the deviation value between the simulated forming surface and the target surface as a state vector input into the deep reinforcement learning model. The deep reinforcement learning model is used to perform reinforcement learning of the support strategy, and in combination with the springback value of the simulated forming surface, the simulated support path corresponding to the next simulation main working path and the current reward value are updated; the current simulation main working path and the current simulation support path are respectively cyclically updated according to the next simulation main working path and the corresponding simulation support path, and the robot arm is cyclically controlled to perform incremental forming according to the updated current simulation main working path and the current simulation support path, and the simulated forming surface is cyclically updated; the state vector is cyclically updated according to the updated simulated forming surface and the target surface, and the model parameters of the deep reinforcement learning model are adjusted according to the updated state vector and the reward value until the convergence conditions of the deep reinforcement learning model are met.
[0008] Optionally, the digital simulation environment is applied to perform simulated forming according to the current simulation main working path and the current simulation support path to obtain the simulated forming surface and the simulated forming surface springback value corresponding to the current simulation main working path, including: converting the coordinates and directions of the discrete points of the current simulation main working path and the current simulation support path into robot motion instructions according to the robot grammar rules; using simulation software to construct a simulation model of sheet deformation, performing simulated forming according to the robot motion instructions, and returning the simulated forming surface and the simulated forming surface springback value corresponding to the current simulation main working path.
[0009] Optionally, the deviation value between the simulated forming surface and the target surface is used as a state vector to input into the deep reinforcement learning model for reinforcement learning of the support strategy, and combined with the rebound value of the simulated forming surface, the simulated support path corresponding to the next simulation main working path and the current reward value are updated, including: obtaining each second reference point on the simulated forming surface corresponding to each first reference point on the target surface, and calculating the error value between each second reference point and the corresponding first reference point to form the state vector; inputting the state vector into the deep reinforcement learning model for reinforcement learning of the support strategy, and outputting the simulated support path corresponding to the next simulation main working path; updating the current reward value according to the simulated forming surface and the rebound value of the simulated forming surface.
[0010] Optionally, the current reward value is obtained according to the simulated forming surface and the springback value of the simulated forming surface, including: the initial value of the current reward value is 0; if the springback value of the simulated forming surface is greater than or equal to a reference value, the current reward value is controlled to decrease by a first preset value; if the springback value of the simulated forming surface is less than the reference value, the current reward value is controlled to increase by a first preset value; if the simulated forming surface fails to form, the current reward value is controlled to be a second preset value.
[0011] Based on the same inventive concept, an embodiment of the present invention further proposes a dual-point incremental forming manufacturing device based on deep reinforcement learning, comprising: a path acquisition unit, used to acquire a three-dimensional model to be manufactured and to hierarchically acquire a plurality of main working paths and a plurality of candidate support paths corresponding to each of the main working paths, and to select an initial main working path and a support path corresponding to the initial main working path as the initial current main working path and the current support path according to the forming direction; an incremental forming unit, used to cyclically control the mechanical arms of the master and slave robots to perform incremental forming in an actual application environment according to the current main working path and the current support path, and to acquire a forming surface corresponding to the current main working path; a reinforcement learning unit, used to use the deviation value between the forming surface and the target surface as a state vector, apply a pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically output the support path corresponding to the next main working path, and cyclically update the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path, until the incremental forming of the three-dimensional model is completed.
[0012] Based on the same inventive concept, an embodiment of the present invention further proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method when executing the program.
[0013] Based on the same inventive concept, an embodiment of the present invention further proposes a computer storage medium, in which at least one executable instruction is stored, and the executable instruction enables a processor to execute the aforementioned method.
[0014] The beneficial effects of the present invention are as follows: as can be seen from the above, an embodiment of the present invention provides a dual-point incremental forming manufacturing method and device based on deep reinforcement learning, the method comprising: obtaining a three-dimensional model to be manufactured and layering the three-dimensional model to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, selecting an initial main working path and a support path corresponding to the initial main working path as the initial current main working path and the current support path according to the forming direction; according to the current main working path and the selected current support path, cyclically controlling the mechanical arms of the master and slave robots in an actual application environment to perform incremental forming, and obtaining a forming surface corresponding to the current main working path; using the deviation value between the forming surface and the target surface as a state vector, applying a pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically outputting the support path corresponding to the next main working path, and cyclically updating the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed, and the support strategy of the slave robot can be adjusted, with high flexibility and high forming accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 It is a schematic flow chart of a double-point incremental forming manufacturing method based on deep reinforcement learning in an embodiment of the present invention;
[0017] Figure 2 is a schematic diagram of a global support strategy in an embodiment of the present invention;
[0018] Figure 3 It is a schematic diagram of a local front support strategy in an embodiment of the present invention;
[0019] Figure 4 Schematic diagram of a follow-up support strategy in an embodiment of the present invention;
[0020] Figure 5 A schematic diagram of a local peripheral support strategy in an embodiment of the present invention;
[0021] Figure 6 It is a schematic diagram of double-point incremental forming in an embodiment of the present invention;
[0022] Figure 7 is an example diagram of a formed curved surface in an embodiment of the present invention;
[0023] Figure 8 Schematic diagram of the structure of a double-point incremental forming manufacturing device based on deep reinforcement learning in an embodiment of the present invention;
[0024] Fig. 9 Schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0026] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present invention should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present invention do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0027] The embodiment of the present invention provides a double-point incremental forming manufacturing method based on deep reinforcement learning. Figure 1 As shown, the double-point incremental forming manufacturing method based on deep reinforcement learning includes:
[0028] Step S11: Acquire the three-dimensional model to be manufactured and layer the three-dimensional model to obtain multiple main working paths and multiple candidate support paths corresponding to each of the main working paths, and select an initial main working path and a support path corresponding to the initial main working path as the initial current main working path and current support path according to the forming direction.
[0029] In an embodiment of the present invention, the three-dimensional surface model to be manufactured is imported into the Grasshopper programming environment based on Rhino. In step S11, optionally, the three-dimensional model to be manufactured is first obtained, and the three-dimensional model is layered along the forming direction with a preset layer thickness using an offset function along the surface to obtain a first preset number of curved paths. The preset layer thickness depends on the size of the forming tool head of the main robot, which is generally in the range of 0.5-2mm. The forming direction can be along the x-axis direction, and of course it can also be other directions, which is not limited here. The first preset number is determined by the preset layer thickness and the size of the three-dimensional model. The smaller the preset layer thickness and the larger the three-dimensional model, the larger the first preset number. The curved path is a 2D curved path.
[0030] Then, for each of the curved paths, a second preset number of discrete points are divided according to the preset point spacing, and a main working path corresponding to the curved path is generated according to the discrete points. The preset point spacing can be set as needed, generally in the range of 1-3mm. The second preset number is determined by the preset point spacing and the size of the curved path. A series of planes are established in the direction of the tangent plane of the curved surface according to each discrete point. The center point of the plane is the moving position of the forming tool head, and the Z axis of the plane is the forming direction of the forming tool head. Each discrete point serves as a forming path point, and the combination of the moving position and forming direction of the forming tool head corresponding to each discrete point on any curved path constitutes a main working path.
[0031] Finally, for each of the main working paths, multiple candidate support paths corresponding to the main working path are obtained according to multiple support strategies, and the support strategy is one of a global support strategy, a local peripheral support strategy, a local frontal support strategy and a following support strategy.
[0032] According to the global support strategy, the outer contour of the forming area is offset by a first preset distance to generate a global support curve, and the forming path points are mapped onto the global support curve to generate a global support path. The first preset distance can be set as needed, preferably the radius size of the forming tool head. Figure 2 As shown, the slave robot moves the support tool head along the part boundary.
[0033] According to the local front support strategy, the main working path of the forming area is mirrored with the metal sheet surface plane, and the forming path point plane is reversed to generate a local front support path. Figure 3 As shown, the supporting tool head of the slave robot is directly opposite to the forming tool head of the master robot.
[0034] According to the follow-up support strategy, the main working path of the forming area is mirrored with the metal sheet surface plane, the forming path point plane is reversed, and the forming path point list is offset forward by 3 items to generate the follow-up support path. Figure 4 As shown, in the case of partial frontal support, the support tool head of the slave robot lags behind the forming tool head of the master robot by a second preset distance. The second preset distance can be set as required, preferably the diameter of the forming tool head.
[0035] According to the local peripheral support strategy, the following support path is offset backward by one layer to generate a local peripheral support path. Figure 5 As shown, the support tool head of the slave robot directly follows the relative offset path of the forming tool head of the master robot, forming a forming gap between the forming tools.
[0036] The above four support strategies have their own advantages and disadvantages. The local support strategy (including the local lower support strategy and the local peripheral support strategy) can strengthen the forming of local details, but the overall error will be large; the global support strategy can unify the overall error, but the detail forming error is large; the follow-up support strategy can improve the surface quality of the forming. In the subsequent reinforcement learning environment, the robot's behavior space is used to execute the four support strategies and select one of the four candidate support paths.
[0037] An initial main working path is selected from multiple main working paths according to the forming direction as the initial current main working path, and one of the four candidate supporting paths corresponding to the initial main working path is randomly selected as the initial current supporting path.
[0038] Step S12: According to the current main working path and the selected current supporting path, the manipulator arms of the master and slave robots are respectively cyclically controlled in an actual application environment to perform incremental forming, and a forming surface corresponding to the current main working path is obtained.
[0039] In an embodiment of the present invention, the behavior space of the deep reinforcement learning model is the support strategy of the slave robot. The output of the deep reinforcement learning model includes 4 dimensions, each of which corresponds to a support path of the slave robot, namely, a global support path, a local front support path, a local peripheral support path, and a follow support path. For example, the deep reinforcement learning model includes an input layer, 3 hidden layers, and an output layer. The state vector input by the input layer includes multiple parameters, for example, including 16 parameters S:{d1; d2; ...d15; d16}, the three hidden layers are 64 neurons, 32 neurons, and 16 neurons, respectively, and the output layer outputs 4 parameters A:{g; l1; l2; f}, where g represents the probability that the support path corresponding to the next main working path is a global support path (global), l1 represents the probability that the support path corresponding to the next main working path is a local front support path (local1), l2 represents the probability that the support path corresponding to the next main working path is a local peripheral support path (local2), and f represents the probability that the support path corresponding to the next main working path is a follow support path (follow). Among the four parameters outputted, only one parameter is not 0, and the other parameters are 0, indicating that the supporting path corresponding to the next main working path is the supporting path corresponding to the parameter that is not 0.
[0040] Before step S12, a digital simulation environment that matches the actual application environment of the three-dimensional model to be manufactured is constructed in Grasshopper; the three-dimensional model is simulated in the digital simulation environment, and the deep reinforcement learning model is trained based on the simulation results to obtain the pre-trained deep reinforcement learning model.
[0041] In a digital simulation environment that matches the actual application environment, the coordinates and directions of each discrete point are converted into robot motion instructions according to the robot grammar rules (KRL). In the convex area, the slave robot is the support, and in the concave area, the master robot is the support. The simulation model of sheet deformation is built using LS Dyna simulation software, which communicates with the Grasshopper simulation environment through Socket, receives the path data of the forming tool head and the support tool head, and returns the formed surface and the springback value of the formed surface after forming.
[0042] When training the deep reinforcement learning model, first select the initial main working path as the current simulation main working path according to the forming direction, and randomly select one of the candidate multiple support paths as the initial current simulation support path according to the current simulation main working path. Then, according to the current simulation main working path and the current simulation support path, the digital simulation environment is applied to perform simulation forming, and the simulation forming surface and the simulation forming surface springback value corresponding to the current simulation main working path are obtained. Specifically, according to the robot grammar rules, the coordinates and directions of the discrete points of the current simulation main working path and the current simulation support path are converted into robot motion instructions; the LSDyna simulation software is used to build a simulation model of sheet deformation, and simulation forming is performed according to the robot motion instructions, and the simulation forming surface and the simulation forming surface springback value corresponding to the current simulation main working path are returned.
[0043] Then, the deviation value between the simulated forming surface and the target surface is used as a state vector to input the deep reinforcement learning model for reinforcement learning of the support strategy, and combined with the springback value of the simulated forming surface, the simulated support path corresponding to the next simulation main working path and the current reward value are updated. Optionally, each second reference point on the simulated forming surface corresponding to each first reference point on the target surface is obtained, and the error value between each second reference point and the corresponding first reference point is calculated to form the state vector; the state vector is input into the deep reinforcement learning model for reinforcement learning of the support strategy, and the simulated support path corresponding to the next simulation main working path is output; the current reward value is updated according to the simulated forming surface and the springback value of the simulated forming surface. The initial value of the current reward value is 0. If the springback value of the simulated forming surface is greater than or equal to the reference value, the current reward value is controlled to decrease by a first preset value; if the springback value of the simulated forming surface is less than the reference value, the current reward value is controlled to increase by a first preset value; if the forming of the simulated forming surface fails, the current reward value is controlled to be a second preset value. The reference value, the first preset value and the second preset value can be set as required. Preferably, the reference value is 10%, the first preset value is 0.1, and the second reference value is -1.0.
[0044] Finally, the current simulation main working path and the current simulation support path are cyclically updated according to the next simulation main working path and the corresponding simulation support path, and the robot arm is cyclically controlled to perform incremental forming according to the updated current simulation main working path and the current simulation support path, and the simulation forming surface is cyclically updated; the state vector is cyclically updated according to the updated simulation forming surface and the target surface, and the model parameters of the deep reinforcement learning model are adjusted according to the updated state vector and the reward value until the convergence condition of the deep reinforcement learning model is met. The convergence condition can be that a preset number of training times is reached, or the error between the forming surface and the target surface reaches or is less than the target value. At this point, a trained deep reinforcement learning model is obtained for incremental forming in a real environment.
[0045] After obtaining the pre-trained deep reinforcement learning model, in step S12, the master and slave robot arms are respectively controlled to perform incremental forming in the actual application environment according to the current main working path and the current supporting path, and the forming surface corresponding to the current main working path is obtained. Specifically, Figure 6 As shown, the metal sheet is fixed by a fixture, the mechanical arm of the master robot is controlled to run along the current main working path, and the mechanical arm of the slave robot is controlled to run along the current support path, so as to complete the incremental forming in the actual application environment to obtain the forming surface corresponding to the current main working path. In the incremental forming process, by exchanging the master and slave roles of the master and slave robots, concave and convex shapes can be formed in the same part. The forming path is consistent with the main working path. When the master robot performs the forming work, the metal sheet is gradually subjected to local plastic deformation along the main working path direction of the forming tool head.
[0046] Step S13: Use the deviation value between the forming surface and the target surface as the state vector, apply the pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically output the support path corresponding to the next main working path, and cyclically update the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.
[0047] In an embodiment of the present invention, the deviation value between the forming surface and the target surface is used as a state vector to input the deep reinforcement learning model to perform reinforcement learning of the support strategy, and output a support path corresponding to the next main working path. In step S12, the springback value of the forming surface is also obtained. During the reinforcement learning process, the current reward value is updated according to the forming surface and the springback value of the forming surface. Optionally, the target surface is provided with a third preset number of first reference points. After obtaining the forming surface, a third preset number of second reference points corresponding to each first reference point are set on the forming surface, and the error value between each second reference point and the corresponding first reference point is calculated to form a state vector. For example, Figure 7 As shown in the figure, 16 second reference points are set on the forming surface, and the 16 reference points are evenly distributed on the forming surface, corresponding to the 16 first reference points on the target surface. The corresponding state vector is S:{d1; d2; ...d15; d16}. The state vector is input into the deep reinforcement learning model for reinforcement learning of the support strategy, and the support path corresponding to the next main working path is output so as to carry out the incremental forming of the next forming surface.
[0048] After the reinforcement learning of the support strategy by the deep reinforcement learning model outputs the support path corresponding to the next main working path, the current main working path and the current support path are updated according to the next main working path and the support path corresponding to the next main working path. That is, the current main working path is updated to the next main working path, and the current support path is updated to the support path corresponding to the next main working path. Then return to step S12, and control the robot arm to perform incremental forming according to the updated current main working path and the current support path to generate a new forming surface. This cycle is repeated until the incremental forming of the three-dimensional model is completed.
[0049] In an embodiment of the present invention, after using the trained deep reinforcement learning model to control the robotic arm to perform incremental forming in a real environment, a three-dimensional scanner can be used to digitize the forming results and perform error calculation, and further, the model parameters of the deep reinforcement learning model can be adjusted according to the calculated error.
[0050] The double-point incremental forming manufacturing method based on deep reinforcement learning in the embodiment of the present invention uses two robots (such as KR-210) to cooperate with each other in hardware. The main robot performs the forming work and gradually produces local plastic deformation of the metal sheet along the tool path direction. The slave robot has four support strategies. Each step is performed, and the support path corresponding to one of the support strategies is selected through the deep reinforcement learning model adjustment, wherein each path completed by the forming tool head of the main robot is a step. In the training stage, a digital simulation environment is first established, and a deep neural network is constructed. The deep neural network outputs the support strategy of the slave robot according to the error value between the forming surface and the target surface in the simulation environment, and optimizes the parameters of the deep neural network according to the effect of the support strategy. If the robot finally completes the entire path and the error of the formed part meets the target requirement, the task is completed. The double-point incremental forming manufacturing method based on deep reinforcement learning in the embodiment of the present invention has strong real-time performance and high flexibility, and can effectively improve the precision of the formed part and reduce the experimental cost.
[0051] The double-point incremental forming manufacturing method based on deep reinforcement learning of the embodiment of the present invention obtains the three-dimensional model to be manufactured and hierarchically obtains multiple main working paths and multiple candidate support paths corresponding to each main working path, selects the initial main working path and a support path corresponding to the initial main working path as the initial current main working path and the current support path according to the forming direction; according to the current main working path and the selected current support path, the mechanical arms of the master and slave robots are respectively cyclically controlled in the actual application environment to perform incremental forming, and the forming surface corresponding to the current main working path is obtained; the deviation value between the forming surface and the target surface is used as the state vector, and the pre-trained deep reinforcement learning model is applied to perform reinforcement learning of the support strategy, and the support path corresponding to the next main working path is cyclically output, and the current main working path and the current support path are cyclically updated according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed, and the support strategy of the slave robot can be adjusted, with high flexibility and high forming accuracy.
[0052] The above specific embodiments of the present invention are described. In some cases, the actions or steps recorded in the embodiments of the present invention can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0053] Based on the same concept, an embodiment of the present invention also provides a double-point incremental forming manufacturing device based on deep reinforcement learning. Figure 8As shown, the double-point incremental forming manufacturing device based on deep reinforcement learning includes: a path acquisition unit, an incremental forming unit and a reinforcement learning unit. Among them,
[0054] a path acquisition unit, used to acquire a three-dimensional model to be manufactured and to perform layering on the three-dimensional model to acquire a plurality of main working paths and a plurality of candidate supporting paths corresponding to each of the main working paths, and to select an initial main working path and a supporting path corresponding to the initial main working path as an initial current main working path and a current supporting path according to a forming direction;
[0055] An incremental forming unit, used to cyclically control the mechanical arms of the master and slave robots to perform incremental forming in an actual application environment according to the current main working path and the current supporting path, and obtain a forming surface corresponding to the current main working path;
[0056] A reinforcement learning unit is used to use the deviation value between the forming surface and the target surface as a state vector, apply a pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically output a support path corresponding to the next main working path, and cyclically update the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.
[0057] For the convenience of description, the above device is described as various modules according to their functions. Of course, when implementing the embodiment of the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0058] The device of the above embodiment is applied to the corresponding method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0059] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of the above embodiments is implemented.
[0060] An embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the method described in any of the above embodiments.
[0061] Fig. 9A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 901, a memory 902, an input / output interface 903, a communication interface 904, and a bus 905. The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other in communication within the device through the bus 905.
[0062] The processor 901 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided by the method embodiment of the present invention.
[0063] The memory 902 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 902 can store an operating system and other application programs. When the technical solution provided by the method embodiment of the present invention is implemented by software or firmware, the relevant program code is stored in the memory 902 and called and executed by the processor 901.
[0064] The input / output interface 903 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0065] The communication interface 904 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0066] The bus 905 includes a path that transmits information between various components of the device (eg, the processor 901 , the memory 902 , the input / output interface 903 , and the communication interface 904 ).
[0067] It should be noted that, although the above device only shows the processor 901, the memory 902, the input / output interface 903, the communication interface 904 and the bus 905, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiment of the present invention, and does not necessarily include all the components shown in the figure.
[0068] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application is limited to these examples. Based on the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the present application as described above, which are not provided in detail for the sake of simplicity.
[0069] This application is intended to cover all such substitutions, modifications and variations that fall within the broad scope of the embodiments of the present invention. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of this application.
Claims
1. A double-point incremental forming manufacturing method based on deep reinforcement learning, characterized in that: The method comprises: Acquire a three-dimensional model to be manufactured and perform layering on the three-dimensional model to acquire a plurality of main working paths and a plurality of candidate supporting paths corresponding to each of the main working paths, and select an initial main working path and a supporting path corresponding to the initial main working path as an initial current main working path and a current supporting path according to a forming direction; According to the current main working path and the selected current supporting path, the mechanical arms of the master and slave robots are respectively controlled cyclically in the actual application environment to perform incremental forming, and a forming surface corresponding to the current main working path is obtained; Before the mechanical arms of the master and slave robots are respectively cyclically controlled to perform incremental forming in the actual application environment according to the current main working path and the selected current support path, the method includes: constructing a digital simulation environment in Grasshopper that matches the actual application environment of the three-dimensional model to be manufactured; simulating the three-dimensional model in the digital simulation environment, training the deep reinforcement learning model in combination with the simulation results, and obtaining the pre-trained deep reinforcement learning model; The three-dimensional model is simulated in the digital simulation environment, and the deep reinforcement learning model is trained in combination with the simulation results to obtain the pre-trained deep reinforcement learning model, including: selecting an initial main working path as the current simulation main working path according to the forming direction, and randomly selecting one of the candidate multiple support paths as the initial current simulation support path according to the current simulation main working path; applying the digital simulation environment to perform simulation forming according to the current simulation main working path and the current simulation support path, and obtaining a simulated forming surface and a simulated forming surface springback value corresponding to the current simulation main working path; using the deviation value between the simulated forming surface and the target surface as a state vector to input into the deep reinforcement learning model. The support strategy is reinforced by the model, and the simulation support path and the current reward value corresponding to the next simulation main working path are updated in combination with the springback value of the simulation forming surface; the current simulation main working path and the current simulation support path are cyclically updated according to the next simulation main working path and the corresponding simulation support path, and the robot arm is cyclically controlled to perform incremental forming according to the updated current simulation main working path and the current simulation support path, and the simulation forming surface is cyclically updated; the state vector is cyclically updated according to the updated simulation forming surface and the target surface, and the model parameters of the deep reinforcement learning model are adjusted according to the updated state vector and the reward value until the convergence condition of the deep reinforcement learning model is met; The deviation value between the forming surface and the target surface is used as the state vector, and a pre-trained deep reinforcement learning model is applied to perform reinforcement learning of the support strategy. The support path corresponding to the next main working path is cyclically output, and the current main working path and the current support path are cyclically updated according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.
2. The method according to claim 1, characterized in that: The step of obtaining a three-dimensional model to be manufactured and layering the three-dimensional model to obtain a plurality of main working paths and a plurality of candidate supporting paths corresponding to each of the main working paths includes: Acquire a three-dimensional model to be manufactured, and apply a surface offset function to layer the three-dimensional model along a forming direction with a preset layer thickness to obtain a first preset number of curved paths; For each of the curved paths, a second preset number of discrete points are divided according to a preset point spacing, and a main working path corresponding to the curved path is generated according to the discrete points; For each of the main working paths, multiple candidate support paths corresponding to the main working path are obtained according to multiple support strategies, and the support strategy is one of a global support strategy, a local peripheral support strategy, a local frontal support strategy and a following support strategy.
3. The method according to claim 1, characterized in that: The step of applying the digital simulation environment to perform simulation forming according to the current simulation main working path and the current simulation supporting path, and obtaining a simulation forming curved surface and a springback value of the simulation forming curved surface corresponding to the current simulation main working path, comprises: Converting the coordinates and directions of the discrete points of the current simulation main working path and the current simulation support path into robot motion instructions according to robot grammar rules; A simulation model of sheet deformation is constructed using simulation software, and simulation forming is performed according to the robot motion instructions, and a simulation forming surface and a simulation forming surface springback value corresponding to the current simulation main working path are returned.
4. The method according to claim 1, characterized in that: The method uses the deviation value between the simulated forming surface and the target surface as a state vector to input the deep reinforcement learning model to perform reinforcement learning of the support strategy, and combines the springback value of the simulated forming surface to update the simulated support path corresponding to the next simulation main working path and the current reward value, including: Acquire each second reference point on the simulated forming surface that corresponds to each first reference point on the target surface, and calculate the error value between each second reference point and the corresponding first reference point to form the state vector; Inputting the state vector into the deep reinforcement learning model to perform reinforcement learning of the support strategy, and outputting a simulation support path corresponding to the next simulation main working path; The current return value is updated according to the simulated forming surface and the springback value of the simulated forming surface.
5. The method according to claim 4, characterized in that: The obtaining the current return value according to the simulated forming surface and the springback value of the simulated forming surface includes: The initial value of the current return value is 0, and if the springback value of the simulated forming surface is greater than or equal to the reference value, the current return value is controlled to decrease by a first preset value; If the simulated forming surface springback value is less than the reference value, controlling the current feedback value to increase by a first preset value; If the simulated curved surface forming fails, the current return value is controlled to be a second preset value.
6. A double-point incremental forming manufacturing device based on deep reinforcement learning, characterized in that: The device comprises: a path acquisition unit, used to acquire a three-dimensional model to be manufactured and to perform layering on the three-dimensional model to acquire a plurality of main working paths and a plurality of candidate supporting paths corresponding to each of the main working paths, and to select an initial main working path and a supporting path corresponding to the initial main working path as an initial current main working path and a current supporting path according to a forming direction; An incremental forming unit, used to cyclically control the mechanical arms of the master and slave robots to perform incremental forming in an actual application environment according to the current main working path and the current supporting path, and obtain a forming surface corresponding to the current main working path; Before the mechanical arms of the master and slave robots are respectively cyclically controlled to perform incremental forming in the actual application environment according to the current main working path and the selected current support path, the method includes: constructing a digital simulation environment in Grasshopper that matches the actual application environment of the three-dimensional model to be manufactured; simulating the three-dimensional model in the digital simulation environment, training the deep reinforcement learning model in combination with the simulation results, and obtaining the pre-trained deep reinforcement learning model; The three-dimensional model is simulated in the digital simulation environment, and the deep reinforcement learning model is trained in combination with the simulation results to obtain the pre-trained deep reinforcement learning model, including: selecting an initial main working path as the current simulation main working path according to the forming direction, and randomly selecting one of the candidate multiple support paths as the initial current simulation support path according to the current simulation main working path; applying the digital simulation environment to perform simulation forming according to the current simulation main working path and the current simulation support path, and obtaining a simulated forming surface and a simulated forming surface springback value corresponding to the current simulation main working path; using the deviation value between the simulated forming surface and the target surface as a state vector to input into the deep reinforcement learning model. The support strategy is reinforced by the model, and the simulation support path and the current reward value corresponding to the next simulation main working path are updated in combination with the springback value of the simulation forming surface; the current simulation main working path and the current simulation support path are cyclically updated according to the next simulation main working path and the corresponding simulation support path, and the robot arm is cyclically controlled to perform incremental forming according to the updated current simulation main working path and the current simulation support path, and the simulation forming surface is cyclically updated; the state vector is cyclically updated according to the updated simulation forming surface and the target surface, and the model parameters of the deep reinforcement learning model are adjusted according to the updated state vector and the reward value until the convergence condition of the deep reinforcement learning model is met; A reinforcement learning unit is used to use the deviation value between the forming surface and the target surface as a state vector, apply a pre-trained deep reinforcement learning model to perform reinforcement learning of the support strategy, cyclically output a support path corresponding to the next main working path, and cyclically update the current main working path and the current support path according to the next main working path and the support path corresponding to the next main working path until the incremental forming of the three-dimensional model is completed.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A computer storage medium, characterized in that: The storage medium stores at least one executable instruction, and the executable instruction enables the processor to execute the method as described in any one of claims 1-5.
Citation Information
Patent Citations
AUV (Autonomous Underwater Vehicle) three-dimensional path planning method based on reinforcement learning
CN109540151A
Salamander robot path tracking hierarchical control method based on reinforcement learning
CN111552301A
Cited By
Design and manufacturing method and system of low-temperature valve bidirectional damping regulation surface microstructure
CN122712764A