Mold machining path optimization method and device based on reinforcement learning and storage medium

By using a reinforcement learning-based method for optimizing mold machining paths, the problem of optimizing machining paths for complex curved surfaces and high-precision molds in existing technologies has been solved. This method achieves adaptive optimization of machining paths, improving cutting stability and machining quality.

CN122047660APending Publication Date: 2026-05-15GUIZHOU RADIO & TV UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUIZHOU RADIO & TV UNIV
Filing Date
2026-04-20
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing mold processing path optimization methods are insufficient to handle complex curved surface structures and high precision requirements, resulting in fluctuations in cutting load and affecting the consistency of processing quality.

Method used

A reinforcement learning-based approach is adopted. By acquiring the mold processing task description data, the processing path is initialized and deconstructed to generate an initial processing path unit set. Action exploration is carried out in the reinforcement learning interactive environment, and the strategy is optimized based on the real-time processing status feedback to generate a processing path that meets the accuracy requirements.

Benefits of technology

It achieves adaptive optimization of mold processing path, improves cutting stability and surface finish, and increases processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047660A_ABST
    Figure CN122047660A_ABST
Patent Text Reader

Abstract

The invention provides a mold machining path optimization method and device based on reinforcement learning and a storage medium, and belongs to the technical field of intelligent manufacturing. The method comprises the following steps: acquiring an initial processing task data set comprising a three-dimensional geometric structure, process sequence constraints and precision requirements; carrying out deconstruction processing on the set to generate an initial processing path unit set which corresponds to each process and has adjustable allowance; inputting the unit set into a reinforcement learning environment, and generating a preliminary processing action strategy set through action exploration; executing the strategy and collecting real-time processing state feedback data, and calculating a processing quality reward value according to the data; evaluating the advantages and disadvantages of the strategy according to the reward value, and generating path optimization gradient information used for indicating the adjustment direction of each adjustable node; and finally, iteratively updating the strategy based on the gradient information until an optimization strategy set meeting the processing precision requirement is obtained. According to the method, adaptive optimization of the processing path is realized by introducing the adjustable nodes and reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and in particular to a method, apparatus and storage medium for optimizing mold processing paths based on reinforcement learning. Background Technology

[0002] In the field of mold processing technology, the planning and optimization of machining paths are among the core factors determining machining quality. A reasonable machining path can effectively improve the stability of the cutting process and enhance the surface quality of the machined material. Currently, common methods for optimizing mold machining paths typically involve obtaining the geometric design model and processing requirements of the mold to be processed. Process programmers then manually create the machining path based on experience or generate an initial machining path using fixed path templates in computer-aided manufacturing software. Individual parameters in the path are then locally adjusted through a limited number of simulations. However, this path optimization method, which relies on static preset rules and manual experience adjustments, struggles to respond promptly to dynamically changing cutting stress states and other factors during the machining process, especially when dealing with mold machining tasks with complex curved surfaces and high precision requirements. This leads to fluctuations in cutting load during actual execution of the generated machining path, affecting the consistency of the mold's surface quality. Summary of the Invention

[0003] This invention provides a method, apparatus, and storage medium for optimizing mold processing paths based on reinforcement learning.

[0004] In a first aspect, embodiments of the present invention provide a mold processing path optimization method based on reinforcement learning, the method comprising: Obtain the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirements information corresponding to each processing step. The initial processing task description data set is processed by processing path initialization deconstruction to generate an initial processing path unit set with adjustable margin corresponding to each processing step. The initial set of processing path units is input into the reinforcement learning interactive environment for processing action exploration, generating a preliminary set of processing action strategies for the initial set of processing path units; In a reinforcement learning interactive environment, a set of preliminary processing action strategies is executed and corresponding real-time processing status feedback data is collected. Based on the real-time processing status feedback data, the processing quality reward value corresponding to the set of preliminary processing action strategies is calculated. The initial processing action strategy set is evaluated based on the processing quality reward value to generate path optimization gradient information to indicate the direction of strategy adjustment. Based on the path optimization gradient information, the initial processing action strategy set is iteratively updated to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

[0005] Secondly, embodiments of the present invention provide a mold processing path optimization device, comprising: The data acquisition module is used to acquire the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirement information corresponding to each processing step. The data deconstruction module is used to perform processing path initialization deconstruction on the initial processing task description data set, and generate an initial processing path unit set with adjustable margins corresponding to each processing step; The reinforcement learning module is used to input the initial set of processing path units into the reinforcement learning interactive environment for processing action exploration and to generate a preliminary set of processing action strategies for the initial set of processing path units. The feedback calculation module is used to execute the initial processing action strategy set in the reinforcement learning interactive environment and collect the corresponding real-time processing status feedback data, and calculate the processing quality reward value corresponding to the initial processing action strategy set based on the real-time processing status feedback data. The strategy evaluation module is used to evaluate the merits of the initial processing action strategy set based on the processing quality reward value, and generate path optimization gradient information to indicate the direction of strategy adjustment. The iterative update module is used to perform iterative update of the preliminary processing action strategy set based on path optimization gradient information, so as to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

[0006] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the methods described above.

[0007] The embodiments of this application have the following beneficial effects: This invention acquires initial machining task description data containing three-dimensional geometry, process sequence constraints, and accuracy requirements, and performs initial machining path deconstruction to generate an initial machining path unit set with adjustable margins, providing a dynamically adjustable basic path structure for subsequent optimization. The initial machining path unit set is input into a reinforcement learning interactive environment to explore machining actions and generate a preliminary machining action strategy set. While executing this strategy set, real-time machining status feedback data is collected, and machining quality reward values ​​are calculated based on this feedback data, achieving immediate quantitative evaluation of the strategy execution effect. The merits of the strategies are evaluated based on the machining quality reward values ​​to generate path optimization gradient information indicating the adjustment direction. The strategy set is iteratively updated based on this gradient information until a machining path optimization strategy set that meets the accuracy requirements is obtained. This invention, through a cyclical mechanism of action exploration, status feedback, reward calculation, gradient generation, and strategy iteration within a reinforcement learning framework, enables the machining path to adaptively adjust and continuously optimize based on real-time feedback machining quality indicators, effectively improving the cutting stability, surface finish, and machining efficiency of the machining path, and realizing intelligent autonomous optimization of mold machining paths. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the structure of the computer system provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the mold processing path optimization method based on reinforcement learning provided in the embodiments of this application.

[0009] Figure 3 This is a schematic diagram of the mold processing path optimization device provided in the embodiments of this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] The embodiments of the present invention are applied to computer systems, such as servers, tablet computers, laptop computers, desktop computers, etc., but are not limited thereto, and no limitations are imposed on the embodiments of this application.

[0012] The computer system implementing the reinforcement learning-based mold processing path optimization method provided in the embodiments of this application will be described next. See also Figure 1 , Figure 1 This is a schematic diagram of the structure of the computer system provided in the embodiments of this application. Figure 1The computer system shown includes at least one processor 210, memory 250, at least one network interface 220, and an external interface 230. The various components in the computer system 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 240.

[0013] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0014] External interface 230 may include, for example, one or more speakers and / or one or more visual displays. External interface 230 may also include one or more input devices 432, such as a keyboard, mouse, microphone, touch screen display, camera, etc.

[0015] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.

[0016] The memory 250 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.

[0017] In some embodiments, memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0018] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 252 is used to reach external devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 253 is configured to enable the display of information (e.g., external interface for operating peripheral devices and displaying content and information) via one or more output devices 231 (e.g., display screen, speaker, etc.) associated with external interface 230; The input processing module 254 is used to detect and translate one or more user inputs or interactions from one or more input devices 232.

[0019] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A mold processing path optimization device 255 stored in memory 250 is shown, which can be software in the form of programs and plug-ins.

[0020] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the skill release device in the virtual scene provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the mold processing path optimization method based on reinforcement learning provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0021] The following describes the mold processing path optimization method based on reinforcement learning provided in the embodiments of this application. In practical implementation, the mold processing path optimization method based on reinforcement learning provided in the embodiments of this application can be implemented by a computer system. See also Figure 3 , Figure 3 This is a flowchart illustrating the reinforcement learning-based mold processing path optimization method provided in this application embodiment. Next, we will combine... Figure 3 The steps shown are explained.

[0022] Step S100: Obtain the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirements information corresponding to each processing step.

[0023] Specifically, the digital 3D model file of the mold to be processed can be read from the database of the computer system, and the 3D geometric structure description information can be obtained from it. The 3D geometric structure description information includes the mold surface contour defined in the form of non-uniform rational B-spline surfaces, the topological connection relationship between each surface, the overall size boundary of the mold, and the geometric parameters of internal features, including grooves, holes, and bosses. Additionally, the processing sequence constraint information can be obtained from the system. This information is represented by a directed graph data structure, where each node in the directed graph represents a processing step, and the directed edges between nodes indicate the order in which the steps are executed. For example, a roughing node pointing to a semi-finishing node indicates that roughing must be completed before semi-finishing. Furthermore, the processing accuracy requirements corresponding to each processing step can be extracted from the process database. This information can exist in the form of a structured data table, where each row of the data table corresponds to a processing step. The column fields include the dimensional tolerance grade, surface roughness parameter, shape tolerance, and position tolerance required for the step. The surface roughness parameter is the arithmetic mean deviation value of the profile, the shape tolerance includes straightness and flatness, and the position tolerance includes parallelism and perpendicularity. The above three-dimensional geometric structure description information, machining process sequence constraint information, and machining accuracy requirements corresponding to each machining process together form the initial machining task description data set. This set can be stored in the system as a comprehensive data container. The data inside the container is organized in the form of key-value pairs, where the key is the data type identifier and the value is the corresponding specific data content.

[0024] Step S200: Perform processing path initialization destructuring on the initial processing task description data set to generate an initial processing path unit set with adjustable margins corresponding to each processing step.

[0025] In one implementation, step S200 may specifically include the following steps S210 to S260: Step S210: Parse the three-dimensional geometric structure description information in the initial processing task description data set, and extract the surface contour surface topology features of the mold to be processed. The surface contour surface topology features include the curvature change gradient distribution of the surface and the adjacency relationship graph structure between the surfaces.

[0026] When parsing the 3D geometric structure description information, all surface elements in the description information are traversed. For each surface, its mathematical definition is obtained. For non-uniform rational B-spline surfaces, its control point mesh, node vectors, and weighting factors are extracted. By calculating the first and second partial derivative vectors of the surface at each point in the parameter domain, the curvature value at that point is solved. The rate of change of curvature value along the surface parameter direction is defined as the curvature gradient. The curvature gradient of all sampling points on the surface is calculated, and the gradient values ​​of the sampling points constitute the curvature gradient distribution of the surface. Simultaneously, the boundary connection relationships between surfaces are analyzed, identifying two surfaces sharing the same boundary curve. An undirected edge is established between them. The attributes of this edge include the geometric type of the shared boundary curve, which is distinguished as a smooth connection or a sharp-corner connection, and the parameter range of the boundary curve in its respective surface parameter domain. All surfaces are treated as nodes, and the connections between all surfaces are treated as edges. Together, they form an adjacency graph that represents the topological structure of the mold surface. This graph is stored in memory using an adjacency list data structure. The nodes of the graph record the unique identifier of the surface and its curvature gradient distribution data, and the edges of the graph record the connection relationship between adjacent surfaces.

[0027] Step S220: Divide the surface contour topology features into processing regions according to the processing sequence constraint information to obtain a sequence of processing region sub-blocks arranged in the order of process execution, with each processing region sub-block corresponding to one processing process.

[0028] Based on the directed graph structure defined in the processing sequence constraint information, the processing objects and scope of each process are determined. The adjacency graph extracted in step S210 is mapped to the process information. For each process, the set of surfaces to be processed in that process is selected from the adjacency graph. The selection principle is based on the predefined correspondence between processing features and surfaces in the process knowledge base. For example, a groove feature may be composed of multiple adjacent surfaces, which form a connected subgraph in the adjacency graph. All surfaces and their adjacency relationships within this connected subgraph are extracted to obtain a processing region sub-block. According to the node order of the directed graph in the process sequence constraint information, each processing region sub-block is sorted to obtain a sequence of processing region sub-blocks arranged according to the process execution order. Each processing region sub-block in this sequence not only contains surface geometric information but also is associated with the corresponding processing accuracy requirements, such as the surface roughness and dimensional tolerances that the process needs to achieve.

[0029] Step S230: For each processing area sub-block, generate a tool contact path dot matrix based on the processing accuracy requirement information corresponding to the processing procedure. The spacing between adjacent path points in the tool contact path dot matrix is ​​adjusted according to the surface roughness index in the processing accuracy requirement information.

[0030] In one implementation, step S230 may specifically include the following steps S231 to S236: Step S231: Extract the surface parameter equations of the processing area sub-blocks, perform isoparametric line sampling on the surface parameter equations, and obtain the initial set of sampling point coordinates distributed along the U and V directions of the surface.

[0031] The parametric equations of the surfaces contained within the sub-blocks of the processing region are extracted from their data structure. For non-uniform rational B-spline surfaces, the parametric equations are represented as vector functions with u and v as parameters. Within the parameter domain of the surface, the values ​​of u and v can range from 0 to 1. A parametric grid is formed by dividing the surface along the u-direction and the v-direction with fixed parameter increments. Each intersection of the parametric grid corresponds to a set of parameter values. These values ​​are substituted into the parametric equations of the surface to calculate the corresponding three-dimensional spatial coordinates. All calculated three-dimensional spatial coordinates constitute an initial set of sampled point coordinates distributed along the U and V directions of the surface. This set is organized as a two-dimensional array, where the row index corresponds to the parameter position in the u-direction and the column index corresponds to the parameter position in the v-direction.

[0032] Step S232: Analyze the machining accuracy requirements information corresponding to the machining process, obtain the maximum allowable residual height value corresponding to the surface roughness index, and calculate the allowable range of chord height error between adjacent tool contact points based on the maximum allowable residual height value.

[0033] The machining accuracy requirements are analyzed, and the surface roughness index, namely the arithmetic mean deviation of the profile, is extracted. Based on the empirical conversion relationship between the arithmetic mean deviation and the residual height, the maximum allowable residual height is determined; this conversion relationship is stored in the process parameter knowledge base. There is a geometric correspondence between the residual height and the chord height error between adjacent tool contacts. The chord height error refers to the maximum perpendicular distance between a straight line segment connecting adjacent tool contacts and the theoretical curved surface. Using the geometry of the tool, such as a ball end mill or flat end mill, and the residual height value, the allowable range of chord height error can be derived through geometric calculations. This calculation process is based on the principle of circular arc approximation, simplifying the contact area between the tool and the workpiece as an arc. The corresponding chord length is calculated based on the arc radius and the residual height; the chord height corresponding to this chord length is the allowable range of chord height error.

[0034] Step S233: Based on the allowable range of chord height error, perform adaptive point cloud density adjustment processing on the initial sampling point coordinate set. Increase the sampling point density in areas with large surface curvature gradient changes and decrease the sampling point density in areas with small surface curvature gradient changes to obtain a density-optimized sampling point coordinate set.

[0035] The initial set of sample point coordinates is traversed. For every two adjacent sample points, the actual chord height error between the straight line segment connecting them and the theoretical surface is calculated. When calculating the actual chord height error, multiple intermediate points can be taken on the parameter line between the two sample points. The distances from these intermediate points to the straight line segment are calculated, and the maximum value is taken as the actual chord height error. This actual chord height error is compared with the allowable range of chord height error obtained in step S232. If the actual chord height error exceeds the preset allowable range, it indicates that the current sample point density is insufficient, and a new sample point needs to be inserted on the parameter line between the two points. The position of the new point is determined by linear parameter interpolation, that is, the parameter values ​​of the two points are averaged to obtain the parameter value of the new point, which is then substituted into the surface equation to calculate the coordinates. If the actual chord height error is much smaller than the lower limit of the allowable range, and the surface curvature change gradient is small, then it is possible to consider deleting intermediate sample points to simplify the path. After performing the above judgment and adjustment on all adjacent sample points, a density-optimized set of sample point coordinates is obtained.

[0036] Step S234: Based on the boundary curve equation of the processing area sub-block, perform spatial position mapping processing on the density-optimized set of sampling point coordinates to generate the distance parameter and orientation parameter of each sampling point relative to the boundary curve.

[0037] The boundary curve equation of the processing area sub-block is extracted; this boundary curve is the contour line of the processing area. For each sampling point in the density-optimized sampling point coordinate set, its shortest distance to the boundary curve is calculated. This distance is obtained by solving the problem of finding the shortest distance from a point to a curve, specifically using an iterative method, such as Newton's iteration method. The point on the boundary curve closest to the sampling point is found, and then the Euclidean distance between the two points is calculated; this distance value is the distance parameter. Simultaneously, the orientation of the sampling point relative to the boundary curve is determined. The orientation parameter is determined by judging which side of the boundary curve the sampling point is located on. For example, for a closed boundary, it can be determined whether the sampling point is inside or outside the boundary; for an open boundary, it can be determined whether the sampling point is on the left or right side of the boundary. The orientation parameter is represented in the form of enumerated values, such as inside, outside, left, right, etc. The distance parameters and orientation parameters of all sampling points constitute the spatial position mapping information of that point.

[0038] Step S235: Arrange the coordinates of the sampling points after spatial position mapping according to the parameter order of the U and V directions to generate a tool contact path matrix with topological connection relationship.

[0039] Based on the parameter mesh established in step S231, the sampling points, after density optimization and spatial location mapping, are refilled into the mesh. Since sampling points may have been added or deleted in step S233, the original regular mesh structure may be disrupted; therefore, it is necessary to re-establish the topological connections between the sampling points. Specifically, for each parameter value in the u direction, all corresponding sampling points in the v direction are sorted in ascending order of the v parameter value, forming a row sequence. For each parameter value in the v direction, all corresponding sampling points in the u direction are sorted in ascending order of the u parameter value, forming a column sequence. Through this double sorting, the row and column indices of each sampling point in the point matrix are determined, and adjacency relationships are naturally established between sampling points in adjacent rows and adjacent columns, thereby generating a knife-contact path point matrix with topological connections.

[0040] Step S236: Sort the tool contact point coordinates in the tool contact point path matrix with topological connection relationship according to the machining tool path direction to obtain the tool contact point path matrix arranged in sequence along the machining tool path direction.

[0041] Obtain the machining direction set in the process plan, such as reciprocating or unidirectional toolpath. Based on the tool contact path point matrix with topological connections in step S235, traverse and sort the points in the matrix according to the toolpath direction. For reciprocating toolpath, the first row takes points sequentially in the direction of increasing v parameter, the second row takes points in the opposite direction of decreasing v parameter, and so on, forming a continuous, back-and-forth path. For unidirectional toolpath, each row takes points in the same direction. After machining each row, the tool is lifted to the starting position of the next row and machining continues in the same direction. Through this sorting method, all tool contacts are connected into one or more continuous trajectories. Each tool contact has a unique predecessor node and successor node in the trajectory, ultimately resulting in a tool contact path point matrix arranged sequentially along the machining direction.

[0042] Step S240: Optimize the path connection order of the tool contact path point matrix to generate an initial path segment set connecting all tool contact path points, and reserve at least one path adjustment node in each initial path segment as an adjustable margin for subsequent processing actions.

[0043] In one implementation, step S240 may specifically include the following steps S241 to S246: Step S241: Extract the spatial three-dimensional coordinates of all tool contacts in the tool contact path point matrix, calculate the Euclidean distance between every two adjacent tool contacts as the candidate length of the path segment, and sort the candidate lengths of the path segment according to the machining tool path direction to obtain the candidate sequence of path segments along the machining tool path direction.

[0044] Read the spatial three-dimensional coordinates of each tool contact point from the tool contact path point matrix. Following the tool movement sequence determined in step S236, pair adjacent tool contacts together. For each pair of adjacent tool contacts, calculate the Euclidean distance between them using the three-dimensional distance formula (refer to the general formula). The calculated distance value serves as the candidate length for that path segment. Arrange all path segments according to their order of appearance on the tool movement path, forming a sequence. Each element in the sequence records a path segment, including its starting coordinates, ending coordinates, and the calculated candidate length. This sequence is the candidate sequence of path segments along the machining tool movement direction.

[0045] Step S242: For each candidate path segment in the candidate path segment sequence, determine whether the path segment crosses the restricted area within the processing area sub-block. If it crosses the restricted area, remove the path segment from the candidate path segment sequence and reconnect the tool contacts at both ends of the path segment to other tool contacts that bypass the restricted area.

[0046] Obtain the restricted area data defined within the processing area sub-block. Restricted areas are typically described as closed curves on spatial polyhedra or parametric surfaces. For each path segment in the candidate sequence, discretize it into a series of dense points, and then determine if any of these points are located inside a restricted area. The determination of whether a point is inside a restricted area is made using a spatial ray method or a method based on the signed distance field. If a point is found to be inside a restricted area, the path segment is determined to cross the restricted area and is removed from the candidate sequence. Subsequently, for the two tool contact points at both ends of the path segment, the connection path between them needs to be replanned. Find two points on the boundary of the restricted area, respectively serving as connection points from the start point to the boundary and from the boundary to the end point, such that the three paths from the start point to the first boundary point, from the first boundary point to the second boundary point, and from the second boundary point to the end point do not cross the restricted area, and the total path length is as short as possible. In this way, the original direct connection is changed to a polyline connection that bypasses the restricted area.

[0047] Step S243: Connect adjacent path segments in the candidate sequence of path segments after the restricted area removal process to generate an initial continuous path trajectory covering all knife contacts in the knife contact path point matrix.

[0048] After step S242, each path segment in the candidate path segment sequence no longer crosses the restricted area. These path segments are then connected end-to-end according to their order in the sequence, meaning the endpoint coordinates of the previous path segment and the starting coordinates of the next path segment should be the same tool contact point. If a tool contact point generates a new branch path by bypassing the restricted area, the branch path needs to be integrated into the main path according to the topological relationship. This ultimately forms a continuous path trajectory starting from the first tool contact point, passing through all tool contacts in sequence, and finally reaching the last tool contact point. This trajectory consists of a series of end-to-end connected straight line segments, each corresponding to a machining motion, and the entire trajectory covers all tool contacts in the tool contact point matrix.

[0049] Step S244: Mark a path adjustment node at every preset number of tool contact points in the initial continuous path trajectory, and set the spatial coordinates of the tool contact points marked as path adjustment nodes to an adjustable state, allowing subsequent machining actions to offset these coordinates.

[0050] On the generated initial continuous path trajectory, tool contacts are selected as path adjustment nodes at fixed intervals. For example, one tool contact can be selected as an adjustment node every five tool contacts, or the interval can be dynamically adjusted according to the complexity of the path, decreasing the interval in areas with large curvature changes to increase the density of adjustment nodes. Selected tool contacts are marked as path adjustment nodes, and a status flag is added to their corresponding data structure, set to "adjustable". Simultaneously, storage space is expanded for this node to record possible subsequent coordinate offsets. Unselected tool contacts remain as fixed nodes, their coordinates remaining unchanged during subsequent optimization. Based on this, the path, while maintaining a fixed overall structure, possesses a certain degree of local elastic deformation capability.

[0051] Step S245: Based on the initial continuous path trajectory and the marked positions of the path adjustment nodes, generate an initial path segment set consisting of fixed path segments and adjustable path nodes.

[0052] The initial continuous path trajectory generated in step S243 is merged with the path adjustment node information marked in step S244. Two types of elements are distinguished in the trajectory data: one type consists of line segments connecting two fixed nodes or connecting a fixed node and an adjustment node; these are called fixed path segments, and their geometry is determined in the initial state. The other type consists of path nodes marked as adjustable; these nodes are themselves components of the trajectory, but their coordinate values ​​are set as variables. The entire trajectory data structure is reorganized, using path adjustment nodes as key dividing points to divide the trajectory into multiple segments. Within each segment, if both the starting and ending points are fixed nodes, the segment is completely fixed; if the starting or ending point contains an adjustment node, the position of the adjustment node within that segment can change within a certain range, thus affecting the shape of the segment. This organization method forms an initial path segment set composed of fixed path segments and adjustable path nodes.

[0053] Step S246: Based on the spatial relationship between fixed path segments and adjustable path nodes in the initial path segment set, construct a processing path node topology with adjustable degrees of freedom.

[0054] The initial set of path segments is abstracted into a graph data structure. Nodes in the graph are divided into two categories: fixed nodes, corresponding to fixed contact points, and adjustable nodes, corresponding to path adjustment points. Edges in the graph correspond to path segments, connecting two nodes. For any edge, if both ends are fixed nodes, the edge is a fixed edge; if at least one of the ends is an adjustable point, the edge is a flexible edge, and its spatial shape changes with the coordinates of the adjustable point. During graph construction, the spatial coordinates, node type, and information of all edges connected to each node are recorded. For adjustable points, their allowed offset range is additionally recorded, typically represented by a spatial sphere or ellipsoid centered on the node.

[0055] Step S250: Associate and store the initial set of path segments with the spatial coordinate data of the path adjustment nodes to generate a directed graph structure containing the position coordinates of the path adjustment nodes and the connection relationships between the path adjustment nodes.

[0056] The geometric data in the initial path segment set generated in step S245 and the coordinate data of the path adjustment nodes marked in step S244 are integrated into a unified data structure. This data structure adopts the form of a directed graph because the machining path has a clear direction of movement. The nodes of the directed graph store the spatial three-dimensional coordinates of the path adjustment nodes. Each node also has an attribute list recording information such as the node's sequence number in the machining path, the process identifier it belongs to, and whether it is a start or end point of the path. The edges of the directed graph store the connection relationship between two nodes. The direction of the edge is consistent with the machining tool path direction, and the edge attributes include the length of the path segment, the preset machining speed value, and whether it passes through the edge of a restricted area. All path adjustment nodes and their connection relationships constitute a complete directed graph, which is stored in computer memory in the form of an adjacency matrix or an adjacency list.

[0057] Step S260: Extract the information of the preceding and following adjacent nodes of each path adjustment node in the processing path based on the directed graph structure, determine the adjustable degree of freedom range of each path adjustment node according to the preceding and following adjacent node information, and obtain the initial processing path unit set with adjustable margin corresponding to each processing operation.

[0058] Traverse the directed graph generated in step S250. For each path adjustment node in the graph, obtain its preceding and following nodes in the processing path by querying its incoming and outgoing edges. The preceding node may be a fixed node or another adjustment node, and the same applies to the following node. Based on the spatial positions of the preceding and following adjacent nodes, the adjustable degree of freedom range of the current adjustment node can be calculated. Determining the degree of freedom range requires considering several constraints: First, the adjustment node cannot move too far from the straight line connecting its preceding and following adjacent nodes to avoid excessive chord height error; second, the adjustment node cannot cross the preceding and following adjacent nodes, causing the path order to be disordered; finally, the movement of the adjustment node cannot cause the path segment to cross the restricted area. Combining these constraints, a spatial convex polyhedron is calculated for each adjustment node, and any point within this polyhedron can be used as the node's legal new position. Store this adjustable degree of freedom range along with the node's initial coordinates and the information of the preceding and following adjacent nodes to form the initial processing path unit with adjustable margin corresponding to this process. Summarize these units for all processes to obtain the complete set of initial processing path units.

[0059] Step S300: Input the initial processing path unit set into the reinforcement learning interactive environment for processing action exploration and processing, and generate a preliminary processing action strategy set for the initial processing path unit set.

[0060] In this step, a reinforcement learning interactive environment is first constructed, which encapsulates the physical simulation model of mold processing. The initial set of processing path units generated in step S200 is used as the initial state input of the environment. Within the environment, a set of actions that the agent can perform are defined; these actions correspond to modifications to the coordinates of path adjustment nodes. The agent generates new processing paths by trying different actions and changing the positions of the path adjustment nodes. The environment performs simulated processing based on the new paths and calculates a reward value to evaluate the quality of the paths. The agent's goal is to learn a set of action sequences that yield high reward values ​​through continuous trial and error, i.e., how to adjust each path node. After multiple rounds of exploration, the agent outputs a set of action choices for the current input initial state. This set of actions constitutes a preliminary modification scheme for the initial set of processing path units, i.e., the preliminary processing action strategy set.

[0061] In one implementation, step S300 may specifically include the following steps S310 to S360: Step S310: Construct a state space representation in the reinforcement learning interactive environment, using the current spatial coordinates of each adjustable path node in the initial processing path unit set as the basis vector of the state space, and encoding the processing sequence constraint information as the legality constraint conditions for state transition.

[0062] In a reinforcement learning interactive environment, a data structure for the state space is defined. This state space is represented as a multi-dimensional vector, where the dimension of the vector is equal to the total number of adjustable path nodes in the initial set of processing path units multiplied by three, because each node requires three coordinate values ​​to describe its spatial position. For example, if there are N adjustable nodes, the state vector is a floating-point array of length 3N, where the first three elements correspond to the x, y, and z coordinates of the first adjustable node, the next three elements correspond to the second node, and so on. Simultaneously, the processing sequence constraints are encoded and converted into a series of constraint functions regarding node coordinates. These functions exist in the form of Boolean expressions, taking a set of node coordinates as input and outputting whether the coordinate combination satisfies the processing sequence. For example, a constraint might stipulate that a path node belonging to a later process cannot appear before a node belonging to a previous process; this order relationship is achieved by comparing the node's index position or spatial position relative to the path.

[0063] Step S320: Use the coordinate offset direction and offset magnitude of each adjustable path node as the basic motion primitives of the motion space. The offset direction includes the tool lifting direction along the normal of the machining surface and the feed direction along the tangent of the machining surface.

[0064] Define the basic operational units that a reinforcement learning agent can execute, i.e., action primitives. For each adjustable path node, define a set of possible coordinate offset operations. The offset direction is decomposed into two main components: one component is along the normal direction of the machining surface where the node is located, called the tool lift direction, which changes the cutting depth or lifts the tool away from the workpiece surface; the other component is along the tangent direction of the machining surface where the node is located, called the feed direction, which changes the tool's sliding position on the workpiece surface. Each offset direction corresponds to multiple possible offset magnitudes, such as small offset, medium offset, and large offset. Therefore, for each adjustable node, the action space contains multiple discrete action primitives, such as "small offset along the positive normal direction", "medium offset along the negative normal direction", and "large offset along the positive tangent direction". All action primitives of all nodes together constitute the action space of the entire system.

[0065] Step S330: Call the exploration policy network of the reinforcement learning interactive environment to evaluate the action value of the state space representation. Based on the evaluation results, select a set of candidate action primitives in the current state from the action space. Each candidate action primitive corresponds to a coordinate offset operation of an adjustable path node.

[0066] In one implementation, step S330 may specifically include the following steps S331 to S336: Step S331: Normalize the current spatial coordinates of each adjustable path node in the state space representation to generate a normalized state feature vector. The dimension of the state feature vector is consistent with the number of adjustable path nodes.

[0067] The current spatial coordinates of all adjustable path nodes are extracted from the state space representation. These coordinates are raw floating-point numbers. Since the coordinates of different nodes may fall within different numerical ranges (e.g., some nodes at the top of the mold have larger z-coordinates, while others at the bottom have smaller z-coordinates), directly using the raw coordinates would make neural network training difficult. Therefore, the coordinate values ​​need to be normalized. The normalization method involves iterating through all nodes and calculating the mean and standard deviation of each of the x, y, and z coordinate components across the entire set. Then, for each coordinate component of each node, the mean of that component is subtracted, and the result is divided by the standard deviation to obtain the normalized value. The three normalized coordinate values ​​of all nodes are then concatenated in a fixed order into a one-dimensional vector. The length of this vector is three times the number of adjustable path nodes. This vector is the normalized state feature vector.

[0068] Step S332: Input the normalized state feature vector into the feature extraction layer of the exploration policy network, and extract the spatial correlation features between nodes in the state feature vector through multi-layer convolution operations to generate a hidden state feature representation with global context information.

[0069] The normalized state feature vector generated in step S331 is input into the first part of the exploration policy network, namely the feature extraction layer. This feature extraction layer consists of multiple stacked one-dimensional convolutional layers. Since the elements in the state feature vector are coordinate values ​​arranged in node order, one-dimensional convolution can perform sliding convolution operations between the coordinates of adjacent nodes, thereby capturing the spatial correlation between local nodes, such as the coordinate change trend of several adjacent nodes. As the number of convolutional layers increases, the receptive field gradually expands, and subsequent convolutional layers can capture the correlation between nodes at greater distances. Each convolutional layer is followed by an activation function, such as a linear rectified unit, to increase the non-linear expressive power of the network. After multiple convolutional operations, the original state feature vector is transformed into a new feature vector, in which each element contains information about its surrounding and even global nodes. This new feature vector is called the hidden state feature representation.

[0070] Step S333: Input the hidden state feature representation into the action value output layer of the exploration policy network, calculate the expected cumulative reward value of each action primitive in the action space in the current state, and generate the action value estimate corresponding to each action primitive.

[0071] The hidden state feature representation generated in step S332 is input into the second part of the exploration policy network, namely the action value output layer. This output layer typically consists of several fully connected layers. The first fully connected layer maps the hidden state feature representation to an intermediate dimension, then processes it through an activation function before inputting it into the second fully connected layer. The number of output neurons in the last fully connected layer is equal to the total number of action primitives in the action space. For each output neuron, the calculated value is the estimated action value of the corresponding action primitive in the current state. This estimate is a scalar representing the expected cumulative reward that can be obtained by performing the action from the current state and then following the current policy.

[0072] Step S334: Sort the action value estimates, select the top preset number of action primitives with the highest action value estimates as preliminary candidate action primitives, and record the adjustable path node identifier corresponding to each preliminary candidate action primitive.

[0073] After obtaining the estimated action value of all action primitives, these values ​​are sorted from largest to smallest. Based on a preset number of candidates, for example, selecting the top ten highest-value actions each time, the top ten action primitives are extracted from the sorted results. These selected action primitives constitute an initial set of candidate action primitives. For each action primitive in the set, its corresponding adjustable path node identifier is parsed out; this identifier is a unique number used to distinguish different nodes. Simultaneously, the offset direction specified by the action primitive is recorded, such as whether it is a positive normal or a negative tangential direction, and the offset magnitude, such as whether it is small or large. All selected action primitives and their corresponding node identifiers, offset directions, and magnitudes are stored in a temporary candidate list.

[0074] Step S335: Perform conflict detection on the preliminary candidate action primitives. If two or more preliminary candidate action primitives act on the same adjustable path node, retain the one with the highest action value estimate and remove the other conflicting action primitives.

[0075] The initial candidate action primitive list generated in step S334 is traversed, checking if multiple action primitives with the same adjustable path node identifier exist in the list. For example, if both a small-amplitude normal movement and a medium-amplitude tangential movement for node 5 appear in the list, these two action primitives constitute a conflict because they attempt to perform different operations on the same node. For a conflicting node, the action value estimates of all action primitives acting on that node are compared, and the one with the largest value is retained, while the remaining conflicting action primitives are removed from the candidate list. This conflict detection mechanism ensures that in the final candidate set, each adjustable path node has at most one action to be executed, avoiding logical contradictions and execution chaos.

[0076] Step S336: Perform constraint encoding on the preliminary candidate action primitives after conflict detection according to the processing sequence constraints, and generate a set of constraint action primitives that match the processing sequence constraints, which will serve as the final set of candidate action primitives.

[0077] For the initial candidate action primitives retained after conflict detection, further screening is performed using machining process sequence constraints. These constraints are pre-encoded into a series of executable functions. These functions take the current state space and a candidate action primitive as input and output whether the action is allowed. For example, a constraint might stipulate that large movements along the negative normal direction are not allowed in the finishing process, as this could lead to overcutting. For each candidate action primitive, its corresponding constraint function is called for checking. If an action primitive violates any constraint, it is removed from the candidate list. After constraint encoding, the remaining action primitives are all legal actions that have high value estimates and satisfy all machining process constraints. These action primitives constitute the final set of candidate action primitives, awaiting combination and execution.

[0078] Step S340: Combine and encapsulate the selected candidate action primitives to generate an action instruction sequence containing multiple adjustable path node coordinate offset instructions. Each coordinate offset instruction in the action instruction sequence corresponds uniquely to an adjustable path node in the initial processing path unit set.

[0079] The candidate action primitives determined in step S336 are combined. Each candidate action primitive is essentially an operation instruction targeting a specific node. These instructions are encapsulated in a certain logical order to form an action instruction sequence. This sequence is an ordered list, where each element corresponds to a coordinate offset instruction. The data structure of each coordinate offset instruction contains three key fields: the first field is an adjustable path node identifier, used to uniquely identify the target of the instruction; the second field is the offset direction vector, a three-dimensional unit vector pointing in the direction of offset, such as the normal or tangential direction; the third field is the offset magnitude value, a scalar representing the distance to be moved in that direction. All instructions are combined to form a complete action instruction sequence, which will be used to drive path node updates in the simulation environment.

[0080] Step S350: Execute the action instruction sequence through the reinforcement learning interactive environment to generate the coordinate correction amount for each adjustable path node in the initial processing path unit set.

[0081] In one implementation, step S350 may specifically include the following steps S351 to S356: Step S351: Parse each coordinate offset instruction in the action instruction sequence, extract the adjustable path node identifier corresponding to each coordinate offset instruction and its target offset direction and target offset magnitude, where the target offset magnitude is represented as the displacement distance in the target offset direction.

[0082] After receiving a sequence of action instructions, the reinforcement learning interactive environment initiates an instruction parsing process. This process iterates through each instruction in the sequence, breaking down each instruction into its fields. First, it reads the adjustable path node identifier field from the instruction, obtaining an integer value that points to a specific node object stored in memory. Second, it reads the offset direction vector field from the instruction, obtaining an array containing three floating-point numbers representing the direction cosines on the x, y, and z axes, respectively. The magnitude of this vector is one, ensuring that it only represents direction and not magnitude. Finally, it reads the offset magnitude field from the instruction, obtaining a positive floating-point number representing the physical distance to be moved in that direction, in units consistent with the model coordinate system, such as millimeters.

[0083] Step S352: Perform vector superposition calculation on the current spatial coordinates corresponding to the adjustable path node identifier, the target offset direction, and the target offset magnitude to generate the updated spatial coordinates of each adjustable path node.

[0084] Based on the node identifier parsed in step S351, the environment locates the current spatial coordinate data of that node stored in memory. This data is also a floating-point array containing three components: x, y, and z. Then, vector superposition calculation is performed. Each component of the target offset direction vector is multiplied by the target offset magnitude to obtain a displacement vector, which also has three components. The three components of the current spatial coordinates are added to the three components of the displacement vector, i.e., the x-coordinate is added to the x-component of the displacement vector, the y-coordinate is added to the y-component of the displacement vector, and the z-coordinate is added to the z-component of the displacement vector, resulting in three new values. These three new values ​​constitute the updated spatial coordinates of the node. After the calculation is completed, the environment temporarily stores these new coordinates in a buffer and does not immediately overwrite the original data.

[0085] Step S353: Perform coordinate replacement operation on the corresponding adjustable path nodes in the initial processing path unit set according to the updated spatial coordinates, while keeping the coordinates of the fixed path nodes unchanged, to obtain an intermediate processing path unit set with updated path nodes.

[0086] After calculating the updated coordinates of all nodes involved in the instructions in step S352, the environment begins to perform a data update operation. It iterates through all path nodes in the initial processing path unit set, determining the node type for each node. If the node type is an adjustable path node and its identifier is included in the instruction list parsed in step S351, then the node's current spatial coordinates in memory are replaced with the updated coordinates calculated in step S352. If the node type is an adjustable path node but not in the instruction list, its coordinates remain unchanged. If the node type is a fixed path node, its coordinates are forcibly kept unchanged regardless of whether it appears in the instruction list, because fixed nodes cannot be modified. After all nodes are updated, the resulting processing path unit set is the intermediate processing path unit set, in which the positions of some adjustable nodes have changed.

[0087] Step S354: Evaluate the smoothness of the connection relationship between path nodes in the intermediate processing path unit set, detect the curvature change gradient between adjacent updated path nodes, and generate smoothness adjustment coefficients based on the curvature change gradient.

[0088] After obtaining the set of intermediate processing path units, the smoothness of the path needs to be evaluated. Traverse all adjacent path nodes. For each pair of three adjacent nodes, denoted as point A, point B, and point C, where point B is a node that may have been updated. Calculate vectors AB and BC, representing the path's direction before and after point B. Calculate the angle between these two vectors; the angle reflects the degree of inflection at point B. A smaller angle indicates a smoother transition; a larger angle indicates a more abrupt transition. Compare the calculated angle with a preset ideal angle range. If the angle exceeds this range, the path is not smooth enough at that point. Based on the degree of deviation, calculate a smoothness adjustment coefficient, a value between 0 and 1. The greater the deviation of the angle from the ideal range, the closer the coefficient is to 0, indicating a greater need for smoothing; if the angle is within the ideal range, the coefficient is close to 1, indicating good smoothness and no adjustment is needed.

[0089] Step S355: Fine-tune the coordinates of the updated path nodes according to the smoothness adjustment coefficient to keep the path curvature change between adjacent path nodes within the preset smoothness range.

[0090] For the updated nodes identified in step S354 whose path turning angles exceed the ideal range, coordinate fine-tuning is required. The goal of fine-tuning is to reduce the angle between vectors AB and BC. The method involves searching for a new position point B' near the current position of point B, making the angle between vectors AB' and B'C closer to a straight angle. The search range is limited to a spherical region centered on point B with a preset step size as its radius. Within this region, multiple candidate points are sampled at a certain step size, and a new angle is calculated for each candidate point. The candidate point whose angle is closest to the ideal value is selected as B'. A smoothness adjustment coefficient is used to control the search step size and range; a smaller coefficient allows for a larger search range to achieve better smoothness. Through this fine-tuning, the smoothness of the path is improved while minimizing deviation from the original instruction intent.

[0091] Step S356: Integrate the final coordinates of all path nodes in the intermediate processing path unit set after smoothing fine-tuning to generate complete path description data containing updated path node coordinates and fixed path node coordinates as coordinate correction amount.

[0092] After the fine-tuning process in step S355, the coordinates of all adjustable path nodes are finally determined. Now, these final coordinates need to be integrated with the coordinates of the fixed path nodes that have never changed. Following the original order of the processing path, all path nodes are traversed, and for each node, its final spatial coordinates are read from memory. The node's index and corresponding coordinate values ​​are written in pairs into a new data structure. This data structure is an ordered list, and the order of the list is exactly the same as the order in which the processing path is traversed. Each element in the list contains two parts: a unique identifier for the node and the node's three-dimensional coordinates. This complete ordered list is the final result of this action, describing the updated entire processing path.

[0093] Step S360: Update the state space representation based on the coordinate correction to obtain a preliminary set of processing action strategies containing the corrected path nodes.

[0094] The coordinate correction values ​​generated in step S356 are fed back to the state manager of the reinforcement learning interactive environment. The state manager receives this complete path description data and uses it to replace the previously stored state space representation. Specifically, the state vectors originally stored in the state manager, i.e., the list of coordinates of all adjustable path nodes, are replaced one by one with the new coordinates of the corresponding nodes in the coordinate correction values. At the same time, other state-related information, such as the adjacency relationships of nodes, may need to be recalculated due to the change in node coordinates. These geometric attributes, such as the length and direction of path segments, are also updated synchronously in the state manager. After the update is completed, the environment enters a new state, which fully reflects the processing path after executing the action instruction sequence. This new state, together with the action instruction sequence executed to generate this state, constitutes a preliminary set of processing action strategies, which records the transition process from the initial state to the current state.

[0095] Step S400: Execute the initial processing action strategy set in the reinforcement learning interactive environment and collect the corresponding real-time processing status feedback data. Calculate the processing quality reward value corresponding to the initial processing action strategy set based on the real-time processing status feedback data.

[0096] In one implementation, step S400 may specifically include the following steps S410 to S460: Step S410: In the reinforcement learning interactive environment, simulate the cutting motion of the tool according to the corrected path nodes in the initial machining action strategy set, and generate a real-time tool force data sequence during the simulated machining process.

[0097] The machining process simulator in the reinforcement learning interactive environment begins executing the initial machining action strategy. First, it interpolates and generates a continuous tool center point trajectory based on the corrected path node coordinates. Then, the simulator calls its internal cutting force prediction model, which is based on mechanical mechanics theory and considers factors such as tool geometry, workpiece material properties, depth of cut, width of cut, and instantaneous cutting thickness. For each tiny movement step on the trajectory, the model calculates the current contact area between the tool and the workpiece, and solves for the three-dimensional cutting force components acting on the tool based on the geometry of the contact area and the cutting parameters. As the simulation progresses, the force values ​​calculated at each time step are recorded, forming a data sequence arranged in chronological order. This sequence is the real-time tool force data sequence. Each data point in the sequence includes a timestamp at that moment and the force values ​​in the x, y, and z directions.

[0098] Step S420: Analyze the real-time tool force data sequence, extract the tool force peak, force fluctuation amplitude and force average as machining stability evaluation indicators, and compare the machining stability evaluation indicators with the preset stability threshold range to generate stability bonus components.

[0099] In one implementation, step S420 may specifically include the following steps S421 to S426: Step S421: Perform peak detection processing on the real-time tool force data sequence, identify all local maxima in the sequence and extract the maximum value as the tool force peak value, compare the tool force peak value with the preset force peak value safety threshold, and generate a negative reward value if it exceeds the safety threshold.

[0100] The real-time tool stress data sequence is traversed. For each data point in the sequence, its value is compared with the values ​​of its immediate and next neighbors. If a point's value is greater than both its preceding and following neighbors, it is marked as a local maximum. After traversal, a set of all local maxima is obtained. From this set, the point with the largest value is identified; this value is the peak tool stress. This peak value is compared with a preset peak stress safety threshold retrieved from the process knowledge base. If the peak value is less than or equal to the safety threshold, no penalty is applied. If the peak value exceeds the safety threshold, a penalty value is calculated based on the proportion of the exceedance. The larger the proportion, the larger the penalty value. This penalty value is converted into a negative reward value, i.e., the corresponding value is deducted from the total reward.

[0101] Step S422: Calculate the standard deviation of all data points in the real-time tool force data sequence as the force fluctuation amplitude, and compare the force fluctuation amplitude with the preset fluctuation amplitude tolerance threshold. The greater the fluctuation amplitude exceeds the tolerance threshold, the greater the negative reward value generated.

[0102] Calculate the arithmetic mean of all data points in the real-time tool stress data sequence. Then, for each data point in the sequence, calculate the square of the difference between it and the mean. Sum the squares of all differences to obtain a total. Divide this total by the number of data points to obtain the variance. Take the square root of the variance to obtain the standard deviation, which is used as the stress fluctuation amplitude. Compare this calculated fluctuation amplitude with a preset fluctuation amplitude tolerance threshold. If the fluctuation amplitude is less than or equal to the tolerance threshold, no penalty is applied. If the fluctuation amplitude exceeds the tolerance threshold, a penalty value is calculated based on the proportion of the excess; the larger the proportion, the larger the penalty value. This penalty value is also converted into a negative reward value and deducted from the total reward.

[0103] Step S423: Calculate the arithmetic mean of all data points in the real-time tool force data sequence as the force average. Compare the force average with the preset ideal force average range. The greater the distance between the force average and the ideal force average range, the greater the negative reward value generated.

[0104] The arithmetic mean of all data points in the real-time tool stress data sequence is calculated, which is the sum of the force values ​​at all points divided by the total number of points. A preset ideal stress mean range is retrieved from the process knowledge base, defined by a lower limit and an upper limit. The calculated stress mean is compared to this range. If the stress mean falls within the range, no penalty is applied. If the stress mean is less than the lower limit, the difference between the lower limit and the stress mean is calculated; the larger the difference, the larger the penalty. If the stress mean is greater than the upper limit, the difference between the stress mean and the upper limit is calculated; the larger the difference, the larger the penalty. Based on the magnitude of the difference, a negative reward value is generated according to a preset proportional function.

[0105] Step S424: Based on the comparison results of the peak force, force fluctuation amplitude, and average force of the tool with their respective thresholds or intervals, three independent penalty factors are generated, and the three penalty factors are merged according to a preset weight ratio to generate a comprehensive stability penalty value.

[0106] The three negative reward values ​​generated from the comparison results in steps S421, S422, and S423 are treated as three independent penalty factors, denoted as P1, P2, and P3. Each penalty factor is a non-negative numerical value. Three preset weight coefficients, corresponding to the relative importance of peak value, fluctuation, and average value, are obtained from the process knowledge base and denoted as W1, W2, and W3, respectively. The sum of these three weight coefficients is one. A combined calculation is performed: P1 is multiplied by W1, P2 by W2, and P3 by W3. The three products are then added together, and the sum is a comprehensive stability penalty value. This comprehensive penalty value comprehensively reflects the overall performance of the processing process across multiple stability dimensions; the larger the penalty value, the worse the stability.

[0107] Step S425: Convert the stability penalty value into a stability reward component. The stability reward component is negatively correlated with the stability penalty value, that is, the larger the stability penalty value, the smaller the stability reward component.

[0108] Since reward functions are typically designed so that a larger reward value indicates better performance, while a larger overall stability penalty value calculated in step S424 indicates worse stability, it is necessary to convert the penalty value into a reward component. This conversion is achieved using a monotonically decreasing function. For example, a base reward value can be set, and then the overall stability penalty value can be subtracted from it. Alternatively, a negative exponential function can be used, such that the larger the penalty value, the smaller the converted reward component, approaching zero. After this conversion, when the original penalty value was large, the corresponding stability reward component will be very small; while when the penalty value is zero, the corresponding stability reward component will reach its maximum value. This achieves a positive correlation between the reward value and processing stability.

[0109] Step S426: Map the stability reward component to the range of standard reward values ​​to obtain the normalized stability reward component.

[0110] The stability reward component obtained after the transformation in step S425 may have an uncertain numerical range, which may vary depending on different processing tasks. To facilitate subsequent weighted summation with other reward components, it needs to be mapped to a unified standard range, such as between 0 and 1. Normalization is then performed. First, the possible value range of the stability reward component is determined, with the maximum value corresponding to a penalty of zero and the minimum value corresponding to a penalty reaching the theoretical upper limit. Then, a linear normalization method is used: the calculated stability reward component is subtracted from the possible minimum value, and then divided by the difference between the possible maximum and minimum values. After this calculation, the original stability reward component is mapped to a closed interval between 0 and 1.

[0111] Step S430: Obtain the three-dimensional machining surface morphology data generated after the simulation machining is completed from the reinforcement learning interactive environment, compare the three-dimensional machining surface morphology data with the surface roughness index in the machining accuracy requirement information, calculate the surface roughness deviation value, and generate the surface quality reward component based on the surface roughness deviation value.

[0112] After the simulation of the machining process is completed, the simulator generates a dataset describing the geometry of the machined surface, i.e., three-dimensional machined surface topography data. This data usually exists in the form of point clouds or height fields, recording the spatial coordinates of a large number of sampled points on the workpiece surface. From this dataset, according to international standards, such as ISO 4287, the surface roughness parameter is calculated, specifically the profile arithmetic mean deviation value. During the calculation, a sampling length is intercepted along a certain direction on the machined surface, and the height data on this profile line is extracted. The arithmetic mean of the absolute values ​​of the deviations of the heights of each point on the profile relative to the centerline is calculated. The calculated actual profile arithmetic mean deviation value is compared with the target surface roughness value specified in the machining accuracy requirement information obtained in step S100, and the absolute value of the difference between the two is taken as the surface roughness deviation value. The smaller the deviation value, the better the quality of the machined surface. Based on this deviation value, a surface quality bonus component is generated through a preset mapping function, such as a linear inverse proportional function. The smaller the deviation value, the larger the bonus component.

[0113] Step S440: Obtain processing time consumption data during the simulated processing process from the reinforcement learning interactive environment, compare the processing time consumption data with the preset standard processing time benchmark value, calculate the time saving ratio, and generate processing efficiency reward components based on the time saving ratio.

[0114] During the simulated machining process, the simulator's internal timer module records the total time elapsed from the start of tool movement to the end of machining. This time includes cutting time and possible idle travel time; this total time is the machining time consumption data. A preset standard machining time benchmark value is retrieved from the process knowledge base. This benchmark value is a reference time calculated based on typical machining parameters and paths. The actual machining time is compared with the standard machining time benchmark value. If the actual time is less than the benchmark value, the time saving percentage is calculated, which is the difference between the benchmark value and the actual time, divided by the benchmark value. If the actual time is greater than or equal to the benchmark value, the time saving percentage is set to zero. Based on this saving percentage, a machining efficiency reward component is generated through a linear function. The higher the saving percentage, the larger the reward component, thereby encouraging strategies that can shorten machining time.

[0115] Step S450: Based on the contact relationship data between the tool path and the workpiece material during the simulated machining process, calculate the cumulative cutting load of the tool per unit machining path length, and generate a tool wear bonus component based on the cumulative cutting load.

[0116] During the simulated machining process, the simulator tracks the contact state between the tool and the workpiece in real time, generating contact relationship data. This data records the area where the tool participates in cutting and the magnitude of the cutting load in that area at each moment. By integrating over the entire machining process, the cumulative cutting load of the tool along the entire path can be calculated, such as the integral of the total cutting force over time or the integral of the total cutting power over time. Simultaneously, the total machining path length is extracted from the path data. Dividing the cumulative cutting load by the total machining path length yields the average cutting load per unit path length, which reflects the stress intensity on the tool. Higher stress intensity generally leads to faster tool wear. This average cutting load is compared with a preset ideal load range, and the degree of deviation is calculated. A tool wear reward component is generated based on the degree of deviation; the smaller the deviation, the more balanced the tool stress, the less tool wear, and the larger the reward component.

[0117] Step S460: Perform a weighted summation calculation on the stability reward component, surface quality reward component, machining efficiency reward component, and tool wear reward component to generate a comprehensive machining quality reward value, and store the machining quality reward value in association with the preliminary machining action strategy set.

[0118] The normalized stability reward component generated in step S426, the surface quality reward component generated in step S430, the machining efficiency reward component generated in step S440, and the tool wear reward component generated in step S450 are aggregated together. A set of preset weight coefficients is obtained from the process knowledge base, corresponding to the relative importance of the four optimization objectives: stability, surface quality, machining efficiency, and tool wear. The sum of these four weight coefficients is one. A weighted summation calculation is performed, multiplying each reward component by its corresponding weight coefficient, and then adding the four products together. The resulting sum is a comprehensive machining quality reward value. Finally, the calculated machining quality reward value is associated with the initial machining action strategy set that generated the reward value, and stored in the experience replay buffer as a key-value pair, serving as data samples for training the reinforcement learning model.

[0119] Step S500: Evaluate the merits of the preliminary processing action strategy set based on the processing quality reward value, and generate path optimization gradient information to indicate the direction of strategy adjustment.

[0120] In one implementation, step S500 may specifically include the following steps S510-S560: Step S510: Pair the initial processing action strategy set with the processing quality reward value to construct training sample pairs containing strategy description data and corresponding reward values. The strategy description data in each training sample pair contains the coordinate information of all adjustable path nodes in the initial processing action strategy set.

[0121] The initial processing action policy set is read from the experience replay buffer, containing the final spatial coordinates of all adjustable path nodes under that policy. Simultaneously, the processing quality reward value associated with that policy is read. These two sets of data are combined to form a training sample pair. In this pair, the policy description data is a high-dimensional vector, with its elements being the coordinates of all adjustable path nodes arranged in order. The reward value is a scalar. Such a sample pair, consisting of an input vector and an output scalar, constitutes a training instance in supervised learning. Multiple such sample pairs are collected to train or update the policy evaluation network.

[0122] Step S520: Input the training samples into the policy evaluation network of the reinforcement learning interactive environment, calculate the expected reward value of the initial processing action policy set in the current state through the value function of the policy evaluation network, and compare the difference between the expected reward value and the actual processing quality reward value.

[0123] The policy evaluation network is a deep neural network with an input layer dimension identical to that of the policy description data, and an output layer consisting of a single neuron that outputs a scalar value. The policy description data from the training sample pairs constructed in step S510 is input into the policy evaluation network. The network performs forward propagation computation through its multiple hidden layers, ultimately generating a numerical value at the output layer. This value represents the expected reward predicted by the policy evaluation network based on its currently learned knowledge for that policy description data. Then, this predicted expected reward value is compared with the actual processing quality reward value from the training sample pairs, and the difference between the two is calculated. This difference is typically measured using the mean squared error function, which calculates the square of the difference between the predicted and actual values.

[0124] Step S530: Based on the difference between the expected return value and the actual processing quality reward value, calculate the contribution of the coordinates of each adjustable path node in the preliminary processing action strategy set to the reward value, and generate the local contribution coefficient corresponding to each adjustable path node.

[0125] In one implementation, step S530 may specifically include the following steps S531 to S536: Step S531: Perform small positive and negative perturbations on the coordinates of each adjustable path node in the preliminary processing action policy set to generate two perturbation policy variants, and input the two policy variants into the policy evaluation network to obtain the corresponding perturbation reward value.

[0126] For each adjustable path node in the initial processing action policy set, perform the following operations: First, copy the current policy description data to obtain a copy. In the copy, increase the x-coordinate value of the node by a small positive number, such as a preset small perturbation, while keeping the coordinates of all other nodes unchanged. This generates a policy variant with positive perturbation. Copy the original policy description data again, and in another copy, decrease the x-coordinate value of the node by the same small perturbation to generate a policy variant with negative perturbation. Input these two perturbation policy variants into the policy evaluation network used in step S520, respectively. The network performs forward propagation calculations on each variant and outputs the corresponding predicted reward value, denoted as the reward value after positive perturbation and the reward value after negative perturbation, respectively. Repeat the above process for the y-coordinate and z-coordinate of the node.

[0127] Step S532: Calculate the positive difference between the reward value after positive disturbance and the original processing quality reward value, and the negative difference between the reward value after negative disturbance and the original processing quality reward value. Use the average of the positive and negative differences as the coordinate change sensitivity index of the adjustable path node.

[0128] For the same coordinate component of the same adjustable path node, such as the x-coordinate, two new reward values ​​are obtained after step S531. The positive difference is obtained by subtracting the original processing quality reward value recorded in step S510 from the reward value after positive perturbation. This difference reflects the direction and magnitude of the reward value change when the x-coordinate of the node increases by a small amount. The negative difference is obtained by subtracting the original processing quality reward value from the reward value after negative perturbation, reflecting the change in the reward value when the x-coordinate decreases by a small amount. Then, the arithmetic mean of these two differences is calculated. This average eliminates the randomness of unidirectional perturbation and more robustly reflects the sensitivity of the reward value to small changes in the x-coordinate of the node. This average is the sensitivity index of the x-coordinate component of the node. The sensitivity indices are also calculated separately for the y-coordinate and z-coordinate.

[0129] Step S533: Conduct a correlation analysis between the coordinate change sensitivity index and the difference between the expected return value and the actual processing quality reward value, calculate the weight ratio of the coordinate change sensitivity index in the total difference, and obtain the initial contribution coefficient of each adjustable path node.

[0130] Step S520 calculates the total difference between the expected return value and the actual reward value. Step S532 calculates a sensitivity index for each coordinate component of each node. Now, the absolute values ​​of the sensitivity indices of all coordinate components of all nodes are summed to obtain a total sensitivity. For a specific coordinate component of a node, the absolute value of its sensitivity index is divided by the total sensitivity to obtain a ratio. This ratio reflects the relative importance of the change in that coordinate component of the node in explaining the total difference; this ratio is the initial contribution coefficient of that coordinate component of the node. If a node's sensitivity index is positive, it indicates that its current coordinate position makes a positive contribution to obtaining a high reward; if it is negative, it indicates a negative contribution. The initial contribution coefficients of the three coordinate components of the same node are combined to obtain the initial contribution coefficient vector of that node.

[0131] Step S534: Normalize the initial contribution coefficients so that the sum of the contribution coefficients of all adjustable path nodes is equal to the preset normalization constant value, and obtain the normalized local contribution coefficients.

[0132] The initial contribution coefficients calculated in step S533 may vary in magnitude and range depending on the processing task, making them unsuitable for comparison across different tasks and as unified gradient information. Therefore, normalization is required. First, calculate the sum of squares of the initial contribution coefficients for all coordinate components of all adjustable path nodes, then take the square root to obtain the magnitude of the vector formed by these coefficients. Then, divide each initial contribution coefficient by this magnitude. After this processing, the vector formed by all coefficients becomes a unit vector with a magnitude of one. Each component in this unit vector is the normalized local contribution coefficient. The sign of the normalized coefficients retains the directional information of the original contribution, while the absolute value indicates the relative importance of the node at the unit scale.

[0133] Step S535: Determine the direction of coordinate adjustment for the adjustable path node based on the sign of the local contribution coefficient. A positive contribution coefficient indicates that the current coordinate adjustment direction is conducive to increasing the reward value, while a negative contribution coefficient indicates that the current coordinate adjustment direction is not conducive to increasing the reward value and needs to be reversed.

[0134] For each coordinate component of each adjustable path node, analyze the sign of its normalized local contribution coefficient. If the coefficient is positive, it means that under the current strategy, a small perturbation of that coordinate component in the positive direction will increase the reward value; or, the current coordinate value of the node in that component is relatively small relative to its optimal position. Therefore, to increase the reward value, the coordinate component of the node should be adjusted in the positive direction. Conversely, if the coefficient is negative, it means that perturbing the coordinate component in the positive direction will decrease the reward value, while perturbing it in the negative direction will increase the reward value. Therefore, the coordinate component of the node should be adjusted in the negative direction. In this way, the sign of the contribution coefficient clearly indicates the direction of optimization adjustment for each node's coordinates.

[0135] Step S536: Associate and encapsulate the local contribution coefficient with the corresponding coordinate adjustment direction information to generate complete contribution description data for each adjustable path node.

[0136] The normalized local contribution coefficients generated in step S534 are integrated with the coordinate adjustment direction information determined in step S535. For each adjustable path node, a data structure is created containing three entries, corresponding to the x, y, and z coordinate components, respectively. Each entry records two pieces of information: one is the adjustment direction, represented as a unit direction vector, for example, the positive x-direction is represented as (1,0,0), and the negative x-direction is represented as (-1,0,0); the other is the adjustment step size factor, which is the absolute value of the normalized local contribution coefficient. The magnitude of this absolute value indicates the urgency or intensity of adjustment in that direction. All these data structures for all nodes are combined to form a list, which is the complete contribution description data corresponding to each adjustable path node.

[0137] Step S540: Perform directional decomposition on the local contribution coefficients, decomposing the contribution coefficient of each adjustable path node into a normal component along the normal direction of the machining surface and a tangential component along the tangential direction of the machining surface, to obtain the two-dimensional gradient direction vector of each adjustable path node.

[0138] For each adjustable path node, obtain the normal unit vector and tangential unit vector at the location of the node on the machining surface. The normal unit vector is obtained by calculating and normalizing the cross product of the partial derivative vectors of the surface at that point, while the tangential unit vector is determined based on the projection of the machining path direction onto the tangential plane of the surface. Project the three-dimensional adjustment direction vector from the complete contribution description data of the node generated in step S536 onto the normal unit vector and tangential unit vector, respectively. The projection is calculated by multiplying the adjustment direction vector by the normal unit vector; the result is the contribution of the normal component, with the sign indicating positive or negative adjustment along the normal. Similarly, multiplying the adjustment direction vector by the tangential unit vector yields the contribution of the tangential component. Thus, the contribution information, originally in the three-dimensional Cartesian coordinate system, is decomposed into physically meaningful normal and tangential directions, forming a two-dimensional gradient direction vector. The two components of this two-dimensional vector represent the magnitude of adjustment required in the normal and tangential directions, respectively.

[0139] Step S550: Integrate the two-dimensional gradient direction vectors of all adjustable path nodes to generate a multi-dimensional gradient vector with the same dimension as the number of adjustable path nodes.

[0140] Following the fixed order of all adjustable path nodes in the path, the two-dimensional gradient direction vectors calculated for each node in step S540 are arranged sequentially. For each node, its two-dimensional vector has two components: a normal component and a tangential component. Therefore, if there are N adjustable nodes, the integrated multidimensional gradient vector is a vector of length 2N. The first two elements of this vector correspond to the normal and tangential gradient components of the first node, the next two elements correspond to the normal and tangential gradient components of the second node, and so on. This multidimensional gradient vector fully describes the degree of movement that each adjustable path node should make in its respective local surface coordinate system to improve processing quality, starting from the current strategy.

[0141] Step S560: Normalize the multidimensional gradient vector to obtain path optimization gradient information used to indicate the direction of policy adjustment.

[0142] The multidimensional gradient vector generated in step S550 may have components with significantly different values, which could lead to instability in the optimization process if used directly. Therefore, it needs to be normalized. The magnitude of the entire multidimensional gradient vector is calculated by summing the squares of each component and then taking the square root. Then, each component in the vector is divided by this magnitude. After normalization, the new multidimensional gradient vector becomes a unit vector with a magnitude of one. This normalized vector retains the relative proportions between the components in the original gradient information but unifies the scale of all components to a standard range. This normalized multidimensional gradient vector is the final path optimization gradient information.

[0143] Step S600: Based on the path optimization gradient information, perform iterative update processing on the preliminary processing action strategy set to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

[0144] In one implementation, step S600 may specifically include the following steps S610 to S660: Step S610: Input the path optimization gradient information into the policy update module of the reinforcement learning interactive environment, parse the multi-dimensional gradient vector in the path optimization gradient information, and obtain the gradient direction indication and gradient magnitude value corresponding to each adjustable path node.

[0145] The policy update module receives the path optimization gradient information generated in step S560, which is a normalized multi-dimensional vector. The module first splits this long vector into segments corresponding to each node according to a preset node order. For each adjustable path node, two values ​​are extracted from its two corresponding consecutive vector elements: a normal gradient component and a tangential gradient component. These two components together constitute the gradient direction indicator for that node, specifically represented as a two-dimensional vector whose direction indicates the direction of movement within the normal-tangential plane. Simultaneously, the magnitude of these two components, i.e., their absolute values, represents the gradient magnitude in that direction. A larger magnitude indicates a higher urgency for the node to adjust in that direction, and a greater potential contribution to improving the overall reward value.

[0146] Step S620: Determine the coordinate adjustment direction of each adjustable path node according to the gradient direction indication. If the gradient direction indication is positive, keep the current adjustment direction unchanged. If the gradient direction indication is negative, reverse the current adjustment direction and generate the updated adjustment direction for each adjustable path node.

[0147] For each adjustable path node, the signs of the normal and tangential gradient components resolved in step S610 are analyzed. In the normal direction, if the normal gradient component is positive, it means that to increase the reward value, the node should be moved along the positive direction of the surface normal; therefore, the updated normal adjustment direction is set to the positive direction of the surface normal. If the normal gradient component is negative, the updated normal adjustment direction is set to the negative direction of the surface normal. Similarly, in the tangential direction, if the tangential gradient component is positive, the updated tangential adjustment direction is set to the positive direction of the machining toolpath tangent; if it is negative, it is set to the negative direction of the tangent. In this way, the abstract gradient sign is transformed into a concrete, geometrically meaningful spatial movement direction. These two directions are perpendicular to each other and together constitute the updated adjustment direction of the node on the local surface.

[0148] Step S630: Calculate the coordinate adjustment step size of each adjustable path node based on the gradient magnitude value, and multiply the gradient magnitude value by the preset learning rate coefficient to obtain the actual displacement distance of each adjustable path node.

[0149] The gradient magnitude values ​​resolved in step S610, i.e., the absolute values ​​of the normal and tangential components, represent the urgency of the adjustment, but not the final physical distance to be moved. To obtain the actual step size, a preset learning rate coefficient needs to be introduced. This learning rate coefficient is a global hyperparameter used to control the step size of each iteration, preventing excessively large update steps from causing policy oscillations or non-convergence. For each adjustable path node, the absolute value of its normal gradient component is multiplied by the learning rate coefficient to obtain the actual displacement distance of the node in the normal direction. Similarly, the absolute value of its tangential gradient component is multiplied by the learning rate coefficient to obtain the actual displacement distance in the tangential direction. In this way, each node obtains specific movement distance values ​​in two mutually perpendicular directions.

[0150] Step S640: Update the current coordinates of each adjustable path node in the preliminary processing action strategy set according to the updated adjustment direction and actual displacement distance, and generate an iterative processing action strategy set containing the updated path node coordinates.

[0151] For each adjustable path node, two pieces of information are defined: the direction of movement, i.e., the positive and negative normal and tangential directions determined in step S620; and the distance of movement, i.e., the normal displacement distance and tangential displacement distance determined in step S630. The movements in these two directions are then combined. First, the unit normal vector and unit tangential vector of the node are obtained from its current position on the surface. Then, the displacement vectors are calculated: the unit normal vector is multiplied by the normal displacement distance to obtain a normal displacement vector; the unit tangential vector is multiplied by the tangential displacement distance to obtain a tangential displacement vector. These two displacement vectors are then vector-summed to obtain a total displacement vector. Finally, the node's current spatial coordinates are added to this total displacement vector to obtain the updated spatial coordinates of the node. This process is performed on all adjustable nodes, while the coordinates of fixed nodes remain unchanged. After updating all nodes, the resulting new path data constitutes the set of iterative processing action strategies.

[0152] Step S650: Calculate the constraint satisfaction of the updated path node coordinates in the iterative processing action strategy set according to the processing sequence constraints, generate the path node coordinate correction amount that satisfies the constraints, and adjust the updated path node coordinates according to the correction amount.

[0153] The node coordinates in the iterative machining action strategy set generated in step S640 are generated under gradient guidance, but may unintentionally violate certain machining sequence constraints. For example, moving a node may cause a path segment that should belong to the roughing stage to intrude into the finishing area, or cause the toolpath to self-intersect. Therefore, constraint verification is required. Traverse all updated path nodes, and for each step sequence constraint, check whether the current node coordinates meet the conditions. For each violated constraint, calculate the amount that needs to be corrected for the node coordinates to bring it back within the allowed range. This correction might involve pulling an out-of-bounds node back to the boundary, or adjusting the positions of its preceding and following nodes to straighten the path sequence. After applying corrections to all nodes that violate constraints, a set of path node coordinate corrections is obtained. These corrections are then used to fine-tune the coordinates generated in step S640, ensuring that the final path follows the gradient optimization direction and fully complies with all process constraints.

[0154] Step S660: Input the adjusted iterative processing action strategy set back into the reinforcement learning interactive environment for iterative processing of processing action exploration, processing state feedback collection, processing quality reward value calculation and path optimization gradient information generation, until the processing quality reward value obtained from multiple iterations tends to be stable and meets the processing accuracy requirements. Use the iterative processing action strategy set obtained from the last iteration as the processing path optimization strategy set.

[0155] In one implementation, step S660 may specifically include the following steps S661 to S666: Step S661: At the beginning of each iteration loop, compare the current iteration number with the preset maximum iteration number threshold. If the current iteration number has reached the maximum iteration number threshold, terminate the loop and output the set of iterative processing action strategies obtained in the current iteration.

[0156] Before each iteration, a counter is set to record the number of iterations already completed. The value of this counter is compared to a preset maximum iteration threshold. This threshold is an integer used to prevent the optimization process from going indefinitely; for example, it could be set to 10,000 iterations. If the current iteration count equals or exceeds this threshold, it means the optimization has reached the preset computational resource limit, and it must stop even if it may not have fully converged. At this point, the loop is forcibly terminated, and the set of iterative processing strategies generated in the current iteration is output as the final result. This mechanism ensures that the algorithm has a definite stopping point under all circumstances.

[0157] Step S662: If the current iteration count has not reached the maximum iteration count threshold, the adjusted iterative processing action strategy set is input into the reinforcement learning interactive environment, the processing action exploration process is executed to generate a new preliminary processing action strategy set, and the corresponding real-time processing status feedback data is collected.

[0158] If the judgment result of step S661 is that the current iteration number is less than the maximum iteration number threshold, the loop continues. The iterative processing action policy set updated and constrained in the previous round is used as input and passed to the reinforcement learning interactive environment. The interactive environment starts a new round of simulation and, following the process from steps S300 to S350, performs processing action exploration. That is, the agent selects actions according to the current policy, the environment executes the actions and updates the path nodes, generating a new preliminary processing action policy set. At the same time, following the process in step S400, the environment collects real-time processing status feedback data during the execution of the new policy, such as force data, time data, and surface morphology data. This data is temporarily stored for the next reward calculation.

[0159] Step S663: Calculate the new processing quality reward value based on the newly collected real-time processing status feedback data, and compare the new processing quality reward value with the processing quality reward value of the previous iteration to calculate the change in reward value between adjacent iterations.

[0160] Following the methods described in steps S410 to S460 of step S400, the newly collected real-time processing status feedback data from step S662 is processed and analyzed to calculate a new processing quality reward value. The processing quality reward value stored at the end of the previous iteration is read from memory. The newly calculated reward value is subtracted from the previous reward value to obtain the difference. The absolute value of this difference is the change in reward value between two adjacent iterations. This change is a key indicator for measuring whether the optimization process has converged.

[0161] Step S664: Compare the change in reward value with a preset convergence threshold. If the change in reward value is less than the preset convergence threshold and the condition is met for a preset number of consecutive iterations, then the processing quality reward value is determined to be stable.

[0162] The change in reward value calculated in step S663 is compared with a preset convergence threshold, which is a very small positive number. If the current change is less than this threshold, it means that the increase in reward value from the last iteration to the current iteration is very small. The system also needs to determine whether this phenomenon is persistent. Therefore, a counter is set up to record the number of iterations that continuously satisfy the condition that the change is less than the threshold. If the condition is met this time, and the previous consecutive iterations, such as five consecutive iterations, also satisfy this condition, then it can be determined that the processing quality reward value has stabilized and no longer fluctuates significantly. If the current change is not less than the threshold, then this consecutive counter is reset.

[0163] Step S665: After determining that the processing quality reward value tends to stabilize, compare the current set of processing action strategies with each indicator in the processing accuracy requirement information, and calculate the satisfaction parameter of each indicator.

[0164] Once step S664 determines that the reward value tends to stabilize, the system needs to verify whether the processing strategy in this stable state truly meets the process requirements. From the initial processing task description data set obtained in step S100, specific indicators from the processing accuracy requirement information are extracted, such as the arithmetic mean deviation of surface roughness profile and dimensional tolerance values. For the processing action strategy set obtained in the current iteration, simulation analysis is used to calculate the actual achievable performance indicators. Each actually calculated indicator is compared with its corresponding requirement indicator. For each indicator, if the actual value is better than or equal to the requirement value, the satisfaction parameter for that indicator is recorded as 100%; if the actual value is worse than the requirement value, the degree of deviation of the actual value from the requirement value is calculated, and this deviation degree is converted into a satisfaction parameter between 0 and 100%.

[0165] Step S666: If the satisfaction parameters of each indicator reach the preset satisfaction threshold, the loop is terminated and the current iteration processing action strategy set is output as the processing path optimization strategy set; if the processing accuracy requirement information is not met or the change in reward value is not less than the convergence threshold, the next iteration loop is executed.

[0166] Check the satisfaction parameters of all indicators calculated in step S665. Determine if the satisfaction parameter of each indicator has reached a preset satisfaction threshold, such as 95%. If all indicators are satisfied, it means that the current strategy has converged not only in terms of reward value but also in terms of specific physical indicators. At this point, terminate the entire iteration loop and output the current set of iterative processing action strategies as the final result, i.e., the set of processing path optimization strategies. Conversely, if even if the reward value has stabilized, but some indicators have not reached the satisfaction threshold, it means that the optimization has fallen into a local optimum but has not met the process requirements. In this case, the loop should not be terminated, but other measures need to be taken, such as adjusting the exploration strategy, and the next iteration should continue. If the change in reward value has not yet fallen below the convergence threshold, the next iteration should also continue.

[0167] The following description continues to illustrate the exemplary structure of the mold processing path optimization device 255 in the application scenario provided in this application embodiment as a software module. In some embodiments, such as... Figure 3 As shown, the software module in the skill release device 255 stored in the virtual scene of the memory 450 may include: The data acquisition module 2551 is used to acquire the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirement information corresponding to each processing step. Data deconstruction module 2552 is used to perform processing path initialization deconstruction processing on the initial processing task description data set to generate an initial processing path unit set with adjustable margin corresponding to each processing step; The reinforcement learning module 2553 is used to input the initial processing path unit set into the reinforcement learning interactive environment for processing action exploration and to generate a preliminary processing action strategy set for the initial processing path unit set. The feedback calculation module 2554 is used to execute the set of preliminary processing action strategies in the reinforcement learning interactive environment and collect the corresponding real-time processing status feedback data, and calculate the processing quality reward value corresponding to the set of preliminary processing action strategies based on the real-time processing status feedback data. The strategy evaluation module 2555 is used to evaluate the merits of the preliminary processing action strategy set based on the processing quality reward value, and generate path optimization gradient information to indicate the direction of strategy adjustment. The iterative update module 2556 is used to perform iterative update processing on the preliminary processing action strategy set based on the path optimization gradient information to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

[0168] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the reinforcement learning-based mold processing path optimization method provided in this application. For example, ... Figure 2 The method for optimizing mold processing paths based on reinforcement learning is shown.

[0169] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or CD-ROM, etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0170] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for optimizing mold processing paths based on reinforcement learning, characterized in that, The method includes: Obtain the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirement information corresponding to each processing step. The initial processing task description data set is subjected to processing path initialization destructuring to generate an initial processing path unit set with adjustable margin corresponding to each processing step. The initial set of processing path units is input into a reinforcement learning interactive environment for processing action exploration, thereby generating a preliminary set of processing action strategies for the initial set of processing path units. In the reinforcement learning interactive environment, the set of preliminary processing action strategies is executed and corresponding real-time processing status feedback data is collected. Based on the real-time processing status feedback data, the processing quality reward value corresponding to the set of preliminary processing action strategies is calculated. The initial processing action strategy set is evaluated based on the processing quality reward value to generate path optimization gradient information that indicates the direction of strategy adjustment. Based on the path optimization gradient information, the initial processing action strategy set is iteratively updated to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

2. The mold processing path optimization method based on reinforcement learning according to claim 1, characterized in that, The step of performing processing path initialization and deconstruction processing on the initial processing task description data set to generate an initial processing path unit set with adjustable margins corresponding to each processing step includes: The three-dimensional geometric structure description information in the initial processing task description data set is parsed, and the surface contour surface topology features of the mold to be processed are extracted. The surface contour surface topology features include the curvature change gradient distribution of the surface and the adjacency relationship graph structure between the surfaces. Based on the processing sequence constraint information, the surface contour topology features are divided into processing regions to obtain a sequence of processing region sub-blocks arranged in the order of process execution, with each processing region sub-block corresponding to one processing process; For each processing area sub-block, a tool contact path dot matrix is ​​generated based on the processing accuracy requirement information corresponding to the processing procedure. The spacing between adjacent path points in the tool contact path dot matrix is ​​adjusted according to the surface roughness index in the processing accuracy requirement information. The path connection order of the tool contact path point matrix is ​​optimized to generate an initial path segment set that connects all tool contact path points, and at least one path adjustment node is reserved in each initial path segment as an adjustable margin for subsequent processing actions. The initial path segment set is associated with the spatial coordinate data of the path adjustment nodes and stored together to generate a directed graph structure containing the position coordinates of the path adjustment nodes and the connection relationships between the path adjustment nodes. Based on the directed graph structure, extract the information of the preceding and following adjacent nodes of each path adjustment node in the processing path, determine the adjustable degree of freedom range of each path adjustment node according to the preceding and following adjacent node information, and obtain the initial processing path unit set with adjustable margin corresponding to each processing operation.

3. The mold processing path optimization method based on reinforcement learning according to claim 2, characterized in that, The step of generating a tool contact path dot matrix for each processing area sub-block based on the processing accuracy requirement information corresponding to the processing procedure includes: Extract the surface parameter equations of the sub-blocks of the processing area, and perform isoparametric line sampling on the surface parameter equations to obtain the initial set of sampling point coordinates distributed along the U and V directions of the surface; Analyze the machining accuracy requirement information corresponding to the machining process, obtain the maximum allowable residual height value corresponding to the surface roughness index, and calculate the allowable range of chord height error between adjacent tool contacts based on the maximum allowable residual height value; Based on the allowable range of chord height error, the initial set of sampling point coordinates is subjected to adaptive point cloud density adjustment processing. The sampling point density is increased in areas with large surface curvature gradient and decreased in areas with small surface curvature gradient, resulting in a density-optimized set of sampling point coordinates. Based on the boundary curve equation of the processing area sub-block, the spatial position mapping process is performed on the density-optimized set of sampling point coordinates to generate the distance parameter and orientation parameter of each sampling point relative to the boundary curve. The coordinates of the sampling points after spatial location mapping are arranged according to the parameter order of the U and V directions to generate a tool contact path point matrix with topological connection relationship; The tool contact point coordinates in the tool contact path point matrix with topological connection are sorted according to the machining tool path direction to obtain a tool contact path point matrix arranged sequentially along the machining tool path direction.

4. The mold processing path optimization method based on reinforcement learning according to claim 2, characterized in that, The optimization of the path point connection order of the tool contact path point matrix to generate an initial set of path segments connecting all tool contact path points includes: Extract the spatial three-dimensional coordinates of all tool contacts in the tool contact path point matrix, calculate the Euclidean distance between every two adjacent tool contacts as the candidate length of the path segment, and sort the candidate lengths of the path segment according to the machining tool path direction to obtain the candidate sequence of path segments along the machining tool path direction. For each path segment candidate in the path segment candidate sequence, determine whether the path segment crosses the restricted area within the processing area sub-block. If it crosses the restricted area, remove the path segment from the path segment candidate sequence and reconnect the tool contacts at both ends of the path segment to other tool contacts that bypass the restricted area. The adjacent path segments in the candidate sequence of path segments after the restricted area removal process are connected end to end to generate an initial continuous path trajectory covering all knife contacts in the knife contact path point matrix. In the initial continuous path trajectory, a path adjustment node is marked at every preset number of tool contact points. The spatial coordinates of the tool contact points marked as path adjustment nodes are set to an adjustable state, allowing subsequent machining operations to offset these coordinates. Based on the initial continuous path trajectory and the marked positions of the path adjustment nodes, an initial path segment set consisting of fixed path segments and adjustable path nodes is generated. Based on the spatial relationship between fixed path segments and adjustable path nodes in the initial path segment set, a processing path node topology graph with adjustable degrees of freedom is constructed.

5. The mold processing path optimization method based on reinforcement learning according to claim 1, characterized in that, The step of inputting the initial processing path unit set into the reinforcement learning interactive environment for processing action exploration processing, and generating a preliminary processing action strategy set for the initial processing path unit set, includes: In the reinforcement learning interactive environment, a state space representation is constructed, and the current spatial coordinates of each adjustable path node in the initial processing path unit set are used as the basis vector of the state space. The processing sequence constraint information is encoded as the legality constraint conditions of state transition. The coordinate offset direction and its offset magnitude of each adjustable path node are used as the basic motion primitives of the motion space. The offset direction includes the tool lifting direction along the normal of the machining surface and the feed direction along the tangent of the machining surface. The exploration policy network of the reinforcement learning interactive environment is invoked to evaluate the action value of the state space representation. Based on the evaluation result, a set of candidate action primitives in the current state are selected from the action space. Each candidate action primitive corresponds to a coordinate offset operation of an adjustable path node. The selected candidate action primitives are combined and encapsulated to generate an action instruction sequence containing multiple adjustable path node coordinate offset instructions. Each coordinate offset instruction in the action instruction sequence corresponds uniquely to an adjustable path node in the initial processing path unit set. By executing the action instruction sequence through a reinforcement learning interactive environment, coordinate corrections are generated for each adjustable path node in the initial processing path unit set; The state space representation is updated based on the coordinate correction, resulting in a set of preliminary processing action strategies containing the corrected path nodes.

6. The mold processing path optimization method based on reinforcement learning according to claim 5, characterized in that, The exploration policy network that invokes the reinforcement learning interactive environment evaluates the action value of the state space representation, and selects a set of candidate action primitives from the action space in the current state based on the evaluation results, including: The current spatial coordinates of each adjustable path node in the state space representation are normalized to generate a normalized state feature vector. The dimension of the state feature vector is consistent with the number of adjustable path nodes. The normalized state feature vector is input into the feature extraction layer of the exploration strategy network. Through multi-layer convolution operations, the spatial correlation features between nodes in the state feature vector are extracted to generate a hidden state feature representation with global context information. The hidden state feature representation is input into the action value output layer of the exploration policy network to calculate the expected cumulative reward value of each action primitive in the action space in the current state, and to generate the action value estimate corresponding to each action primitive. The estimated action values ​​are sorted, and the top preset number of action primitives with the highest estimated action values ​​are selected as preliminary candidate action primitives. The adjustable path node identifier corresponding to each preliminary candidate action primitive is recorded. Conflict detection is performed on the preliminary candidate action primitives. If two or more preliminary candidate action primitives act on the same adjustable path node, the one with the highest action value estimate is retained, and the other conflicting action primitives are removed. Based on the processing sequence constraints, the preliminary candidate action primitives after conflict detection are constrained and encoded to generate a set of constraint action primitives that match the processing sequence constraints, which serves as the final set of candidate action primitives. The step of executing the action instruction sequence through a reinforcement learning interactive environment to generate coordinate corrections for each adjustable path node in the initial processing path unit set includes: Parse each coordinate offset instruction in the action instruction sequence, extract the adjustable path node identifier corresponding to each coordinate offset instruction and its target offset direction and target offset magnitude, wherein the target offset magnitude is represented as the displacement distance in the target offset direction; The current spatial coordinates corresponding to the adjustable path node identifier are vector-superimposed with the target offset direction and target offset magnitude to generate the updated spatial coordinates of each adjustable path node. Based on the updated spatial coordinates, the coordinates of the adjustable path nodes in the initial processing path unit set are replaced, while keeping the coordinates of the fixed path nodes unchanged, to obtain an intermediate processing path unit set with updated path nodes. The smoothness of the connection relationship between path nodes in the intermediate processing path unit set is evaluated, the curvature change gradient between adjacent updated path nodes is detected, and a smoothness adjustment coefficient is generated based on the curvature change gradient. The coordinates of the updated path nodes are fine-tuned according to the smoothness adjustment coefficient to keep the path curvature change between adjacent path nodes within a preset smoothness range. The final coordinates of all path nodes in the intermediate processing path unit set after smoothing fine-tuning are integrated to generate complete path description data containing updated path node coordinates and fixed path node coordinates as coordinate correction amounts.

7. The mold processing path optimization method based on reinforcement learning according to claim 1, characterized in that, The process of executing the initial processing action strategy set in the reinforcement learning interactive environment and collecting corresponding real-time processing status feedback data, and calculating the processing quality reward value corresponding to the initial processing action strategy set based on the real-time processing status feedback data, includes: In a reinforcement learning interactive environment, the cutting tool is simulated to perform cutting motion according to the modified path nodes in the initial machining action strategy set, generating a real-time tool force data sequence during the simulated machining process; The real-time tool force data sequence is analyzed, and the peak force, force fluctuation amplitude, and average force are extracted as machining stability evaluation indicators. The machining stability evaluation indicators are then compared with a preset stability threshold range to generate a stability bonus component. The three-dimensional machining surface morphology data generated after the simulation machining is completed is obtained from the reinforcement learning interactive environment. The three-dimensional machining surface morphology data is compared with the surface roughness index in the machining accuracy requirement information to calculate the surface roughness deviation value, and a surface quality reward component is generated based on the surface roughness deviation value. The processing time consumption data during the simulated processing process is obtained from the reinforcement learning interactive environment. The processing time consumption data is compared with the preset standard processing time benchmark value, the time saving ratio is calculated, and a processing efficiency reward component is generated based on the time saving ratio. Based on the contact relationship data between the tool path and the workpiece material during the simulated machining process, the cumulative cutting load of the tool per unit machining path length is calculated, and a tool wear bonus component is generated based on the cumulative cutting load. The stability reward component, surface quality reward component, machining efficiency reward component, and tool wear reward component are weighted and summed to generate a comprehensive machining quality reward value, and the machining quality reward value is associated with and stored in the preliminary machining action strategy set.

8. The mold processing path optimization method based on reinforcement learning according to claim 7, characterized in that, The process involves analyzing the real-time tool force data sequence, extracting the peak force, force fluctuation amplitude, and average force as machining stability evaluation indicators, and comparing these indicators with a preset stability threshold range to generate a stability bonus component, including: Peak detection processing is performed on the real-time tool force data sequence to identify all local maxima in the sequence and extract the maximum value as the tool force peak value. The tool force peak value is compared with a preset force peak value safety threshold. If the safety threshold is exceeded, a negative reward value is generated. The standard deviation of all data points in the real-time tool force data sequence is calculated as the force fluctuation amplitude. The force fluctuation amplitude is compared with a preset fluctuation amplitude tolerance threshold. The greater the fluctuation amplitude exceeds the tolerance threshold, the greater the negative reward value generated. The arithmetic mean of all data points in the real-time tool force data sequence is calculated as the force average. The force average is compared with the preset ideal force average range. The greater the distance between the force average and the ideal force average range, the greater the negative reward value generated. Based on the comparison results of the peak force, force fluctuation amplitude, and average force of the tool with their respective thresholds or intervals, three independent penalty factors are generated, and the three penalty factors are combined according to a preset weight ratio to generate a comprehensive stability penalty value. The stability penalty value is converted into a stability reward component, and the stability reward component is negatively correlated with the stability penalty value, that is, the larger the stability penalty value, the smaller the stability reward component. The stability reward component is mapped to the range of standard reward values ​​to obtain the normalized stability reward component.

9. A mold processing path optimization device, characterized in that, include: The data acquisition module is used to acquire the initial processing task description data set corresponding to the mold processing task. The initial processing task description data set includes the three-dimensional geometric structure description information of the mold to be processed, the processing sequence constraint information, and the processing accuracy requirement information corresponding to each processing step. The data deconstruction module is used to perform processing path initialization deconstruction processing on the initial processing task description data set to generate an initial processing path unit set with adjustable margins corresponding to each processing step. The reinforcement learning module is used to input the initial processing path unit set into the reinforcement learning interactive environment for processing action exploration and to generate a preliminary processing action strategy set for the initial processing path unit set. The feedback calculation module is used to execute the set of preliminary processing action strategies in the reinforcement learning interactive environment and collect the corresponding real-time processing status feedback data, and calculate the processing quality reward value corresponding to the set of preliminary processing action strategies based on the real-time processing status feedback data. The strategy evaluation module is used to evaluate the merits of the initial processing action strategy set based on the processing quality reward value, and generate path optimization gradient information to indicate the direction of strategy adjustment. The iterative update module is used to perform iterative update processing on the preliminary processing action strategy set based on the path optimization gradient information to obtain a processing path optimization strategy set that meets the processing accuracy requirements.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.