Reinforcement learning-based automatic fiber placement trajectory planning method for composite pressure vessel

By triangulating and refining the local mesh of composite pressure vessels, and combining reinforcement learning and soft actor-commentator algorithms, the trajectory planning of composite pressure vessels was optimized, solving the problems of wrinkle defects and fiber angle deviation in traditional methods, and achieving efficient trajectory planning and structural performance improvement.

CN122113610APending Publication Date: 2026-05-29XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
Filing Date
2026-02-09
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

When manufacturing composite pressure vessels, traditional methods struggle to simultaneously optimize wrinkle defects and fiber angle deviations, leading to a trade-off between structural performance and manufacturability.

Method used

A reinforcement learning-based approach is used to triangulate and refine the local mesh of composite pressure vessels to construct a discrete map, integrate minimum turning radius constraints, and train an agent using a soft actor-critic algorithm to generate the optimal layup trajectory and optimize wrinkle defects and fiber angle deviations.

Benefits of technology

It significantly reduced wrinkle defects by 35.2%, and controlled fiber angle deviation between -2° and +1.5°, improving the geometric modeling efficiency and structural performance of complex curved surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a composite pressure vessel automatic fiber placement trajectory planning method based on reinforcement learning, comprising the following steps: triangulating a geometric model of a composite pressure vessel and identifying high-curvature areas, locally refining the high-curvature areas, generating a high-precision triangular discrete model, calculating the geometric properties of each grid element and constructing a discrete map; constructing a reinforcement learning interactive environment based on the discrete map, integrating the minimum turning radius constraint of the automatic fiber placement process, updating the state on the triangular grid after the agent performs an action and calculating the path turning angle; designing an observation space, an action space and a multi-objective reward function, training the agent in the reinforcement learning interactive environment through a soft actor-critic algorithm, iteratively optimizing under the guidance of the multi-objective reward function, and learning an optimal trajectory planning strategy that minimizes wrinkle defects and controls fiber angle deviation. The application can improve the quality of the automatic fiber placement trajectory, and takes into account manufacturability and structural performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control technology for composite materials, and relates to, but is not limited to, an automatic fiber placement trajectory planning method for composite pressure vessels based on reinforcement learning. Background Technology

[0002] Composite pressure vessels (CPVs), as lightweight solutions for storing pressurized gases and liquids, have been widely used in critical fields such as aerospace, defense, automotive, and marine engineering. Automated fiber placement technology overcomes the inherent limitations of traditional winding processes by precisely controlling fiber orientation and placement path, opening up new avenues for the efficient production of composite components with complex shapes.

[0003] Trajectory planning is crucial when manufacturing pressure vessels with complex hyperbolic heads. In traditional methods, the geodesic method can ensure the shortest fiber path, but it is difficult to control the fiber angle, resulting in a large deviation from the principal stress direction and impairing structural performance. While the fixed angle method can ensure the consistency of fiber angle, it is prone to inducing wrinkles in the head area of ​​composite pressure vessels, seriously affecting manufacturing quality and structural integrity.

[0004] Therefore, there is an urgent need for a more comprehensive and intelligent trajectory planning method that can minimize fold defects and align fiber angles by co-optimizing folds and fiber angles, thereby generating high-quality trajectories and achieving a balance between manufacturability and structural performance. Summary of the Invention

[0005] This application provides an automatic fiber placement trajectory planning method for composite material pressure vessels based on reinforcement learning.

[0006] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide an automatic fiber placement trajectory planning method for composite pressure vessels based on reinforcement learning. The method includes: triangulating the geometric model of the composite pressure vessel and identifying high curvature regions; refining the local mesh for the high curvature regions to generate a high-precision triangular discretized model; calculating the geometric attributes of each mesh unit and constructing a discrete map; constructing a reinforcement learning interactive environment based on the discrete map, wherein the reinforcement learning interactive environment integrates the minimum turning radius constraint of the automatic fiber placement process; after the agent performs an action, it completes a state update on the triangular mesh, and calculates the path turning angle after each state update; designing an observation space, an action space, and a multi-objective reward function; in the reinforcement learning interactive environment, training the agent through a soft actor-critic algorithm; the agent outputs a placement angle action based on the observation state, and performs iterative optimization under the guidance of the multi-objective reward function to learn the optimal trajectory planning strategy that minimizes wrinkle defects and controls fiber angle deviation; for the composite pressure vessel to be placed, generating the optimal placement trajectory through the trained agent and evaluating its wrinkle defect rate and fiber angle deviation.

[0007] The technical solution provided in this application triangulates the geometric model of a composite pressure vessel and identifies high-curvature regions. Local mesh refinement is then performed on these high-curvature regions to generate a high-precision triangular discretized model. This allows for the approximation of arbitrary hypercurvature geometries, including variable-curvature heads and transition zones, using triangular mesh units without requiring reparameterization or surface splicing. Furthermore, local refinement enables near-linear growth of degrees of freedom, significantly improving the efficiency and accuracy of geometric modeling for complex surfaces. The geometric properties of each mesh unit are calculated, and a discrete map is constructed. Based on this discrete map, a reinforcement learning interactive environment is built. This environment integrates the minimum turning radius constraint of the automated wire-laying process. After the agent performs an action, it updates its state on the triangular mesh and calculates the path turning angle after each state update. The minimum turning radius, a practical process constraint, is integrated into the interactive environment. Simultaneously, the feasible range of states and actions is represented through meshing, resulting in a well-defined and structurally clear state and action combination space, reducing blind spots during agent training. This paper explores various dimensions of the reinforcement learning process and enhances the reliability of its predictions for the value of each decision, thereby improving the training efficiency and convergence of the reinforcement learning process. It designs an observation space, action space, and a multi-objective reward function. In a reinforcement learning interactive environment, the agent is trained using a soft actor-critic algorithm. Based on the observation state, the agent outputs a placement angle action and iteratively optimizes under the guidance of the multi-objective reward function, learning the optimal trajectory planning strategy to minimize wrinkle defects and control fiber angle deviation. This guides the agent to learn strategies that satisfy manufacturing constraints and optimize structural performance, solving the trade-off between manufacturability and structural performance in traditional methods. For composite pressure vessels to be laid, the trained agent generates the optimal placement trajectory and evaluates its wrinkle defect rate and fiber angle deviation. Ultimately, it achieves a significant performance improvement, reducing wrinkle defects by an average of 35.2% and controlling fiber angle deviation within the range of -2° to +1.5°, demonstrating the comprehensive superiority of the reinforcement learning-based automatic fiber placement trajectory planning method for composite pressure vessels provided in this application.

[0008] Optionally, the step of triangulating the geometric model of the composite pressure vessel and identifying high-curvature regions, refining the local mesh for the high-curvature regions to generate a high-precision triangular discretized model, calculating the geometric properties of each mesh unit, and constructing a discrete map includes: triangulating the geometric model of the composite pressure vessel to generate a triangular surface mesh; calculating the curvature of each mesh unit in the triangular surface mesh and identifying high-curvature regions with curvature values ​​higher than a preset curvature threshold; refining the local mesh for the high-curvature regions by reducing the size of the triangular mesh units in the high-curvature regions and performing conformal adjustments to generate a high-precision triangular discretized model; pre-calculating and storing the geometric properties of each triangular mesh unit in the high-precision triangular discretized model, the geometric properties including node coordinates, unit surface normal vector, curvature metric value, and related geodesic direction; and constructing a structured discrete map by connecting all triangular mesh units in the high-precision triangular discretized model according to their topological connections.

[0009] Optionally, the reinforcement learning interactive environment is constructed based on the discrete map. This environment integrates the minimum turning radius constraint of the automated filament placement process. After the agent performs an action, it updates its state on the triangular mesh and calculates the path turning angle after each state update. This includes: using the discrete map as the geometric basis and pre-setting the starting and target mesh cells for trajectory planning to construct the reinforcement learning interactive environment; in the reinforcement learning interactive environment, the agent determines the adjacent mesh cell to enter at the next moment by calculating the intersection point with the boundary of the current mesh cell based on the local placement angle determined by the current action, thereby completing the state update; obtaining the geometric attributes of the updated state based on the discrete map; integrating the minimum turning radius constraint of the automated filament placement process into the reinforcement learning interactive environment; calculating the path turning angle between the current action direction and the previous action direction after the state update is completed; the environment receives the agent's action instructions; after the state update and constraint check are completed, it returns the updated state and an immediate reward signal to the agent.

[0010] Optionally, the observation space is a 15-dimensional continuous vector, constructed by concatenating five three-dimensional vectors. These five three-dimensional vectors include a position vector representing the current position of the agent, a surface normal vector representing the surface normal of the current grid cell, a principal direction vector representing the local target bearing direction, a geodesic direction vector representing the local shortest path, and a previous direction vector representing the previous placement direction. The action output of the action space is the deflection angle of fiber placement at the current placement point. This deflection angle is constrained by the minimum turning radius constraint of the automated fiber placement process. The minimum turning radius constraint is expressed by the following formula: ; In the formula, Indicates the deflection angle between adjacent path segments; Indicates the maximum deflection angle; Indicates the local path length; The minimum turning radius is represented; the action space is modeled using a Gaussian policy function, and action selection is sampled from the action space through a Gaussian distribution. During training, the Gaussian distribution gradually shrinks according to a predefined decay strategy.

[0011] Optionally, in the reinforcement learning interactive environment, after the current observation state is transferred to the next observation state through the current placement angle action, an immediate reward is calculated using a multi-objective reward function. This multi-objective reward function is determined by weighted summation of path feasibility reward, payload alignment reward, curvature penalty, and terminal reward. The calculation formula for the multi-objective reward function is expressed as follows: ; Indicates an immediate reward; This indicates a path feasibility reward, providing a positive reward for effective placement actions. Provide an equal penalty for a valid placement action. ,Right now ; Indicates the payload alignment reward, relative to the trajectory tangent. and the direction of the target load The angle difference between them is proportional to the cosine of the angle difference, that is... , Indicates the alignment reward coefficient; Indicates curvature penalty when the path turns at an angle. Exceeding the maximum deflection angle In certain situations, a second penalty is imposed; otherwise, no reward or penalty is imposed. , Indicates the curvature penalty coefficient; A sparse but substantial reward is given upon successfully reaching the designated finish line; otherwise, no reward or penalty is imposed. , Indicates the terminal reward value; Indicates the first The reward weight of each reward component. .

[0012] Optionally, the reward weights are dynamically adjusted through the course learning strategy, with path feasibility rewards and... Higher weights guide the agent to learn and generate complete trajectories. During training, the weights of load alignment rewards and curvature penalties are gradually increased to guide the agent to optimize fiber alignment and suppress wrinkle defects. The formula for expressing the reward weights is as follows: ; In the formula, Indicates the first During the training step, the first... The reward weight of each reward component; Indicates the first Initial values ​​for the reward weights of each reward component; Indicates the first The final reward weight of each reward component; This represents the total number of predefined training steps.

[0013] Optionally, the soft actor-critic algorithm is a reinforcement learning method based on the maximum entropy principle, which simultaneously maximizes the agent's expected cumulative reward and policy entropy during training. The objective function of the soft actor-critic algorithm is expressed by the following formula: ; In the formula, where Representing the agent's strategy; Represents the mathematical expectation. Represents a state-action pair. Represents the policy-induced state-action distribution; Indicates the discount factor; This represents the immediate reward function obtained from the environment; This represents a temperature parameter used to control the weight of the policy entropy in the target. The policy entropy is represented by the observation space, the action space, and the multi-objective reward function, which are embedded in the soft actor-critic algorithm framework. The agent collects experience by repeatedly interacting with the training environment. The soft actor-critic algorithm updates the policy network and value network parameters of the agent to obtain the optimal trajectory planning strategy that minimizes wrinkle defects and control fiber angle deviation.

[0014] Secondly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the steps in the above-described reinforcement learning-based automatic fiber placement trajectory planning method for composite pressure vessels.

[0015] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the above-described reinforcement learning-based automatic fiber placement trajectory planning method for composite pressure vessels.

[0016] The beneficial effects of the technical solutions provided in this application include at least the following: This application provides an automatic fiber placement trajectory planning method for composite material pressure vessels based on reinforcement learning. The method triangulates the geometric model of the composite material pressure vessel and identifies high-curvature regions. Local mesh refinement is then performed on these high-curvature regions to generate a high-precision triangular discretized model. This allows for the approximation of arbitrary hypercurvature geometries, including variable-curvature heads and transition zones, using triangular mesh units without requiring reparameterization or surface splicing. Furthermore, local refinement enables near-linear growth of degrees of freedom, significantly improving the efficiency and accuracy of geometric modeling for complex surfaces. The method calculates the geometric attributes of each mesh unit and constructs a discrete map. Based on this discrete map, a reinforcement learning interactive environment is built. This environment integrates the minimum turning radius constraint of the automatic fiber placement process. After the agent performs an action, it updates its state on the triangular mesh and calculates the path turning angle after each state update. The minimum turning radius, a practical process constraint, is integrated into the interactive environment. Simultaneously, the feasible range of states and actions is represented through meshing, resulting in a well-defined and structurally clear state and action combination space. This method reduces the dimension of blind trial and error in the agent during training and makes its prediction of the value of each decision more reliable, thereby effectively improving the training efficiency and convergence effect of the reinforcement learning process. An observation space, action space, and multi-objective reward function are designed. In a reinforcement learning interactive environment, the agent is trained using a soft actor-critic algorithm. Based on the observation state, the agent outputs the layup angle action and iteratively optimizes under the guidance of the multi-objective reward function, learning the optimal trajectory planning strategy to minimize wrinkle defects and control fiber angle deviation. This guides the agent to learn strategies that satisfy manufacturing constraints and optimize structural performance, solving the trade-off between manufacturability and structural performance in traditional methods. For composite pressure vessels to be laid, the trained agent generates the optimal layup trajectory, and its wrinkle defect rate and fiber angle deviation are evaluated. Ultimately, a significant performance improvement is achieved, with an average reduction of 35.2% in wrinkle defects and fiber angle deviation controlled within the range of -2° to +1.5°. This demonstrates the comprehensive superiority of the automatic fiber layup trajectory planning method for composite pressure vessels based on reinforcement learning provided in this application. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 A flowchart illustrating the automatic fiber placement trajectory planning method for composite pressure vessels based on reinforcement learning provided in this application embodiment; Figure 2A schematic diagram illustrating the acquisition of state vectors for automatic filament placement trajectory planning based on triangular meshes, provided in an embodiment of this application; Figure 3 A schematic diagram of an agent-environment interaction framework for automatic filament placement trajectory planning based on reinforcement learning, provided in an embodiment of this application; Figure 4 This is a schematic diagram of the observation space composition of an intelligent agent provided in an embodiment of this application; Figure 5 A geometric representation of a minimum turning radius constraint provided in an embodiment of this application; Figure 6 A schematic diagram of the reward function components for automatic filament placement path planning provided in this application embodiment; Figure 7 This is a schematic diagram of the overall geometry of a pressure vessel model provided in an embodiment of this application; Figure 8 A schematic diagram of the learning curve of a reinforcement learning agent provided in an embodiment of this application; Figure 9 This is a schematic diagram comparing the effects of a traditional trajectory planning method in an embodiment of this application. Figure 10 A schematic diagram comparing trajectories obtained using a traditional fixed-angle method and a reinforcement learning-based method, provided as an embodiment of this application; Figure 11 A schematic diagram illustrating the comprehensive analysis of the angle deviation of the geodesic path fiber provided in this application embodiment; Figure 12 A schematic diagram illustrating the comprehensive analysis of wrinkle defect rate using a fixed angle method, provided in an embodiment of this application; Figure 13 A schematic diagram illustrating the performance analysis of trajectory planning based on reinforcement learning, provided for an embodiment of this application; Figure 14 This application provides a schematic diagram of the fiber angle distribution along an overall laying trajectory. Figure 15 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0020] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0021] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0022] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0023] In view of the current problems in the research on automatic fiber placement trajectory planning of composite pressure vessels in the field of intelligent control technology of composite materials, this application provides an automatic fiber placement trajectory planning method for composite pressure vessels based on reinforcement learning.

[0024] The technical solution of this application is described below, starting with the method embodiments.

[0025] Please refer to Figure 1 It illustrates a flowchart of the automatic fiber placement trajectory planning method for composite pressure vessels based on reinforcement learning provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes at least the following steps S110 to S140.

[0026] Step S110: Triangulate the geometric model of the composite pressure vessel and identify high curvature regions. Refine the local mesh for the high curvature regions to generate a high-precision triangular discretized model. Calculate the geometric properties of each mesh unit and construct a discrete map.

[0027] In this embodiment, the commercial CAD software CATIA is used to triangulate the overall geometric model of the composite pressure vessel, generating a triangular surface mesh. After exporting the triangular surface mesh, secondary processing is performed to calculate the curvature of each mesh element in the triangular surface mesh. High curvature regions with curvature values ​​exceeding a preset curvature threshold are identified, and local mesh refinement is performed on these high curvature regions. By reducing the size of the triangular mesh elements within the high curvature regions and performing conformal adjustments, a high-precision triangular discretization model is generated, maintaining element quality while completing the re-meshing. The data of each triangular mesh element in the high-precision triangular discretization model is organized to construct a reinforcement learning interactive environment. Specifically, for each triangular mesh element in the high-precision triangular discretization model, the geometric attributes of each triangular mesh element are pre-calculated and stored, including node coordinates, unit surface normal vector, curvature metric value, and related geodesic directions. After calculation, all triangular mesh elements in the high-precision triangular discretization model are constructed into a structured discrete map according to their topological connections.

[0028] The technical solution provided in this application discretizes the geometric model of a composite pressure vessel using triangular meshes. Compared with traditional parametric surface descriptions such as NURBS or spline surfaces, triangular mesh discretization offers several advantages for automatic wire-laying path planning on pressure vessels. First, triangular elements can approximate arbitrary hypercurvature geometries, including variable curvature heads and transition regions, without requiring reparameterization or surface splicing; local variations in Gaussian curvature can be captured simply by adjusting vertex positions and connectivity. Second, local refinement is straightforward; only areas with high curvature, tight geodesic deflections, or cut-out interruptions need to have their element size reduced, while maintaining a coarse mesh elsewhere, resulting in a non-uniform but structured computational domain with a near-linear increase in degrees of freedom. Furthermore, the geometric data required for path planning—adjacent regions, local surface normals, and geodesic directions—is simplified into simple operations on sparse adjacency lists and element-by-element measurements, enabling highly efficient computation and easy parallelization. Finally, the establishment of the triangular mesh discretization model provides an interface for subsequent optimization algorithms. Specifically, key manufacturing constraints in the automated fiber placement process, such as allowable fiber angle deviations, fiber bundle gap / overlap limits, and turning radii, can be conveniently defined on each specific triangular mesh cell. This transforms the mathematical description of the entire optimization problem into a sparse constraint matrix and reduces the complexity of traditional global nonlinear programming formulas. Furthermore, this decision-making mechanism based on individual mesh cells enables direct policy mapping, meaning actions are output directly from the current cell without any intermediate continuous parameterization layers, thus reducing potential errors when approximating the ideal policy. Simultaneously, the reward function can be finely defined at the level of each mesh cell. For example, penalties can be imposed for local fiber bundle gaps, fiber bundle overlaps, turning behaviors exceeding the allowable curvature threshold, or deviations between the fiber direction and the main bearing direction within the current cell. These rewards or penalties for individual cells are explicitly accumulated across all cells visited by the agent, thus more accurately allocating the credit of the final result to each specific decision that led to that result. Compared to searching in a continuous parameter space, this clearly structured state and action space, composed of a finite number of grid cells, significantly reduces the dimensionality the agent needs to explore and makes its estimation of the value of each state more stable, thereby effectively improving the convergence speed and performance of the overall reinforcement learning training. Furthermore, the learned policy is interpretable; the agent's preferred placement direction, high-risk regions, and failure modes can all be directly visualized on the triangular mesh model, facilitating the verification and evaluation of the rationality and reliability of the agent's decisions for specific manufacturing constraints and structural performance standards.

[0029] Step S120: Construct a reinforcement learning interactive environment based on the discrete map. The reinforcement learning interactive environment integrates the minimum turning radius constraint of the automatic wire laying process. After the agent performs an action, it completes the state update on the triangular mesh and calculates the path turning angle after each state update.

[0030] In this embodiment, a reinforcement learning interactive environment is constructed based on a discrete map. Specifically, a structured discrete map is loaded as the geometric basis of the reinforcement learning interactive environment. On this discrete map, the starting and target grid cells for the preset trajectory planning are clearly defined, thereby defining the navigation task that the agent needs to complete. During the training and subsequent inference phases, the agent will initiate a query request to the interactive environment at each of its current locations. The interactive environment will respond to this request and return a corresponding state vector, thereby ensuring that the agent can explicitly consider the local geometry of its current location and the fiber placement direction allowed by its process as key factors in each decision of trajectory path planning. For example, please refer to [reference needed]. Figure 2 This illustration shows a schematic diagram of obtaining the state vector for automatic filament placement trajectory planning based on a triangular mesh, provided by an embodiment of this application. Further, a state update logic is defined. In the reinforcement learning interactive environment, each action of the agent (corresponding to a local placement angle) triggers a state update. Specifically, the interactive environment, based on the local placement angle determined by the agent in the current mesh cell according to the current action, calculates the geometric intersection point with the boundary of the current mesh cell to accurately determine the adjacent mesh cell the agent should enter in the next moment, thereby completing the state update. After the state update, the geometric attributes of the updated state (i.e., the mesh cell entered after the update) are quickly obtained by querying the discrete map.

[0031] In this embodiment, the key physical constraints of the automated fiber placement process, namely the minimum turning radius constraint, are integrated into the reinforcement learning interactive environment. After each state update, the interactive environment automatically calculates the path turning angle between the current action direction and the previous action direction to ensure that the curvature of the trajectory meets manufacturing feasibility requirements. Furthermore, an evaluation and feedback mechanism for the interactive environment, namely reward signal generation, is implemented. A multi-objective reward function is integrated within the interactive environment. After each state update, the interactive environment comprehensively evaluates the result of that action step and calculates an immediate reward signal, which is then returned to the agent. Finally, all functions are encapsulated into a standardized reinforcement learning environment interface. This interface receives action instructions from the agent, sequentially drives the state update, constraint check, and reward calculation processes, and ultimately returns the updated state observation, immediate reward value, and other flag information (such as whether the process has terminated) to the agent. This highly integrated simulation environment, deeply incorporating knowledge from the automated fiber placement domain, provides a reliable foundation for the efficient training of the agent. For an example, please refer to [link to example]. Figure 3This illustration shows a schematic diagram of an agent-environment interaction framework for automatic fiber placement trajectory planning based on reinforcement learning, provided in an embodiment of this application. The core objective of the agent is to master the optimal decision-making rule. This objective is achieved through a cyclical process. Specifically, the agent first analyzes its current state information, such as its specific location and orientation on the pressure vessel surface, and selects an action accordingly. This action determines the next placement point on the trajectory. Subsequently, the interactive environment, representing the automatic fiber placement process simulation, updates the agent to a new state based on this action and simultaneously feeds back a reward signal. This signal is mainly used to give positive rewards for actions that promote effective placement and negative penalties for actions that cause wrinkles. Through repeated iterations and continuous optimization of the decision-making strategy to maximize long-term cumulative rewards, the agent gradually learns to generate globally optimal placement trajectories with fewer defects.

[0032] Step S130: Design the observation space, action space, and multi-objective reward function. In the reinforcement learning interactive environment, train the agent using the soft actor-critic algorithm. The agent outputs the laying angle action according to the observation state and performs iterative optimization under the guidance of the multi-objective reward function to learn the optimal trajectory planning strategy that minimizes wrinkle defects and controls fiber angle deviation.

[0033] In this embodiment, to enable the agent to make optimized decisions at each step of trajectory planning, a 15-dimensional continuous observation space is constructed. This observation space is built based on five key three-dimensional vectors, which together constitute a complete description of the agent's current geometric environment and historical state. For an example, please refer to... Figure 4 It illustrates a schematic diagram of the observation space composition of an intelligent agent according to an embodiment of this application, including a position vector, a surface normal vector, a principal direction vector, a geodesic direction vector, and a previous direction vector, wherein the position vector... The surface normal vector represents the agent's current precise position. The position vector represents the normal vector of the current mesh cell surface, describing the orientation of the triangular facet directly below the agent. The position vector and surface normal vector together define the geometric state of the agent on the complex surface. To obtain structurally sound and manufacturable decisions, the agent is equipped with basic orientation cues, including the principal direction vector. This indicates the target bearing direction on the local surface, serving as the primary alignment reference; simultaneously, the geodesic direction vector... Representing the local shortest path provides crucial information for assessing and avoiding potential wrinkle formation. Finally, to maintain trajectory continuity, the state representation incorporates the previous direction vector. The vector represents the direction of the preceding layer. Based on the historical information provided by this vector, the agent can calculate the turning angle and maintain the overall smoothness of the trajectory. This observation space lays a solid foundation for determining the agent's subsequent actions.

[0034] In this embodiment, the motion output of the motion space is the deflection angle of the fiber placement at the current placement point. Furthermore, in automated fiber placement processes, strictly controlling the curvature of the placement trajectory is crucial to preventing defects. Geodesic curvature is a key factor affecting the smooth turning of the placement path, playing a central role in ensuring that the strain experienced by the material during placement does not exceed its tolerance limit. The physical relationship between the maximum strain capacity that the prepreg fiber bundle can withstand and the minimum allowable turning radius during placement is expressed by the following formula: ; In the formula, Indicates maximum strain capacity; Indicates the width of the prepreg filament bundle; This represents the minimum turning radius. During path planning, this material constraint is directly translated into an explicit constraint on the deflection angle between adjacent paved path segments, i.e., the minimum turning radius constraint. The formula for this minimum turning radius constraint is expressed as follows: ; In the formula, Indicates the deflection angle between adjacent path segments; Indicates the maximum deflection angle; Indicates the local path length; This represents the minimum turning radius; please refer to the example. Figure 5 It shows a geometric representation of a minimum turning radius constraint provided in an embodiment of this application, with adjacent path segments. and Deflection angle between Subject to minimum turning radius and effective path length The constraints are then addressed. Finally, the aforementioned geometric and material constraints are integrated into a reinforcement learning framework, under which the agent learns to strictly adhere to both geometric and material constraints while optimizing fiber placement trajectories. The action space for the agent's decision-making regarding placement angles is modeled using a Gaussian policy function. Action selection is sampled from the action space through a Gaussian distribution, allowing the agent to dynamically and extensively explore various possible actions in the early stages of training. As the agent accumulates experience, its decisions become more deterministic and optimized, and the Gaussian distribution gradually shrinks according to a predefined decay strategy, reducing the scope and randomness of the exploration.

[0035] In this embodiment of the application, in the reinforcement learning interactive environment, after the current observation state is transferred to the next observation state through the current placement angle action, an immediate reward is calculated through a multi-objective reward function. This multi-objective reward function balances key manufacturing constraints and structural performance requirements, guiding the agent to generate the optimal automatic fiber placement trajectory. The multi-objective reward function is determined by weighted summation of path feasibility reward, load alignment reward, curvature penalty, and terminal reward. The calculation formula of the multi-objective reward function is expressed by the following formula: ; In the formula, Indicates an immediate reward; The path feasibility bonus is primarily used to ensure that the generated trajectory is confined within the specified manufacturing surface. Provide positive rewards for effective placement actions. Provide an equal penalty for a valid placement action. ,Right now ; This indicates a load alignment bonus, used to encourage fiber trajectories to align with the principal stress directions. Tangent to the trajectory and the direction of the target load The angle difference between them is proportional to the cosine of the angle, thus maximizing the reward under perfect alignment conditions. , Indicates the alignment reward coefficient; Curvature penalty refers to the penalty applied to maintain manufacturability and prevent defects such as fiber wrinkles. Penalize excessive local curvature when the path turns at an angle Exceeding the maximum deflection angle In certain situations, a second penalty is imposed; otherwise, no reward or penalty is imposed. This avoids illegal placement actions and promotes the generation of smooth and continuous trajectories. Indicates the curvature penalty coefficient; This represents the terminal reward. To encourage the agent to generate a complete trajectory, a sparse but considerable reward is given upon successfully reaching the designated endpoint; otherwise, no reward or penalty is imposed. , Indicates the terminal reward value; Indicates the first The reward weight of each reward component. The technical solutions provided in this application's embodiments use various reward components to solve specific problems in trajectory generation. For examples, please refer to... Figure 6This illustration shows a schematic diagram of the reward function components for automatic filament placement path planning provided in an embodiment of this application. Based on path feasibility assessment (a) corresponding to path feasibility reward, load direction alignment measurement (b) corresponding to load alignment reward, turning angle constraint (c) corresponding to curvature penalty, and terminal state recognition (d) corresponding to terminal reward, this integrated reward structure establishes a balanced learning objective to coordinate the competitive requirements of automatic filament placement by appropriately adjusting the reward weights. The system can adjust different aspects of trajectory quality according to requirements.

[0036] In this embodiment, during training, the reward weights are dynamically adjusted through a course learning strategy to prevent the agent from being overwhelmed by multiple complex objectives simultaneously, further stabilizing and accelerating the learning process. In the early stages of training, this allows the agent to achieve learning through rewarding path feasibility and... Higher weights guide the agent to focus on fundamental tasks, such as learning to generate effective and complete trajectories. As training progresses, the focus gradually shifts to more refined objectives, progressively increasing the weights of load alignment rewards and curvature penalties. This guides the agent to optimize fiber alignment and suppress wrinkle defects. This structured learning process improves training efficiency and the robustness of the final policy. The reward weights evolve from their initial values ​​to their final values ​​within a predefined total number of training steps. The formula for expressing the reward weights is as follows: ; In the formula, Indicates the first During the training step, the first... The reward weight of each reward component; Indicates the first Initial values ​​for the reward weights of each reward component; Indicates the first The final reward weight of each reward component; This represents the total number of predefined training steps.

[0037] In this embodiment, the soft actor-critic algorithm is a reinforcement learning method based on the maximum entropy principle. Its key feature is that it simultaneously maximizes the agent's expected cumulative reward and policy entropy during training. Incorporating entropy maximization incentivizes the agent to explore more possibilities, thereby enhancing the stability of the entire training process and effectively preventing premature entrapment in local optima. The objective function of the soft actor-critic algorithm is expressed by the following formula: ; In the formula, where This represents the agent's decision-making strategy; Represents the mathematical expectation. Represents a state-action pair. Indicating in strategy Induced state-action distribution; This represents the discount factor, used to determine the current importance of future rewards; This represents the immediate reward function obtained from the environment at each step. This represents a temperature parameter used to control the weight of the policy entropy in the target. Represents policy entropy, used to quantify the state given by the policy. Next, strategy The randomness or uncertainty of composite material layup is addressed. During the generation of the main layup trajectory, the agent continuously learns and develops a decision-making strategy that adapts to the geometric features described by the discrete mesh environment and meets structural load requirements. Compared to traditional methods, this reinforcement learning method can dynamically evaluate the curvature changes and geometric complexity of the model, adaptively adjusting and improving its strategy, thereby significantly reducing the possibility of defects during manufacturing. Importantly, this reinforcement learning method balances the exploration of new actions with the utilization of known experience through an entropy regularization mechanism, enabling the discovery of stable and reliable layup strategies even when facing extremely complex surface geometry. The agent continuously interacts with the discrete triangular mesh model, and the ultimately learned strategy simultaneously satisfies manufacturing process constraints and the design goal of optimizing structural performance. In a specific embodiment, the observation space, action space, and multi-objective reward function are embedded in the soft actor-critic algorithm framework. The agent collects experience through repeated interaction with the training environment, and the soft actor-critic algorithm updates the agent's policy network and value network parameters, ultimately obtaining the optimal trajectory planning strategy that minimizes wrinkle defects and controls fiber angle deviation.

[0038] Step S140: For the composite pressure vessel to be laid, the optimal laying trajectory is generated by the trained agent, and its wrinkle defect rate and fiber angle deviation are evaluated.

[0039] In one specific embodiment, the geometric and process parameters are first set; for example, please refer to [reference needed]. Figure 7 It shows a schematic diagram of the overall geometry of a pressure vessel model provided in an embodiment of this application, wherein the geometry of its ellipsoidal head is determined by the major semi-axis. and short half shaft By definition, the pole hole is located at the axial height. Based on the physical limitations of the automated fiber placement process, the minimum turning radius of the placement system is set to [value missing]. This leads to the derivation of the maximum permissible deflection angle that must be followed in path planning. The angle is approximately 0.19°. The agent's action space, i.e., the exploration range of the placement angle, is limited to a continuous interval of ±5°. Further, the agent is trained and its convergence is verified. In the above parameterized environment, policy training based on the soft actor-critic algorithm is performed. The agent updates its network parameters by repeatedly interacting with the environment. The randomness of its action exploration decreases according to a predetermined strategy as the number of training steps increases. Please refer to [reference needed]. Figure 8 The diagram illustrates the learning curve of a reinforcement learning agent provided in this embodiment. The cumulative reward obtained by the agent increases rapidly with the number of training iterations and eventually stabilizes at a high level, indicating that the agent has successfully learned an effective decision-making strategy and achieved stable convergence.

[0040] Furthermore, the generated trajectory undergoes quantitative performance evaluation. For the composite pressure vessel to be laid, an optimal laying trajectory is generated by a trained agent. The performance evaluation mainly focuses on two core indicators: wrinkle defect rate and fiber angle deviation. The wrinkle defect rate is calculated as the percentage of all line segments in the trajectory that violate the maximum allowable deflection angle constraint out of the total length. The fiber angle deviation is evaluated by statistically analyzing the distribution of the angle difference between the actual laying direction and the preset target load direction at each point on the trajectory. Finally, the automatic fiber laying trajectory planning method provided in this application embodiment is systematically compared and analyzed with traditional methods. Specifically, it is compared with the fixed angle method and the geodesic method. For qualitative comparisons, please refer to... Figure 9 The diagram illustrates a comparison of the effects of a traditional trajectory planning method provided in this application. It shows that the trajectory generated by the fixed-angle method on the left exhibits significant abrupt changes and discontinuities in direction at the boundary of the triangular mesh, while the trajectory generated by the geodesic method on the right, although globally smooth, deviates severely from the designed principal stress direction. Please refer to... Figure 10 The illustration shows a comparison diagram of trajectories generated by a traditional fixed-angle method and a reinforcement learning-based method, as provided in this application embodiment. The comparison further shows that the trajectory generated by the method provided in this application embodiment is smoother and more reasonable in high-curvature regions. This is because the agent can intelligently adjust its trajectory through high-curvature regions without violating the minimum turning radius constraint. For quantitative analysis, please refer to... Figures 11 to 14 This document illustrates a comprehensive analysis of fiber angle deviation along a geodesic path, a comprehensive analysis of wrinkle defect rate using a fixed-angle method, a performance analysis of trajectory planning based on reinforcement learning, and a schematic diagram of fiber angle distribution for an overall laying trajectory, all provided in this application. A comparison shows that the fiber angle distribution using the geodesic method is discrete and has large deviations, while the method provided in this application can accurately control the angle near the target value. The fixed-angle method exhibits a high defect rate distribution under different target angles. Table 1 compares the defect rates of the traditional method and the reinforcement learning-based method under different target angles.

[0041] Table 1 (Comparison of defect rates between traditional methods and reinforcement learning-based methods from different objective perspectives) ; As can be seen, the automatic fiber layup trajectory planning method based on reinforcement learning provided in this application significantly reduces the average wrinkle defect rate from 15.6% to 10.3% compared with the traditional fixed angle method, a reduction of 35.2%. At the same time, it successfully controls the fiber layup angle deviation within the range of -2° to +1.5°, comprehensively verifying its significant effect in collaboratively optimizing manufacturability and structural performance.

[0042] In summary, the reinforcement learning-based automatic fiber placement trajectory planning method for composite material pressure vessels provided in this application triangulates the geometric model of the composite material pressure vessel and identifies high-curvature regions. Local mesh refinement is then performed on these high-curvature regions to generate a high-precision triangular discretized model. This allows for the approximation of arbitrary hypercurvature geometries, including variable-curvature heads and transition zones, using triangular mesh units without requiring reparameterization or surface splicing. Furthermore, local refinement enables near-linear growth of degrees of freedom, significantly improving the efficiency and accuracy of geometric modeling for complex surfaces. The method calculates the geometric attributes of each mesh unit and constructs a discrete map. Based on this discrete map, a reinforcement learning interactive environment is built. This environment integrates the minimum turning radius constraint of the automatic fiber placement process. After the agent executes an action, it updates its state on the triangular mesh and calculates the path turning angle after each state update. This integrates the actual process constraint of the minimum turning radius into the interactive environment. Simultaneously, the feasible range of states and actions is represented through meshing, resulting in a well-defined and structurally clear set of states and actions. The design of the observation space, action space, and multi-objective reward function reduces the dimensionality of blind trial and error during training, making the prediction of the value of each decision more reliable, thus effectively improving the training efficiency and convergence effect of the reinforcement learning process. The design incorporates observation space, action space, and a multi-objective reward function. In the reinforcement learning interactive environment, the agent is trained using a soft actor-critic algorithm. Based on the observation state, the agent outputs the layup angle action and iteratively optimizes under the guidance of the multi-objective reward function, learning the optimal trajectory planning strategy to minimize wrinkle defects and control fiber angle deviation. This guides the agent to learn strategies that satisfy manufacturing constraints and optimize structural performance, solving the trade-off between manufacturability and structural performance in traditional methods. For composite pressure vessels to be laid, the trained agent generates the optimal layup trajectory, and its wrinkle defect rate and fiber angle deviation are evaluated. Ultimately, a significant performance improvement is achieved, with an average reduction of 35.2% in wrinkle defects and fiber angle deviation controlled within the range of -2° to +1.5°. This demonstrates the comprehensive superiority of the automatic fiber layup trajectory planning method for composite pressure vessels based on reinforcement learning provided in this application.

[0043] It should be noted that, in the embodiments of this application, if the above-mentioned automatic fiber placement trajectory planning method for composite material pressure vessels based on reinforcement learning is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0044] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by a processor, this computer program implements the steps in any of the reinforcement learning-based automatic fiber placement trajectory planning methods for composite pressure vessels described in the above embodiments. Correspondingly, embodiments of this application also provide a computer program product, which, when executed by a processor of an electronic device, is used to implement the steps in any of the reinforcement learning-based automatic fiber placement trajectory planning methods for composite pressure vessels described in the above embodiments.

[0045] Based on the same technical concept, this application provides an electronic device for implementing the reinforcement learning-based automatic fiber placement trajectory planning method for composite material pressure vessels described in the above method embodiments. Figure 15 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 15 As shown, the electronic device 1500 includes a memory 1510 and a processor 1520. The memory 1510 stores a computer program that can run on the processor 1520. When the processor 1520 executes the program, it implements the steps in any of the reinforcement learning-based automatic fiber placement trajectory planning methods for composite pressure vessels according to the embodiments of this application.

[0046] The memory 1510 is configured to store instructions and applications executable by the processor 1520, and can also cache data to be processed or already processed by the processor 1520 and various modules in the electronic device (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0047] When processor 1520 executes the program, it implements the steps of the reinforcement learning-based automatic fiber placement trajectory planning method for composite material pressure vessels described above. Processor 1520 typically controls the overall operation of electronics 1500.

[0048] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.

[0049] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various electronic devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0050] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0051] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0052] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0053] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0054] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0055] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0056] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the device automatic test line to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0057] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0058] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0059] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for automatic fiber placement trajectory planning in composite material pressure vessels based on reinforcement learning, characterized in that, The method includes: The geometric model of the composite pressure vessel is triangulated and high curvature regions are identified. Local mesh refinement is performed on the high curvature regions to generate a high-precision triangular discretized model. The geometric properties of each mesh unit are calculated and a discrete map is constructed. A reinforcement learning interactive environment is constructed based on the discrete map. The minimum turning radius constraint of the automatic filament placement process is integrated in the reinforcement learning interactive environment. After the agent performs an action, it completes the state update on the triangular mesh and calculates the path turning angle after each state update. The design includes an observation space, an action space, and a multi-objective reward function. In the reinforcement learning interactive environment, the agent is trained using a soft actor-critic algorithm. The agent outputs a laying angle action based on the observation state and iteratively optimizes under the guidance of the multi-objective reward function to learn the optimal trajectory planning strategy that minimizes wrinkle defects and controls fiber angle deviation. For composite pressure vessels to be laid, an optimal layup trajectory is generated by a trained agent, and its wrinkle defect rate and fiber angle deviation are evaluated.

2. The method according to claim 1, characterized in that, The process involves triangulation of the geometric model of the composite pressure vessel and identification of high-curvature regions. Local mesh refinement is then performed on these high-curvature regions to generate a high-precision triangular discretized model. The geometric properties of each mesh element are calculated, and a discrete map is constructed. This includes: The geometric model of the composite pressure vessel is triangulated to generate a triangular surface mesh. Calculate the curvature of each grid cell in the triangular surface mesh, and identify high curvature regions where the curvature value is higher than a preset curvature threshold; For the high curvature region, local mesh refinement is performed by reducing the size of the triangular mesh units in the high curvature region and making shape-preserving adjustments to generate a high-precision triangular discretization model; For each triangular mesh element in the high-precision triangular discretization model, the geometric properties of each triangular mesh element are pre-calculated and stored. The geometric properties include node coordinates, unit surface normal vector, curvature metric value, and related geodesic direction. All triangular mesh cells in the high-precision triangular discretization model are constructed into a structured discrete map according to their topological connections.

3. The method according to claim 1, characterized in that, The reinforcement learning interactive environment is constructed based on the discrete map. This environment integrates the minimum turning radius constraint of the automated filament placement process. After the agent performs an action, it updates its state on the triangular mesh and calculates the path turning angle after each state update, including: Using the discrete map as the geometric basis, and pre-setting the starting and target grid cells for trajectory planning, a reinforcement learning interactive environment is constructed. In the reinforcement learning interactive environment, the agent determines the next adjacent grid cell to enter by calculating the intersection point with the boundary of the current grid cell based on the local laying angle determined by the current action, thereby completing the state update, and obtaining the geometric attributes of the updated state based on the discrete map. The minimum turning radius constraint of the automatic filament placement process is integrated into the reinforcement learning interactive environment. After the state update is completed, the path turning angle between the current action direction and the previous action direction is calculated. After the environment receives the action instructions from the agent, and the state update and constraint check are completed, it returns the updated state and an immediate reward signal to the agent.

4. The method according to claim 1, characterized in that, The observation space is a 15-dimensional continuous vector, constructed by splicing together five three-dimensional vectors. The five three-dimensional vectors are: a position vector representing the current position of the agent, a surface normal vector representing the surface normal of the current grid cell, a main direction vector representing the local target bearing direction, a geodesic direction vector representing the local shortest path, and a previous direction vector representing the previous laying direction. The action output of the motion space is the deflection angle of the fiber placement at the current placement point. This deflection angle is limited by the minimum turning radius constraint of the automatic fiber placement process, which is expressed by the following formula: ; In the formula, Indicates the deflection angle between adjacent path segments; Indicates the maximum deflection angle; Indicates the local path length; Indicates the minimum turning radius; The action space is modeled using a Gaussian policy function. Action selection is sampled from the action space using a Gaussian distribution. During training, the Gaussian distribution gradually shrinks according to a predefined decay strategy.

5. The method according to claim 1, characterized in that, In the reinforcement learning interactive environment, after the current observation state is transferred to the next observation state through the current placement angle action, the immediate reward is calculated through a multi-objective reward function. The multi-objective reward function is determined by weighted summation of path feasibility reward, payload alignment reward, curvature penalty, and terminal reward. The calculation formula of the multi-objective reward function is expressed by the following formula: ; In the formula, Indicates an immediate reward; This indicates a path feasibility reward, providing a positive reward for effective placement actions. Provide an equal penalty for a valid placement action. ,Right now ; Indicates the payload alignment reward, relative to the trajectory tangent. and the direction of the target load The angle difference between them is proportional to the cosine of the angle difference, that is... , Indicates the alignment reward coefficient; Indicates curvature penalty when the path turns at an angle. Exceeding the maximum deflection angle In certain situations, a second penalty is imposed; otherwise, no reward or penalty is imposed. , Indicates the curvature penalty coefficient; A sparse but substantial reward is given upon successfully reaching the designated finish line; otherwise, no reward or penalty is imposed. , Indicates the terminal reward value; Indicates the first The reward weight of each reward component. .

6. The method according to claim 5, characterized in that, The reward weights are dynamically adjusted based on the course learning strategy, with path feasibility rewards assigned during the initial training phase. Higher weights guide the agent to learn and generate complete trajectories. During training, the weights of load alignment rewards and curvature penalties are gradually increased to guide the agent to optimize fiber alignment and suppress wrinkle defects. The formula for expressing the reward weights is as follows: ; In the formula, Indicates the first During the training step, the first... The reward weight of each reward component; Indicates the first Initial values ​​for the reward weights of each reward component; Indicates the first The final reward weight of each reward component; This represents the total number of predefined training steps.

7. The method according to claim 1, characterized in that, The soft actor-critic algorithm is a reinforcement learning method based on the maximum entropy principle. During training, it simultaneously maximizes the agent's expected cumulative reward and policy entropy. The objective function of the soft actor-critic algorithm is expressed by the following formula: ; In the formula, where Representing the agent's strategy; Represents the mathematical expectation. Represents a state-action pair. Represents the policy-induced state-action distribution; Indicates the discount factor; This represents the immediate reward function obtained from the environment; This represents a temperature parameter used to control the weight of the policy entropy in the target. Represents policy entropy; The observation space, the action space, and the multi-objective reward function are embedded in the soft actor-critic algorithm framework. The agent collects experience by repeatedly interacting with the training environment. The soft actor-critic algorithm updates the policy network and value network parameters of the agent to obtain the optimal trajectory planning strategy that minimizes wrinkle defects and control fiber angle deviation.

8. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.