Composite material pultrusion frame data acquisition management system based on reinforcement learning
By using sensor arrays and nonlinear Kalman filtering algorithms for data acquisition and digital twin virtual simulation, combined with the SoftActor-Critic algorithm and artificial bee colony hyperparameter optimization, the problems of insufficient real-time synchronization and model adaptability in composite material pultrusion processes were solved. This enabled efficient adaptive control and rapid anomaly warning, improving the stability of the production process and the accuracy of decision-making.
Patent Information
- Application Number
- CN202511094356.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing digital twin solutions in composite material pultrusion processes suffer from insufficient real-time synchronization and model adaptive performance. Hyperparameter optimization methods have long calculation cycles and are difficult to integrate with online control systems. Traditional anomaly detection methods cannot achieve real-time online identification and early warning. Data management platforms and intelligent control lack deep integration, failing to form closed-loop decision support.
A sensor array and a nonlinear Kalman filter algorithm are used for real-time data acquisition to construct a digital twin virtual simulation model. Adaptive control is performed based on the SoftActor-Critic algorithm, and hyperparameter optimization of artificial bee colonies is combined to achieve rapid anomaly early warning and closed-loop decision support.
It enables real-time high-precision parameter acquisition and online simulation of composite material pultrusion process, rapid optimization of adaptive control strategy, improves the stability and reliability of production process, shortens model convergence time by 50%, and improves the response speed and decision accuracy of anomaly identification and early warning.
Smart Images

Figure CN120932789A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent manufacturing technology, and in particular to a data acquisition and management system for composite material pultruded borders based on reinforcement learning. Background Technology
[0002] Pultrusion technology for composite materials has been widely used in aerospace, rail transportation, construction, and new energy fields due to its continuous forming and high specific strength characteristics. With the development of intelligent manufacturing and the Industrial Internet, achieving digital, visual, and intelligent control of the pultrusion process has become a research focus. Digital twin technology provides support for online monitoring and simulation optimization by mapping physical equipment and process flows. Existing digital twin solutions mainly rely on finite element models or mapping based on empirical formulas, which have limitations in real-time synchronization of process parameters and model adaptive performance.
[0003] Reinforcement learning has demonstrated advantages in adaptability and robustness in industrial process control. In particular, the Soft Actor-Critic algorithm, due to its ability to combine exploration and exploitation in a continuous action space, has been applied to the online control of some continuous forming processes. However, existing reinforcement learning control systems often employ fixed or offline optimized hyperparameters, which cannot meet the real-time control requirements of multivariate coupling and dynamic changes in pultrusion processes.
[0004] Hyperparameter optimization methods, such as the artificial bee colony algorithm, can find global near-optimal solutions in complex search spaces, but in most industrial applications they are only used as offline optimization tools. They have long calculation cycles and are difficult to integrate with closed-loop online control systems, resulting in hyperparameter adjustments lagging behind process changes.
[0005] Process anomaly early warning and decision support are crucial for ensuring production stability. Traditional anomaly detection methods based on experience thresholds or statistical analysis suffer from time delays in identifying and responding to sudden process deviations. While modern machine learning methods have improved detection accuracy, they often rely on offline data training and cannot achieve real-time online identification and early warning.
[0006] The data management platform, within a framework that integrates edge computing and cloud computing, provides the foundation for storing and analyzing massive amounts of process data. However, the existing platform lacks deep integration with intelligent control and anomaly early warning systems, thus failing to form a closed-loop decision support system.
[0007] Therefore, how to provide a data acquisition and management system for composite material pultruded borders based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose a data acquisition and management system for composite material pultrusion frames based on reinforcement learning. This invention has the technical advantages of real-time high-precision process parameter acquisition and digital twin online simulation, adaptive control and real-time hyperparameter optimization, rapid anomaly early warning and closed-loop auxiliary decision support.
[0009] The reinforcement learning-based composite material pultruded frame data acquisition and management system according to an embodiment of the present invention includes the following modules:
[0010] The data sensing and acquisition unit, consisting of a sensor array and a nonlinear Kalman filter algorithm, acquires temperature, pressure, speed, and pultrusion force parameters during the composite material pultrusion process and generates a process data stream.
[0011] The digital twin virtual simulation unit receives the process data stream, constructs a digital twin virtual model of the composite material pultrusion process, and outputs virtual-real interaction data.
[0012] The adaptive intelligent control unit, built on the SoftActor-Critic algorithm, uses the product quality stability and production efficiency of the extruded frame as a composite reward function to generate control decision data based on virtual and real interaction data.
[0013] The artificial bee colony hyperparameter optimization unit is used to optimize the hyperparameters of the SoftActor-Critic algorithm based on control decision data using the artificial bee colony algorithm, and output the optimized hyperparameters.
[0014] An adaptive intelligent decision-making and anomaly early warning unit is used to identify and output abnormal process parameter data during the pultrusion of composite materials;
[0015] The data management and decision support unit is used to receive and store the process data stream, control decision data, and abnormal process parameter data, generate a data feature library, and output decision support data.
[0016] Optionally, modules can be integrated using the following methods:
[0017] S1. Collect temperature, pressure, speed and pultrusion force parameters during the composite material pultrusion process through a sensor array, and process them through a nonlinear Kalman filter algorithm to generate a process data stream;
[0018] S2. Input the process data stream into the digital twin virtual simulation unit to construct a digital twin virtual model of the composite material pultrusion process and output virtual-real interaction data;
[0019] S3. Based on the virtual-real interaction data, a reinforcement learning model is established using the SoftActor-Critic algorithm, with the product quality stability and production efficiency of the extruded frame as a composite reward function, to generate control decision data.
[0020] S4. Based on the control decision data, optimize the reinforcement learning model using the artificial bee colony algorithm and output the optimized hyperparameters.
[0021] S5. Analyze the process data stream and virtual-real interaction data to identify and output abnormal process parameter data in the composite material pultrusion process.
[0022] S6. Receive and store the process data stream, control decision data and abnormal process parameter data, construct a data feature library, and output auxiliary decision data.
[0023] Optionally, S1 specifically includes:
[0024] S11. The temperature values of the resin melt and multiple spatial points inside the mold cavity during the composite material pultrusion process are measured in real time and in different regions using a temperature sensor array to form a temperature parameter matrix.
[0025] S12. The pressure values of multiple key sections preset at the mold inlet, outlet and inside the mold cavity during the composite material pultrusion process are measured in real time and in stages by a pressure sensor array to form a pressure parameter matrix.
[0026] S13. The operating speed of the composite material traction device and the instantaneous speed at multiple points on the traction trajectory are measured in real time through a speed sensor array to form a speed parameter matrix;
[0027] S14. The magnitude of the pultrusion force at multiple points during the pultrusion process of composite materials and its dynamic characteristics as the process progress are measured in real time by a pultrusion force sensor array, forming a pultrusion force parameter matrix.
[0028] S15. Synchronously align the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix according to the same timestamp, and align the spatial measurement positions of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix with the length direction of the pultrusion frame as the reference. Normalize and splice the values of each matrix after synchronization and alignment element by element in a unified spatiotemporal coordinate system to obtain a process state feature matrix with unified dimension and scale.
[0029] S16. Input the process state feature matrix into a nonlinear Kalman filter algorithm. Through dynamic iterative updating of the covariance of the state feature matrix, the process parameters are dynamically optimized and calibrated, and a process data stream is generated.
[0030] Optionally, S2 specifically includes:
[0031] S21. Based on the real-time data of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, establish a three-dimensional spatial coordinate system corresponding to the actual composite material pultrusion equipment;
[0032] S22. Based on the structural characteristics of the pultrusion process frame of composite materials, determine the discretization nodes of the model and map each node to a specific position in the three-dimensional spatial coordinate system.
[0033] S23. Using the discretized nodes of the model as the reference position, the real-time data in the temperature parameter matrix, pressure parameter matrix, velocity parameter matrix and pultrusion force parameter matrix are mapped to the corresponding nodes one by one according to the spatial distance weighted interpolation method. The material properties and process characteristics at the discretized nodes of the model are used as constraints. Spatial difference correction and parameter matching between nodes are performed on the mapped real-time data to form a set of node parameter distributions and construct a digital twin virtual model of composite material pultrusion process.
[0034] S24. Based on the temperature, pressure, speed and pultrusion force parameters of the nodes in the digital twin virtual model, calculate the spatial distribution characteristics of the parameters node by node and update them continuously over time to obtain the spatial distribution matrix of the process parameters.
[0035] S25. Based on the differences in the spatial distribution matrix of the process parameters within a continuous time step, virtual and real interactive data that can reflect the dynamic changes in the actual operation of the pultrusion process equipment is formed.
[0036] Optionally, S3 specifically includes:
[0037] S31. Based on the numerical change characteristics of the spatial distribution matrix of process parameters in the virtual-real interaction data, define the state feature vector of the spatial nodes of the pultrusion border on the predetermined process path, and use the state feature vector as the input variable of the reinforcement learning model.
[0038] S32. Based on the product quality stability and production efficiency of the pultruded frame, extract the actual geometric dimension data of multiple continuous feature sections along the length direction of the pultruded frame, and compare them one by one with the design geometric dimension data of the corresponding feature sections to calculate the cross-sectional dimension deviation. Determine the production efficiency parameter by the ratio between the actual length value of the pultruded frame and the preset standard length value within a unit process cycle. Construct a composite reward function using the weighted cumulative sum of the absolute values of the dimension deviations of each feature section and the process cycle weighted integral of the production efficiency parameter.
[0039] S33. Based on the variation gradient magnitude and direction of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix, and pultrusion force parameter matrix, establish a parameter variation space with each characteristic section node on the process path as the reference. Divide the parameter variation space into a temperature control subspace, a pressure control subspace, a speed control subspace, and a pultrusion force control subspace. Discretize the parameter adjustment magnitude of each control subspace into multiple executable quantization levels to form a multidimensional discrete action vector with clear and spatially independent positions of each characteristic section node.
[0040] S34. Combine the state feature vector and action vector into a state-action pair, input the SoftActor-Critic algorithm of the reinforcement learning model, and generate an initial set of action probability distributions corresponding to each state-action pair through a random policy network.
[0041] S35. Sample each candidate action vector one by one from the initial set of action probability distributions, calculate the expected policy value of each state-action pair, and sort the state-action pairs by value difference.
[0042] S36. Select the action vector corresponding to the state-action pair with the highest expected strategy value after sorting, and use it as the control decision data for the composite material pultrusion process.
[0043] Optionally, the reinforcement learning model specifically includes a spatial segmentation state perception module, an action space dynamic programming module, a state-action multi-layer cross-mapping module, a multi-objective policy evaluation module, and an adaptive decision output module.
[0044] The spatial segmentation state perception module receives the spatial distribution matrix of process parameters, segments the continuous process path longitudinally along the pultrusion process path into multiple spatial perception segments, and generates an independent spatial perception state vector for each segment.
[0045] The motion space dynamic planning module maps multidimensional discrete motion vectors from temperature control subspace, pressure control subspace, speed control subspace, and pultrusion force control subspace to each spatial perception segment, forming segmented independent local motion spaces, and dynamically plans the motion value combination relationship of different spatial perception segments.
[0046] The state-action multi-layer cross-mapping module generates a multi-layer cross-mapping structure based on the state vector of each spatial perception segment and the corresponding local action space, by establishing cross-mapping paths between local state-action mapping relationships and state-action mapping relationships of adjacent spatial perception segments, and forms an overall state-action cross-mapping network that runs through the pultrusion process path.
[0047] The multi-objective strategy evaluation module, based on the state-action pairs in the overall state-action cross-mapping network, constructs a cross-sectional size deviation objective function and a production efficiency objective function, respectively, using the cumulative sum of the absolute values of the deviations in the cross-sectional dimensions of the pultruded frame and the process cycle weighted integral of the production efficiency parameter as input variables, and iteratively calculates the strategy value of each state-action pair through the objective functions.
[0048] The adaptive decision output module receives the policy value calculated by the multi-objective policy evaluation module, determines the optimal state-action pair according to the policy value ranking, and outputs the corresponding action vector as control decision data.
[0049] Optionally, S4 specifically includes:
[0050] S41. Based on the state-action pairs in the control decision data, define the policy network learning rate, state value network learning rate, Q-value network learning rate and policy entropy coefficient of the SoftActor-Critic algorithm as initial hyperparameters, and map the values of each hyperparameter to the spatial coordinates of the optimizable location within the artificial bee colony algorithm.
[0051] S42. Construct an initial bee colony with hyperparameter spatial location coordinates as the search object. Each spatial location coordinate corresponds to a set of explicit hyperparameter combinations. The sum of the policy network error, state value network error, and Q-value network error of the hyperparameter combination is used as the initial quality value of the spatial location.
[0052] S43. Perform multidimensional position perturbation on the spatial coordinates of each position in the initial bee colony, calculate the total error of the hyperparameter combination corresponding to the new position after perturbation in the reinforcement learning model, and use the difference between the quality value of the new position and the quality value of the original position as the evaluation basis for subsequent position selection.
[0053] S44. Based on the difference in quality values before and after the disturbance at each spatial location, define the quality improvement threshold and establish a quality improvement level matrix. Evaluate the quality improvement level of each spatial location in the initial bee colony one by one. Mark the spatial locations where the quality improvement level reaches the predetermined threshold as leader bee locations. Based on the quality improvement level and spatial location coordinate characteristics of the leader bee locations, construct a hierarchical leading region centered on the leader bee locations in the hyperparameter space. Mark the spatial locations where the quality improvement level does not reach the predetermined threshold as follower bee locations. Determine the leading relationship of each follower bee location based on the spatial distance between each follower bee location and the hierarchical leading region, and form a set of leader bee locations and a corresponding set of follower bee locations respectively.
[0054] S45. Construct a multi-dimensional spatial guidance path through the spatial relative relationship between the positions of each leader bee, and adjust the hyperparameter combination of the leader bee positions in sequence according to the guidance path to generate a new hyperparameter combination and update the quality value of each leader bee position.
[0055] S46. Based on the distance vector between the follower bee position and the corresponding leader bee position in multidimensional space, calculate the position coordinate difference in each dimension, define a distance-sensitive adaptive scaling factor, and dynamically adjust the movement step size and direction of the follower bee position in each dimension according to the magnitude and direction of the position coordinate difference. Multiply the distance-sensitive adaptive scaling factor and the position coordinate difference in each dimension to form the adjusted new position coordinates of the follower bee, calculate the hyperparameter combination and quality value corresponding to the new position coordinates, and update the follower bee position set.
[0056] S47. Sort the quality values of the positions of the leader bee and the follower bee, and determine the hyperparameter combination corresponding to the spatial position with the highest quality value as the optimized hyperparameter of the reinforcement learning model.
[0057] Optionally, S45 specifically includes:
[0058] S451. Taking each leader bee position in the leader bee position set as the center position, calculate the Euclidean distance between the center position and all other leader bee positions in the leader bee position set one by one to form a spatial relative distance matrix of leader bee positions.
[0059] S452. Based on the distance values of the relative distance matrix of the leader bee positions, sort the leader bee positions around each center position and define them as multiple hierarchical neighborhoods. The leader bee positions with the shortest distance are defined as the first layer neighborhood, and then the second layer and subsequent layers of neighborhoods are defined in ascending order of distance.
[0060] S453. Based on the spatial position difference between the center position and each leader bee position in the first layer neighborhood, determine the hyperparameter difference matrix between the hyperparameter vector of each center position and the hyperparameter vector of the first layer neighborhood.
[0061] S454. Using the direction and magnitude of each hyperparameter difference in the hyperparameter difference matrix, adjust the hyperparameter vector of each center position dimension by dimension with a preset neighborhood weight factor to form a new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment.
[0062] S455. Using the new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment as the initial vector, and according to the hyperparameter difference matrix between the hyperparameter vectors of the second layer and subsequent layers, the same neighborhood weight factor is adjusted dimension by dimension in turn, and the final hyperparameter combination of the center position after all layers of neighborhood collaborative adjustment is generated.
[0063] S456. Update the hyperparameter values of each leader bee position using the final hyperparameter combination of the center position, and recalculate the quality value of the updated leader bee position to complete the hyperparameter update of the leader bee position set.
[0064] Optionally, S5 specifically includes:
[0065] S51. Based on the real-time measured values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, calculate the absolute difference between the predicted parameter values and the real-time measured values in the corresponding virtual-real interaction data node by node to obtain the node-level parameter real-time difference matrix.
[0066] S52. Based on the distribution characteristics of all parameter difference data in the real-time difference matrix of node-level parameters, construct the spatial distribution offset vector of parameter differences;
[0067] S53. Standardize the values of the spatial distribution offset vectors one by one to form an anomaly sensitivity matrix based on the direction and magnitude of the spatial offset of the parameter difference.
[0068] S54. Define the spatial parameter offset sensitivity threshold. Based on the standardized parameter difference of each node in the anomaly sensitivity matrix, compare the threshold node by node to determine the node with obvious offset, and record the specific spatial location and process parameter type corresponding to the node with obvious offset.
[0069] S55. Based on the spatial location and process parameter type of the determined node, trace the parameter change sequence of the corresponding node in multiple consecutive historical process cycles, calculate the mutation rate of the parameter change sequence of the determined node, and perform a node-by-node mutation rate evaluation.
[0070] S56. Select nodes whose parameter change sequence mutation rate exceeds the predetermined abnormal index and their corresponding temperature, pressure, speed and pultrusion force parameters as abnormal process parameter data.
[0071] Optionally, S6 specifically includes:
[0072] S61. Based on the values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, define the feature vector of each matrix for each process cycle, and use the spatial position coordinates of the pultrusion border and the process timestamp as a joint index to map and store the feature vectors one by one.
[0073] S62. Based on the action vector sequence in the control decision data, define a decision vector associated with the adjustment amount of temperature, pressure, speed and pultrusion force parameters for each spatial location node, and establish the association between the decision vector and the real-time process feature vector of the corresponding spatial location node, and map and store them one by one to the process control decision history storage space.
[0074] S63. Based on the mutation rate of the node parameter change sequence in the abnormal process parameter data, define the abnormal feature vector for each spatial location node, associate the temperature, pressure, speed and pultrusion force parameters corresponding to the node, and map and store them into the abnormal process parameter data set according to the spatial location coordinates and the timestamp of the mutation occurrence.
[0075] S64. Based on the distribution characteristics of process feature vectors, decision vectors and anomaly feature vectors on spatial location nodes and timestamps, construct a multi-dimensional state-decision-anomaly association structure node by node, and generate a node-level data feature library based on the association structure.
[0076] S65. Using the multidimensional correlation structure of each node in the node-level data feature library as input, calculate the state-decision-abnormal change trend between different process cycles node by node to form a correlation change trend matrix of state, decision and abnormality.
[0077] S66. Based on the correlation trend matrix, quantify the correlation sensitivity of temperature, pressure, speed, and pultrusion force parameter changes on the deviation of pultruded frame cross-sectional dimensions and production efficiency node by node, determine and extract the node positions and corresponding process parameter combinations with parameter sensitivity higher than a preset threshold, and generate and output auxiliary decision data.
[0078] The beneficial effects of this invention are:
[0079] (1) The present invention uses a sensor array and a nonlinear Kalman filter algorithm to synchronously collect and filter temperature, pressure, speed and pultrusion force parameters through a data sensing and acquisition unit, forming a spatiotemporally aligned process state feature matrix, which effectively enhances the reliability and integrity of data acquisition.
[0080] (2) This invention constructs a three-dimensional spatial coordinate system and a discretized node model based on process data flow through a digital twin virtual simulation unit, and updates the node parameters in real time. This enables seamless mapping of virtual and real interactive data, significantly improving the synchronization and simulation accuracy of online simulation and feedback, and exhibiting better real-time performance under complex pultrusion process conditions.
[0081] (3) In the adaptive intelligent control unit, the present invention adopts the SoftActor-Critic reinforcement learning model with cross-sectional size deviation and production efficiency as composite reward functions, and combines the artificial bee colony hyperparameter optimization unit to perform real-time collaborative optimization of policy network, value network and entropy coefficient, which effectively shortens the model convergence time by 50% and improves the robustness and adaptability of the control strategy.
[0082] (4) This invention evaluates the real-time difference matrix and mutation rate at the spatial node level through an adaptive intelligent decision-making and anomaly early warning unit, and assists decision-making by combining a data feature library of multi-dimensional state-decision-anomaly association structure. This enables rapid identification and location of abnormal parameters in the pultrusion process, significantly improves the early warning response speed and decision accuracy, thereby enhancing the stability and reliability of the production process. Attached Figure Description
[0083] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0084] Figure 1 This is a system composition framework diagram of the reinforcement learning-based composite material pultruded frame data acquisition and management system proposed in this invention;
[0085] Figure 2 The flowchart of the digital twin virtual simulation and adaptive intelligent control of the composite material pultruded frame data acquisition and management system based on reinforcement learning proposed in this invention is shown.
[0086] Figure 3 This is a flowchart of the anomaly warning process of the reinforcement learning-based composite material pultruded frame data acquisition and management system proposed in this invention. Detailed Implementation
[0087] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0088] refer to Figures 1-3 The reinforcement learning-based composite material pultruded frame data acquisition and management system includes the following modules:
[0089] The data sensing and acquisition unit, consisting of a sensor array and a nonlinear Kalman filter algorithm, acquires temperature, pressure, speed, and pultrusion force parameters during the composite material pultrusion process and generates a process data stream.
[0090] The digital twin virtual simulation unit receives the process data stream, constructs a digital twin virtual model of the composite material pultrusion process, and outputs virtual-real interaction data.
[0091] The adaptive intelligent control unit, built on the SoftActor-Critic algorithm, uses the product quality stability and production efficiency of the extruded frame as a composite reward function to generate control decision data based on virtual and real interaction data.
[0092] The artificial bee colony hyperparameter optimization unit is used to optimize the hyperparameters of the SoftActor-Critic algorithm based on control decision data using the artificial bee colony algorithm, and output the optimized hyperparameters.
[0093] An adaptive intelligent decision-making and anomaly early warning unit is used to identify and output abnormal process parameter data during the pultrusion of composite materials;
[0094] The data management and decision support unit is used to receive and store the process data stream, control decision data, and abnormal process parameter data, generate a data feature library, and output decision support data.
[0095] By integrating a sensor array with a nonlinear Kalman filter algorithm, this invention achieves high-precision real-time acquisition of key parameters in the pultrusion process. Utilizing a digital twin virtual simulation unit, it establishes an online virtual-real mapping of the process. Based on the SoftActor-Critic algorithm, an adaptive intelligent control unit and an artificial bee colony hyperparameter optimization unit dynamically adjust control strategy parameters, achieving precise control of the pultrusion process. An adaptive intelligent decision-making and anomaly warning unit quickly identifies abnormal parameters by comparing real-time acquired data with simulation data. The data management and auxiliary decision-making unit integrates and stores multi-source process data, control decision data, and abnormal parameter data, providing comprehensive data support for subsequent production process optimization and decision-making.
[0096] In this embodiment, the modules are interconnected using the following method:
[0097] S1. Collect temperature, pressure, speed and pultrusion force parameters during the composite material pultrusion process through a sensor array, and process them through a nonlinear Kalman filter algorithm to generate a process data stream;
[0098] S2. Input the process data stream into the digital twin virtual simulation unit to construct a digital twin virtual model of the composite material pultrusion process and output virtual-real interaction data;
[0099] S3. Based on the virtual-real interaction data, a reinforcement learning model is established using the SoftActor-Critic algorithm, with the product quality stability and production efficiency of the extruded frame as a composite reward function, to generate control decision data.
[0100] S4. Based on the control decision data, optimize the reinforcement learning model using the artificial bee colony algorithm and output the optimized hyperparameters.
[0101] S5. Analyze the process data stream and virtual-real interaction data to identify and output abnormal process parameter data in the composite material pultrusion process.
[0102] S6. Receive and store the process data stream, control decision data and abnormal process parameter data, construct a data feature library, and output auxiliary decision data.
[0103] By synchronously acquiring and processing temperature, pressure, speed, and pultrusion force parameters through a sensor array and a nonlinear Kalman filter algorithm, a continuous and accurate process data stream can be generated. The digital twin virtual simulation unit maps the process data stream into online virtual-real interactive data, realizing real-time visualization of the process. The adaptive intelligent control unit, based on the SoftActor-Critic algorithm model and combined with hyperparameters optimized by artificial bee colony, generates and updates control decisions in real time, which helps stabilize the pultrusion process. The adaptive intelligent decision-making and anomaly early warning unit quickly identifies and outputs abnormal process parameters by comparing measured and simulated data. The data management and auxiliary decision-making unit integrates and stores process data, control decisions, and abnormal data, and generates a data feature library.
[0104] In this embodiment, S1 specifically includes:
[0105] S11. The temperature values of the resin melt and multiple spatial points inside the mold cavity during the composite material pultrusion process are measured in real time and in different regions using a temperature sensor array to form a temperature parameter matrix.
[0106] S12. The pressure values of multiple key sections preset at the mold inlet, outlet and inside the mold cavity during the composite material pultrusion process are measured in real time and in stages by a pressure sensor array to form a pressure parameter matrix.
[0107] S13. The operating speed of the composite material traction device and the instantaneous speed at multiple points on the traction trajectory are measured in real time through a speed sensor array to form a speed parameter matrix;
[0108] S14. The magnitude of the pultrusion force at multiple points during the pultrusion process of composite materials and its dynamic characteristics as the process progress are measured in real time by a pultrusion force sensor array, forming a pultrusion force parameter matrix.
[0109] S15. Synchronously align the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix according to the same timestamp, and align the spatial measurement positions of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix with the length direction of the pultrusion frame as the reference. Normalize and splice the values of each matrix after synchronization and alignment element by element in a unified spatiotemporal coordinate system to obtain a process state feature matrix with unified dimension and scale.
[0110] S16. Input the process state feature matrix into a nonlinear Kalman filter algorithm. Through dynamic iterative updating of the covariance of the state feature matrix, the process parameters are dynamically optimized and calibrated, and a process data stream is generated.
[0111] By using multi-point synchronous measurement and spatiotemporal alignment of temperature, pressure, speed and pultrusion force parameters matrix, this invention can accurately reflect the state characteristics of each measurement location during the composite material pultrusion process. Based on the nonlinear Kalman filter algorithm, it dynamically optimizes and calibrates the process state characteristic matrix of a unified dimension to generate a continuous and consistent process data stream, thereby improving data consistency and integrity.
[0112] In this embodiment, S2 specifically includes:
[0113] S21. Based on the real-time data of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, establish a three-dimensional spatial coordinate system corresponding to the actual composite material pultrusion equipment;
[0114] S22. Based on the structural characteristics of the pultrusion process frame of composite materials, determine the discretization nodes of the model and map each node to a specific position in the three-dimensional spatial coordinate system.
[0115] S23. Using the discretized nodes of the model as the reference position, the real-time data in the temperature parameter matrix, pressure parameter matrix, velocity parameter matrix and pultrusion force parameter matrix are mapped to the corresponding nodes one by one according to the spatial distance weighted interpolation method. The material properties and process characteristics at the discretized nodes of the model are used as constraints. Spatial difference correction and parameter matching between nodes are performed on the mapped real-time data to form a set of node parameter distributions and construct a digital twin virtual model of composite material pultrusion process.
[0116] S24. Based on the temperature, pressure, speed and pultrusion force parameters of the nodes in the digital twin virtual model, calculate the spatial distribution characteristics of the parameters node by node and update them continuously over time to obtain the spatial distribution matrix of the process parameters.
[0117] S25. Based on the differences in the spatial distribution matrix of the process parameters within a continuous time step, virtual and real interactive data that can reflect the dynamic changes in the actual operation of the pultrusion process equipment is formed.
[0118] By constructing a three-dimensional spatial coordinate system corresponding to the actual pultrusion equipment based on real-time data of temperature, pressure, speed, and pultrusion force parameter matrices, and mapping process data to discretized structural nodes, this invention achieves a precise correspondence between the physical process state and the digital twin model. By correcting and matching the mapped data through spatial distance weighted interpolation and material feature constraints, this invention improves the accuracy of node parameter distribution. Through continuous temporal updates of node parameters in the digital twin model, this invention obtains high-fidelity virtual-real interactive data for online simulation and intelligent control.
[0119] In this embodiment, S3 specifically includes:
[0120] S31. Based on the numerical change characteristics of the spatial distribution matrix of process parameters in the virtual-real interaction data, define the state feature vector of the spatial nodes of the pultrusion border on the predetermined process path, and use the state feature vector as the input variable of the reinforcement learning model.
[0121] S32. Based on the product quality stability and production efficiency of the pultruded frame, extract the actual geometric dimension data of multiple continuous feature sections along the length direction of the pultruded frame, and compare them one by one with the design geometric dimension data of the corresponding feature sections to calculate the cross-sectional dimension deviation. Determine the production efficiency parameter by the ratio between the actual length value of the pultruded frame and the preset standard length value within a unit process cycle. Construct a composite reward function using the weighted cumulative sum of the absolute values of the dimension deviations of each feature section and the process cycle weighted integral of the production efficiency parameter.
[0122] S33. Based on the variation gradient magnitude and direction of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix, and pultrusion force parameter matrix, establish a parameter variation space with each characteristic section node on the process path as the reference. Divide the parameter variation space into a temperature control subspace, a pressure control subspace, a speed control subspace, and a pultrusion force control subspace. Discretize the parameter adjustment magnitude of each control subspace into multiple executable quantization levels to form a multidimensional discrete action vector with clear and spatially independent positions of each characteristic section node.
[0123] S34. Combine the state feature vector and action vector into a state-action pair, input the SoftActor-Critic algorithm of the reinforcement learning model, and generate an initial set of action probability distributions corresponding to each state-action pair through a random policy network.
[0124] S35. Sample each candidate action vector one by one from the initial set of action probability distributions, calculate the expected policy value of each state-action pair, and sort the state-action pairs by value difference.
[0125] S36. Select the action vector corresponding to the state-action pair with the highest expected strategy value after sorting, and use it as the control decision data for the composite material pultrusion process.
[0126] By extracting state feature vectors of spatial nodes in the pultrusion boundary based on virtual-real interaction data, and constructing a composite reward function by combining the differences between actual and designed geometric dimensions of multiple continuous cross-sections along the length direction, this invention divides discretized action vectors in a four-dimensional control space of temperature, pressure, speed and pultrusion force, forming an initial set of state-action pair probability distributions, and selecting the optimal action vector by ranking the policy value, effectively improving the convergence speed of the reinforcement learning model and the accuracy of control decisions.
[0127] In this embodiment, the reinforcement learning model specifically includes a spatial segmentation state perception module, an action space dynamic programming module, a state-action multi-layer cross-mapping module, a multi-objective policy evaluation module, and an adaptive decision output module.
[0128] The spatial segmentation state perception module receives the spatial distribution matrix of process parameters, segments the continuous process path longitudinally along the pultrusion process path into multiple spatial perception segments, and generates an independent spatial perception state vector for each segment.
[0129] The motion space dynamic planning module maps multidimensional discrete motion vectors from temperature control subspace, pressure control subspace, speed control subspace, and pultrusion force control subspace to each spatial perception segment, forming segmented independent local motion spaces, and dynamically plans the motion value combination relationship of different spatial perception segments.
[0130] The state-action multi-layer cross-mapping module generates a multi-layer cross-mapping structure based on the state vector of each spatial perception segment and the corresponding local action space, by establishing cross-mapping paths between local state-action mapping relationships and state-action mapping relationships of adjacent spatial perception segments, and forms an overall state-action cross-mapping network that runs through the pultrusion process path.
[0131] The multi-objective strategy evaluation module, based on the state-action pairs in the overall state-action cross-mapping network, constructs a cross-sectional size deviation objective function and a production efficiency objective function, respectively, using the cumulative sum of the absolute values of the deviations in the cross-sectional dimensions of the pultruded frame and the process cycle weighted integral of the production efficiency parameter as input variables, and iteratively calculates the strategy value of each state-action pair through the objective functions.
[0132] The adaptive decision output module receives the policy value calculated by the multi-objective policy evaluation module, determines the optimal state-action pair according to the policy value ranking, and outputs the corresponding action vector as control decision data.
[0133] By dividing the pultrusion process path into spatially perceptible segments and generating corresponding state vectors, independent perception of multiple process states is achieved. Through dynamic programming of the action space and multi-layer state-action cross-mapping, a cross-mapping network covering local and global mapping relationships is established. By evaluating multi-objective strategies, the strategy value of cross-sectional deviation and production efficiency is calculated in parallel, and the optimal action vector is output according to the strategy value ranking, thus achieving simultaneous trade-offs for multi-dimensional control objectives. Therefore, this invention maintains the consistency and coherence of state-action mapping between different segments of the process path, effectively improving the accuracy and adaptability of the reinforcement learning model in control decisions under complex pultrusion process conditions.
[0134] In this embodiment, S4 specifically includes:
[0135] S41. Based on the state-action pairs in the control decision data, define the policy network learning rate, state value network learning rate, Q-value network learning rate and policy entropy coefficient of the SoftActor-Critic algorithm as initial hyperparameters, and map the values of each hyperparameter to the spatial coordinates of the optimizable location within the artificial bee colony algorithm.
[0136] S42. Construct an initial bee colony with hyperparameter spatial location coordinates as the search object. Each spatial location coordinate corresponds to a set of explicit hyperparameter combinations. The sum of the policy network error, state value network error, and Q-value network error of the hyperparameter combination is used as the initial quality value of the spatial location.
[0137] S43. Perform multidimensional position perturbation on the spatial coordinates of each position in the initial bee colony, calculate the total error of the hyperparameter combination corresponding to the new position after perturbation in the reinforcement learning model, and use the difference between the quality value of the new position and the quality value of the original position as the evaluation basis for subsequent position selection.
[0138] S44. Based on the difference in quality values before and after the disturbance at each spatial location, define the quality improvement threshold and establish a quality improvement level matrix. Evaluate the quality improvement level of each spatial location in the initial bee colony one by one. Mark the spatial locations where the quality improvement level reaches the predetermined threshold as leader bee locations. Based on the quality improvement level and spatial location coordinate characteristics of the leader bee locations, construct a hierarchical leading region centered on the leader bee locations in the hyperparameter space. Mark the spatial locations where the quality improvement level does not reach the predetermined threshold as follower bee locations. Determine the leading relationship of each follower bee location based on the spatial distance between each follower bee location and the hierarchical leading region, and form a set of leader bee locations and a corresponding set of follower bee locations respectively.
[0139] S45. Construct a multi-dimensional spatial guidance path through the spatial relative relationship between the positions of each leader bee, and adjust the hyperparameter combination of the leader bee positions in sequence according to the guidance path to generate a new hyperparameter combination and update the quality value of each leader bee position.
[0140] S46. Based on the distance vector between the follower bee position and the corresponding leader bee position in multidimensional space, calculate the position coordinate difference in each dimension, define a distance-sensitive adaptive scaling factor, and dynamically adjust the movement step size and direction of the follower bee position in each dimension according to the magnitude and direction of the position coordinate difference. Multiply the distance-sensitive adaptive scaling factor and the position coordinate difference in each dimension to form the adjusted new position coordinates of the follower bee, calculate the hyperparameter combination and quality value corresponding to the new position coordinates, and update the follower bee position set.
[0141] S47. Sort the quality values of the positions of the leader bee and the follower bee, and determine the hyperparameter combination corresponding to the spatial position with the highest quality value as the optimized hyperparameter of the reinforcement learning model.
[0142] By mapping the learning rates of the policy network, state value network, and Q-value network to spatial coordinates, this invention achieves a spatialized expression of hyperparameters. Through multidimensional perturbation of the initial colony position and error-weighted evaluation to select leader and follower bees, this invention constructs a hierarchical leader region and corresponding adaptive collaborative update path. By using a distance-sensitive adaptive scaling factor to finely adjust the follower bee position, this invention effectively balances global exploration and local search, shortens the hyperparameter optimization time, and improves the convergence efficiency and control accuracy of the reinforcement learning model in pultrusion process control.
[0143] In this embodiment, S45 specifically includes:
[0144] S451. Taking each leader bee position in the leader bee position set as the center position, calculate the Euclidean distance between the center position and all other leader bee positions in the leader bee position set one by one to form a spatial relative distance matrix of leader bee positions.
[0145] S452. Based on the distance values of the relative distance matrix of the leader bee positions, sort the leader bee positions around each center position and define them as multiple hierarchical neighborhoods. The leader bee positions with the shortest distance are defined as the first layer neighborhood, and then the second layer and subsequent layers of neighborhoods are defined in ascending order of distance.
[0146] S453. Based on the spatial position difference between the center position and each leader bee position in the first layer neighborhood, determine the hyperparameter difference matrix between the hyperparameter vector of each center position and the hyperparameter vector of the first layer neighborhood.
[0147] S454. Using the direction and magnitude of each hyperparameter difference in the hyperparameter difference matrix, adjust the hyperparameter vector of each center position dimension by dimension with a preset neighborhood weight factor to form a new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment.
[0148] S455. Using the new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment as the initial vector, and according to the hyperparameter difference matrix between the hyperparameter vectors of the second layer and subsequent layers, the same neighborhood weight factor is adjusted dimension by dimension in turn, and the final hyperparameter combination of the center position after all layers of neighborhood collaborative adjustment is generated.
[0149] S456. Update the hyperparameter values of each leader bee position using the final hyperparameter combination of the center position, and recalculate the quality value of the updated leader bee position to complete the hyperparameter update of the leader bee position set.
[0150] By dividing the set of leading bee locations into hierarchical neighborhoods based on spatial relative distance and constructing neighborhood leading paths, this invention achieves hierarchical collaborative updating of hyperparameter locations. By adjusting the position coordinates dimension-by-dimensionally based on the hyperparameter difference direction and hierarchical weight factor in each neighborhood and updating the quality value in real time, this invention can finely optimize hyperparameter combinations and maintain the continuity of updates. This hierarchical leading-following mechanism effectively balances the intensity of global search and local search, shortens the hyperparameter optimization time, and improves the convergence efficiency and parameter stability of reinforcement learning models in the control of composite material pultrusion processes.
[0151] In this embodiment, S5 specifically includes:
[0152] S51. Based on the real-time measured values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, calculate the absolute difference between the predicted parameter values and the real-time measured values in the corresponding virtual-real interaction data node by node to obtain the node-level parameter real-time difference matrix.
[0153] S52. Based on the distribution characteristics of all parameter difference data in the real-time difference matrix of node-level parameters, construct the spatial distribution offset vector of parameter differences;
[0154] S53. Standardize the values of the spatial distribution offset vectors one by one to form an anomaly sensitivity matrix based on the direction and magnitude of the spatial offset of the parameter difference.
[0155] S54. Define the spatial parameter offset sensitivity threshold. Based on the standardized parameter difference of each node in the anomaly sensitivity matrix, compare the threshold node by node to determine the node with obvious offset, and record the specific spatial location and process parameter type corresponding to the node with obvious offset.
[0156] S55. Based on the spatial location and process parameter type of the determined node, trace the parameter change sequence of the corresponding node in multiple consecutive historical process cycles, calculate the mutation rate of the parameter change sequence of the determined node, and perform a node-by-node mutation rate evaluation.
[0157] S56. Select nodes whose parameter change sequence mutation rate exceeds the predetermined abnormal index and their corresponding temperature, pressure, speed and pultrusion force parameters as abnormal process parameter data.
[0158] By calculating the absolute value of the difference between the predicted value and the real-time measurement value of the virtual-real interactive data node by node and constructing a spatial distribution offset vector, this invention can refine the anomaly sensitivity matrix and identify obvious offset nodes node by node by node by combining the standardized offset value and the threshold. Then, based on the mutation rate calculated by the historical process cycle parameter change sequence, this invention realizes the accurate location of abnormal parameter nodes and multi-cycle mutation analysis, thereby improving the accuracy and reliability of abnormal process parameter detection.
[0159] In this embodiment, S6 specifically includes:
[0160] S61. Based on the values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, define the feature vector of each matrix for each process cycle, and use the spatial position coordinates of the pultrusion border and the process timestamp as a joint index to map and store the feature vectors one by one.
[0161] S62. Based on the action vector sequence in the control decision data, define a decision vector associated with the adjustment amount of temperature, pressure, speed and pultrusion force parameters for each spatial location node, and establish the association between the decision vector and the real-time process feature vector of the corresponding spatial location node, and map and store them one by one to the process control decision history storage space.
[0162] S63. Based on the mutation rate of the node parameter change sequence in the abnormal process parameter data, define the abnormal feature vector for each spatial location node, associate the temperature, pressure, speed and pultrusion force parameters corresponding to the node, and map and store them into the abnormal process parameter data set according to the spatial location coordinates and the timestamp of the mutation occurrence.
[0163] S64. Based on the distribution characteristics of process feature vectors, decision vectors and anomaly feature vectors on spatial location nodes and timestamps, construct a multi-dimensional state-decision-anomaly association structure node by node, and generate a node-level data feature library based on the association structure.
[0164] S65. Using the multidimensional correlation structure of each node in the node-level data feature library as input, calculate the state-decision-abnormal change trend between different process cycles node by node to form a correlation change trend matrix of state, decision and abnormality.
[0165] S66. Based on the correlation trend matrix, quantify the correlation sensitivity of temperature, pressure, speed, and pultrusion force parameter changes on the deviation of pultruded frame cross-sectional dimensions and production efficiency node by node, determine and extract the node positions and corresponding process parameter combinations with parameter sensitivity higher than a preset threshold, and generate and output auxiliary decision data.
[0166] By mapping and generating node-level feature vectors for each process cycle based on temperature, pressure, speed, and pultrusion force parameter matrices, and combining the spatial coordinates and timestamps of control decision vectors and anomaly feature vectors, a multi-dimensional state-decision-anomaly correlation structure and node-level data feature library are constructed. Based on the correlation structure, the state-decision-anomaly change trend matrix and parameter sensitivity index for different process cycles are calculated. This invention can accurately identify the parameter combinations that have the greatest impact on pultruded frame cross-sectional deviation and production efficiency, and output auxiliary decision data accordingly, thereby improving the pertinence and response efficiency of production process decisions.
[0167] Example 1:
[0168] To verify the feasibility of this invention in practice, it was applied to a composite material pultrusion production line of a certain enterprise, specifically to the batch production process control and data management of a certain specification of composite material pultruded frame products. Traditionally, this production line relies on manual experience to set process parameters, adjusting temperature, pressure, speed, and pultrusion force control parameters through offline manual analysis and experimentation. This typically requires on-site engineers to repeatedly adjust and observe production conditions to determine the optimal parameter combination. Existing methods generally suffer from low accuracy in real-time process data acquisition, low levels of automation in the production process, lag in process parameter control response, and insufficient product quality stability. This results in poor dimensional consistency of batch-produced composite material pultruded frame products, with a product qualification rate of approximately 86.2% and an average production efficiency of only 3.8 meters per hour, significantly restricting the enterprise's production scale and product quality improvement.
[0169] In this embodiment, a composite material pultrusion frame data acquisition and management system based on the Soft Actor-Critic algorithm combined with artificial bee swarm algorithm, proposed in this invention, is deployed on the production line. First, a data sensing and acquisition unit uses a sensor array to synchronously measure the temperature of the resin melt and multiple spatial points inside the mold cavity, the mold inlet and outlet pressures, the traction device speed, and the pultrusion force in real time, divided into regions. Then, a nonlinear Kalman filter algorithm is used for data filtering to generate a refined process data stream in real time. Next, the process data stream is input into a digital twin virtual simulation unit to establish a digital twin virtual model corresponding to the actual composite material pultrusion equipment, outputting virtual-real interactive data. Finally, a reinforcement learning model is used to... A composite reward function is constructed using the weighted cumulative sum of absolute values of dimensional deviations and the weighted integral of the production efficiency parameter and process cycle, generating control decision data in real time. Then, the hyperparameters of the reinforcement learning model—including the policy network learning rate, state value network learning rate, Q-value network learning rate, and policy entropy coefficient—are optimized using the artificial bee colony algorithm to achieve real-time hyperparameter optimization. Next, the system analyzes the difference between virtual and real interaction data and process data streams in real time, assesses the mutation rate of node parameters, and identifies and outputs abnormal process parameter data in real time. Finally, the process data stream, control decision data, and abnormal process parameter data are stored according to spatial node location and process timestamps to construct a node-level data feature library, outputting auxiliary decision data.
[0170] After implementing the above scheme, 15 consecutive days of production test data showed that the dimensional consistency of the production batches was significantly improved, the product qualification rate increased from 86.2% to 97.8%, and the production efficiency increased from 3.8 meters per hour to 5.6 meters per hour. The real-time response speed of process parameter control during production was reduced from an average of 45 seconds using traditional manual methods to less than 3 seconds, and the accuracy rate of process anomaly identification reached over 96%, greatly reducing the batch scrap rate. The table below shows a comparison of the quality and production efficiency of typical batches of composite material pultruded frame products before and after implementation:
[0171] Table 1 Comparison of Product Quality and Production Efficiency of Pultruded Frame Composite Materials
[0172]
[0173] As shown in Table 1 above, comparing the product quality and production efficiency of composite pultruded frame products, it is evident that the system of this invention has achieved a high level of control precision and response performance in several key production indicators. Taking product dimensional consistency as an example, the dimensional error of each batch after implementation is controlled within ±0.1mm, with the dimensional error of batch -02 after implementation being only ±0.08mm, compared to ±0.28mm before implementation, representing a 71.4% reduction in dimensional error and effectively improving the stability and consistency of product quality. Regarding product qualification rate, the qualification rate of all batches after implementation exceeds 97%, an average increase of approximately 11 percentage points compared to batches before implementation, demonstrating the reliability of the system in accurate data acquisition and real-time intelligent decision-making.
[0174] In terms of production efficiency, the system of this invention significantly improves the automation control response speed and production continuity of the process. The production efficiency of each batch after implementation exceeds 5.5 meters per hour, which is about 47% higher than the average of 3.8 meters per hour before implementation. Among them, the production efficiency of batch -02 after implementation reaches 5.7 meters per hour, indicating that the real-time collaborative optimization strategy of reinforcement learning model and artificial bee swarm algorithm effectively improves the accuracy of process parameter control and response time.
[0175] Furthermore, the system of this invention demonstrates a significant advantage in terms of response speed for identifying abnormal process parameters. After implementation, the response time for abnormalities in each batch is controlled within 3 seconds. For example, the response time for the -02 batch after implementation is only 2.3 seconds, which is more than 90% faster than the average response time of 42 seconds of the traditional method before implementation. This verifies the efficiency of the system in analyzing the changing trends of abnormal parameters in real time through the data feature library, quickly locating the abnormal parameters, and making auxiliary decisions.
[0176] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A data acquisition and management system for pultruded frames of composite materials based on reinforcement learning, characterized in that, Includes the following modules: The data sensing and acquisition unit, consisting of a sensor array and a nonlinear Kalman filter algorithm, acquires temperature, pressure, speed, and pultrusion force parameters during the composite material pultrusion process and generates a process data stream. The digital twin virtual simulation unit receives the process data stream, constructs a digital twin virtual model of the composite material pultrusion process, and outputs virtual-real interaction data. The adaptive intelligent control unit, built on the SoftActor-Critic algorithm, uses the product quality stability and production efficiency of the extruded frame as a composite reward function to generate control decision data based on virtual and real interaction data. The artificial bee colony hyperparameter optimization unit is used to optimize the hyperparameters of the SoftActor-Critic algorithm based on control decision data using the artificial bee colony algorithm, and output the optimized hyperparameters. An adaptive intelligent decision-making and anomaly early warning unit is used to identify and output abnormal process parameter data during the pultrusion of composite materials; The data management and decision support unit is used to receive and store the process data stream, control decision data, and abnormal process parameter data, generate a data feature library, and output decision support data.
2. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 1, characterized in that, The modules are connected in the following way: S1. Collect temperature, pressure, speed and pultrusion force parameters during the composite material pultrusion process through a sensor array, and process them through a nonlinear Kalman filter algorithm to generate a process data stream; S2. Input the process data stream into the digital twin virtual simulation unit to construct a digital twin virtual model of the composite material pultrusion process and output virtual-real interaction data; S3. Based on the virtual-real interaction data, a reinforcement learning model is established using the SoftActor-Critic algorithm, with the product quality stability and production efficiency of the extruded frame as a composite reward function, to generate control decision data. S4. Based on the control decision data, optimize the reinforcement learning model using the artificial bee colony algorithm and output the optimized hyperparameters. S5. Analyze the process data stream and virtual-real interaction data to identify and output abnormal process parameter data in the composite material pultrusion process. S6. Receive and store the process data stream, control decision data and abnormal process parameter data, construct a data feature library, and output auxiliary decision data.
3. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S1 specifically includes: S11. The temperature values of the resin melt and multiple spatial points inside the mold cavity during the composite material pultrusion process are measured in real time and in different regions using a temperature sensor array to form a temperature parameter matrix. S12. The pressure values of multiple key sections preset at the mold inlet, outlet and inside the mold cavity during the composite material pultrusion process are measured in real time and in stages by a pressure sensor array to form a pressure parameter matrix. S13. The operating speed of the composite material traction device and the instantaneous speed at multiple points on the traction trajectory are measured in real time by a speed sensor array to form a speed parameter matrix; S14. The magnitude of the pultrusion force at multiple points during the pultrusion process of composite materials and its dynamic characteristics as the process progress are measured in real time by a pultrusion force sensor array, forming a pultrusion force parameter matrix. S15. Synchronously align the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix according to the same timestamp, and align the spatial measurement positions of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix with the length direction of the pultrusion frame as the reference. Normalize and splice the values of each matrix after synchronization and alignment element by element in a unified spatiotemporal coordinate system to obtain a process state feature matrix with unified dimension and scale. S16. Input the process state feature matrix into a nonlinear Kalman filter algorithm. Through dynamic iterative updating of the covariance of the state feature matrix, the process parameters are dynamically optimized and calibrated, and a process data stream is generated.
4. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S2 specifically includes: S21. Based on the real-time data of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, establish a three-dimensional spatial coordinate system corresponding to the actual composite material pultrusion equipment; S22. Based on the structural characteristics of the pultrusion process frame of composite materials, determine the discretization nodes of the model and map each node to a specific position in the three-dimensional spatial coordinate system. S23. Using the discretized nodes of the model as the reference position, the real-time data in the temperature parameter matrix, pressure parameter matrix, velocity parameter matrix and pultrusion force parameter matrix are mapped to the corresponding nodes one by one according to the spatial distance weighted interpolation method. The material properties and process characteristics at the discretized nodes of the model are used as constraints. Spatial difference correction and parameter matching between nodes are performed on the mapped real-time data to form a set of node parameter distributions and construct a digital twin virtual model of composite material pultrusion process. S24. Based on the temperature, pressure, speed and pultrusion force parameters of the nodes in the digital twin virtual model, calculate the spatial distribution characteristics of the parameters node by node and update them continuously over time to obtain the spatial distribution matrix of the process parameters. S25. Based on the differences in the spatial distribution matrix of the process parameters within a continuous time step, virtual and real interactive data that can reflect the dynamic changes in the actual operation of the pultrusion process equipment is formed.
5. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S3 specifically includes: S31. Based on the numerical change characteristics of the spatial distribution matrix of process parameters in the virtual-real interaction data, define the state feature vector of the spatial nodes of the pultrusion border on the predetermined process path, and use the state feature vector as the input variable of the reinforcement learning model. S32. Based on the product quality stability and production efficiency of the pultruded frame, extract the actual geometric dimension data of multiple continuous feature sections along the length direction of the pultruded frame, and compare them one by one with the design geometric dimension data of the corresponding feature sections to calculate the cross-sectional dimension deviation. Determine the production efficiency parameter by the ratio between the actual length value of the pultruded frame and the preset standard length value within a unit process cycle. Construct a composite reward function using the weighted cumulative sum of the absolute values of the dimension deviations of each feature section and the process cycle weighted integral of the production efficiency parameter. S33. Based on the variation gradient magnitude and direction of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix, and pultrusion force parameter matrix, establish a parameter variation space with each characteristic section node on the process path as the reference. Divide the parameter variation space into a temperature control subspace, a pressure control subspace, a speed control subspace, and a pultrusion force control subspace. Discretize the parameter adjustment magnitude of each control subspace into multiple executable quantization levels to form a multidimensional discrete action vector with clear and spatially independent positions of each characteristic section node. S34. Combine the state feature vector and action vector into a state-action pair, input the SoftActor-Critic algorithm of the reinforcement learning model, and generate an initial set of action probability distributions corresponding to each state-action pair through a random policy network. S35. Sample each candidate action vector one by one from the initial set of action probability distributions, calculate the expected policy value of each state-action pair, and sort the state-action pairs by value difference. S36. Select the action vector corresponding to the state-action pair with the highest expected strategy value after sorting, and use it as the control decision data for the composite material pultrusion process.
6. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 5, characterized in that, The reinforcement learning model specifically includes a spatial segmentation state awareness module, an action space dynamic programming module, a state-action multi-layer cross-mapping module, a multi-objective policy evaluation module, and an adaptive decision output module. The spatial segmentation state perception module receives the spatial distribution matrix of process parameters, segments the continuous process path longitudinally along the pultrusion process path into multiple spatial perception segments, and generates an independent spatial perception state vector for each segment. The motion space dynamic planning module maps multidimensional discrete motion vectors from temperature control subspace, pressure control subspace, speed control subspace, and pultrusion force control subspace to each spatial perception segment, forming segmented independent local motion spaces, and dynamically plans the motion value combination relationship of different spatial perception segments. The state-action multi-layer cross-mapping module generates a multi-layer cross-mapping structure based on the state vector of each spatial perception segment and the corresponding local action space, by establishing cross-mapping paths between local state-action mapping relationships and state-action mapping relationships of adjacent spatial perception segments, and forms an overall state-action cross-mapping network that runs through the pultrusion process path. The multi-objective strategy evaluation module, based on the state-action pairs in the overall state-action cross-mapping network, constructs a cross-sectional size deviation objective function and a production efficiency objective function, respectively, using the cumulative sum of the absolute values of the deviations in the cross-sectional dimensions of the pultruded frame and the process cycle weighted integral of the production efficiency parameter as input variables, and iteratively calculates the strategy value of each state-action pair through the objective functions. The adaptive decision output module receives the policy value calculated by the multi-objective policy evaluation module, determines the optimal state-action pair according to the policy value ranking, and outputs the corresponding action vector as control decision data.
7. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S4 specifically includes: S41. Based on the state-action pairs in the control decision data, define the policy network learning rate, state value network learning rate, Q-value network learning rate and policy entropy coefficient of the SoftActor-Critic algorithm as initial hyperparameters, and map the values of each hyperparameter to the spatial coordinates of the optimizable location within the artificial bee colony algorithm. S42. Construct an initial bee colony with hyperparameter spatial location coordinates as the search object. Each spatial location coordinate corresponds to a set of explicit hyperparameter combinations. The sum of the policy network error, state value network error, and Q-value network error of the hyperparameter combination is used as the initial quality value of the spatial location. S43. Perform multidimensional position perturbation on the spatial coordinates of each position in the initial bee colony, calculate the total error of the hyperparameter combination corresponding to the new position after perturbation in the reinforcement learning model, and use the difference between the quality value of the new position and the quality value of the original position as the evaluation basis for subsequent position selection. S44. Based on the difference in quality values before and after the disturbance at each spatial location, define the quality improvement threshold and establish a quality improvement level matrix. Evaluate the quality improvement level of each spatial location in the initial bee colony one by one. Mark the spatial locations where the quality improvement level reaches the predetermined threshold as leader bee locations. Based on the quality improvement level and spatial location coordinate characteristics of the leader bee locations, construct a hierarchical leading region centered on the leader bee locations in the hyperparameter space. Mark the spatial locations where the quality improvement level does not reach the predetermined threshold as follower bee locations. Determine the leading relationship of each follower bee location based on the spatial distance between each follower bee location and the hierarchical leading region, and form a set of leader bee locations and a corresponding set of follower bee locations respectively. S45. Construct a multi-dimensional spatial guidance path through the spatial relative relationship between the positions of each leader bee, and adjust the hyperparameter combination of the leader bee positions in sequence according to the guidance path to generate a new hyperparameter combination and update the quality value of each leader bee position. S46. Based on the distance vector between the follower bee position and the corresponding leader bee position in multidimensional space, calculate the position coordinate difference in each dimension, define a distance-sensitive adaptive scaling factor, and dynamically adjust the movement step size and direction of the follower bee position in each dimension according to the magnitude and direction of the position coordinate difference. Multiply the distance-sensitive adaptive scaling factor and the position coordinate difference in each dimension to form the adjusted new position coordinates of the follower bee, calculate the hyperparameter combination and quality value corresponding to the new position coordinates, and update the follower bee position set. S47. Sort the quality values of the positions of the leader bee and the follower bee, and determine the hyperparameter combination corresponding to the spatial position with the highest quality value as the optimized hyperparameter of the reinforcement learning model.
8. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 7, characterized in that, Specifically, S45 includes: S451. Taking each leader bee position in the leader bee position set as the center position, calculate the Euclidean distance between the center position and all other leader bee positions in the leader bee position set one by one to form a spatial relative distance matrix of leader bee positions. S452. Based on the distance values of the relative distance matrix of the leader bee positions, sort the leader bee positions around each center position and define them as multiple hierarchical neighborhoods. The leader bee positions with the shortest distance are defined as the first layer neighborhood, and then the second layer and subsequent layers of neighborhoods are defined in ascending order of distance. S453. Based on the spatial position difference between the center position and each leader bee position in the first layer neighborhood, determine the hyperparameter difference matrix between the hyperparameter vector of each center position and the hyperparameter vector of the first layer neighborhood. S454. Using the direction and magnitude of each hyperparameter difference in the hyperparameter difference matrix, adjust the hyperparameter vector of each center position dimension by dimension with a preset neighborhood weight factor to form a new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment. S455. Using the new hyperparameter vector of the center position after the first layer of neighborhood collaborative adjustment as the initial vector, and according to the hyperparameter difference matrix between the hyperparameter vectors of the second layer and subsequent layers, the same neighborhood weight factor is adjusted dimension by dimension in turn, and the final hyperparameter combination of the center position after all layers of neighborhood collaborative adjustment is generated. S456. Update the hyperparameter values of each leader bee position using the final hyperparameter combination of the center position, and recalculate the quality value of the updated leader bee position to complete the hyperparameter update of the leader bee position set.
9. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S5 specifically includes: S51. Based on the real-time measured values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, calculate the absolute difference between the predicted parameter values and the real-time measured values in the corresponding virtual-real interaction data node by node to obtain the node-level parameter real-time difference matrix. S52. Based on the distribution characteristics of all parameter difference data in the real-time difference matrix of node-level parameters, construct the spatial distribution offset vector of parameter differences; S53. Standardize the values of the spatial distribution offset vectors one by one to form an anomaly sensitivity matrix based on the direction and magnitude of the spatial offset of the parameter difference. S54. Define the spatial parameter offset sensitivity threshold. Based on the standardized parameter difference of each node in the anomaly sensitivity matrix, compare the threshold node by node to determine the node with obvious offset, and record the specific spatial location and process parameter type corresponding to the node with obvious offset. S55. Based on the spatial location and process parameter type of the determined node, trace the parameter change sequence of the corresponding node in multiple consecutive historical process cycles, calculate the mutation rate of the parameter change sequence of the determined node, and perform a node-by-node mutation rate evaluation. S56. Select nodes whose parameter change sequence mutation rate exceeds the predetermined abnormal index and their corresponding temperature, pressure, speed and pultrusion force parameters as abnormal process parameter data.
10. The data acquisition and management system for composite material pultruded borders based on reinforcement learning according to claim 2, characterized in that, S6 specifically includes: S61. Based on the values of the temperature parameter matrix, pressure parameter matrix, speed parameter matrix and pultrusion force parameter matrix in the process data stream, define the feature vector of each matrix for each process cycle, and use the spatial position coordinates of the pultrusion border and the process timestamp as a joint index to map and store the feature vectors one by one. S62. Based on the action vector sequence in the control decision data, define a decision vector associated with the adjustment amount of temperature, pressure, speed and pultrusion force parameters for each spatial location node, and establish the association between the decision vector and the real-time process feature vector of the corresponding spatial location node, and map and store them one by one to the process control decision history storage space. S63. Based on the mutation rate of the node parameter change sequence in the abnormal process parameter data, define the abnormal feature vector for each spatial location node, associate the temperature, pressure, speed and pultrusion force parameters corresponding to the node, and map and store them into the abnormal process parameter data set according to the spatial location coordinates and the timestamp of the mutation occurrence. S64. Based on the distribution characteristics of process feature vectors, decision vectors and anomaly feature vectors on spatial location nodes and timestamps, construct a multi-dimensional state-decision-anomaly association structure node by node, and generate a node-level data feature library based on the association structure. S65. Using the multidimensional correlation structure of each node in the node-level data feature library as input, calculate the state-decision-abnormal change trend between different process cycles node by node to form a correlation change trend matrix of state, decision and abnormality. S66. Based on the correlation trend matrix, quantify the correlation sensitivity of temperature, pressure, speed, and pultrusion force parameter changes on the deviation of pultruded frame cross-sectional dimensions and production efficiency node by node, determine and extract the node positions and corresponding process parameter combinations with parameter sensitivity higher than a preset threshold, and generate and output auxiliary decision data.