Large-scale bulk material equipment motion control optimization method based on deep learning

Through the cross-modal attention LSTM-PPO decision-making network, the multi-source information fusion problem of large bulk equipment in complex environments is solved, high-precision control optimization is achieved, the equipment's adaptability and operation stability are improved, and the control deviation and poor adaptability exist in traditional methods are solved.

CN120257849AActive Publication Date: 2025-07-04CHANGCHUN UNIV OF TECH

Patent Information

Application Number
CN202510732837.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Traditional large bulk equipment control methods are difficult to achieve multi-source information fusion in complex environments, resulting in control deviations and equipment failures. The existing control technology has poor adaptability to dynamically changing working conditions and insufficient stability. The control algorithm based on strategy gradients has poor convergence during training.

Method used

The cross-modal attention LSTM-PPO decision-making network is adopted, through the fusion of multi-source sensor data, and multi-source data such as lidar point cloud and visual images, combined with the timing modeling of the LSTM network and PPO algorithm, the motion control strategy is optimized to achieve adaptive decision-making on the dynamic environment.

Benefits of technology

It improves the target identification and tracking capabilities of large bulk equipment in complex environments, enhances the adaptability and control accuracy of the equipment, ensures that the equipment operates within the optimal range, and improves dynamic response characteristics and long-term operation reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257849A_ABST
    Figure CN120257849A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale bulk cargo equipment motion control optimization method based on deep learning, and relates to the field of computer systems, deep learning, intelligent control and industrial automation based on specific calculation models. Aiming at the problems of incomplete environment perception, poor dynamic environment adaptability and the like in a bulk material equipment control system, the method comprises the following steps: firstly, acquiring data and preprocessing through a multi-source sensor; secondly, constructing a cross-modal attention fusion model to realize feature alignment; and finally, an LSTM-PPO hybrid network architecture is designed, an LSTM layer processes time sequence state characteristics, and a PPO algorithm realizes control strategy optimization. Compared with the prior art, the precision and controllability in motion control of traditional large bulk cargo equipment can be improved, the operation efficiency and robustness of the system can be improved more easily, and the method can be widely applied to the fields of logistics, bulk cargo loading and unloading, industrial intelligent manufacturing and production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine learning, deep learning, intelligent manufacturing and industrial automation, and in particular to a motion control optimization method for large-scale bulk material equipment based on deep learning. Background Art

[0002] In modern industrial production, large-scale bulk material equipment is widely used in places such as mines, ports and power plants for loading, unloading and stacking of large-scale materials. Due to the complex and changeable working environment, the traditional single sensor perception method is difficult to fully capture the material status and equipment operating conditions, which may lead to control deviations, thereby affecting operating efficiency and even causing equipment failures. Therefore, realizing intelligent perception and decision-making based on multi-source information fusion is crucial to improving the automation level of bulk material equipment.

[0003] In the field of large-scale bulk material equipment control, precise control of the material conveying process is a key link in industrial production. In order to ensure that it meets the operating efficiency and safety standards, industry specifications have set strict technical indicators for parameters such as equipment motion trajectory and operating speed. Due to the strong coupling characteristics between the various control parameters of the system, the operating state of the equipment is easily disturbed by material characteristics, environmental factors, etc., resulting in control imbalance, which poses a severe challenge to the efficient and stable operation of the equipment. Therefore, realizing intelligent control of large-scale bulk material equipment has become a key issue that needs to be solved to ensure production efficiency and operational safety. This type of motion control optimization is not only a complex nonlinear control problem, but also a system engineering involving multi-sensor fusion and multi-objective optimization. The relevant research on its precise control and stable operation has become a key research direction in the field of industrial automation. Traditional control methods usually rely on single-mode sensor data and empirical models, but in practical applications, vision-based methods are easily affected by light and weather, and lidar performance degrades in dusty environments. On the one hand, the traditional control methods based on physical mechanism modeling in the existing control technology system have poor adaptability to dynamically changing working conditions. On the other hand, advanced control algorithms based on policy gradients generally have technical bottlenecks such as insufficient stability and poor convergence during the training process.

[0004] In view of the above deficiencies of the prior art, a cross-modal attention LSTM-PPO (Long Short-Term Memory-Proximal Policy Optimization) decision network that integrates multi-source data is proposed to optimize motion control. It realizes precise perception and control of complex working conditions, makes up for the advantages of multi-source features, and constructs a more complete environmental perception system. Based on the Proximal Policy Optimization (PPO) algorithm, a cross-modal attention mechanism and a Long Short-Term Memory (LSTM) time series modeling module are introduced. The adaptability of the system to dynamic working conditions is enhanced through time series feature extraction, making the training process more stable and the control strategy better. Summary of the Invention

[0005] The present invention proposes a motion control optimization method for large-scale bulk material equipment based on deep learning. The motion control of large-scale bulk material equipment is optimized through a cross-modal attention mechanism and an LSTM-PPO network model. The cross-modal attention mechanism can effectively integrate multi-source data such as lidar point cloud and visual images. The PPO algorithm can automatically learn the weight distribution of different sensor data, significantly improving the feature fusion effect. In addition, the LSTM network can accurately capture the time series dependence of the equipment operation state and solve the influence of the historical operation state on the current control decision. This method can stabilize the equipment operation parameters within the optimal range, improve its dynamic response characteristics, enhance the system's adaptability, and ensure long-term operation reliability.

[0006] To achieve the above object, the following technical solutions are adopted:

[0007] Step 1: Data collection and preprocessing are carried out through multi-source sensors. Based on the curvature feature and the improved ORB (Oriented FAST and Rotated BRIEF) algorithm, point cloud and image features are extracted respectively to realize the collaborative processing of multi-source data.

[0008] Step 1.1: The lidar scans the stockyard environment to obtain the three-dimensional point cloud data of the stockpile in real time, including the shape, height, and surface contour information of the stockpile. The RGB industrial camera collects the surface image data of the stockpile, extracts the material color distribution, texture features, and particle size information to assist in identifying the material type and stacking state. The inertial sensor monitors the motion state of the equipment in real time, including the spatial position, acceleration, and angular velocity, to compensate for the measurement error caused by the equipment vibration. The laser material flow meter monitors the material flow data in real time to ensure the timeliness and accuracy of the material flow data.

[0009] Step 1.2: For the point cloud data obtained by the lidar, a feature point detection method based on curvature is adopted. For any point in the point cloud , construct its local neighborhood point set , and characterize the local geometric features by calculating the covariance matrix of the neighborhood points :

[0010]

[0011]

[0012] where is a certain point in the neighborhood, is the centroid of the neighborhood point set, is the transpose. Obtain the eigenvalues from the covariance matrix, and then calculate the curvature of point .

[0013]

[0014]

[0015] Retain the points that satisfy , is the threshold, is the local point cloud density, and screen out the key features of the stockpile contour.

[0016] Step 1.3: Use an improved ORB feature detector to extract visual features from the images collected by the RGB industrial camera. Use the FAST feature detection algorithm with an adaptive threshold to calculate the feature point response value:

[0017]

[0018] where represents the detection window centered on , is the pixel intensity value, is the central pixel intensity value. At the same time, dynamically control the number of feature points according to the overall contrast of the image:

[0019] N f = N max [ 1 − exp ( − k φ ) ]

[0020] where is the maximum number of feature points, represents the image contrast metric value, is the adjustment coefficient. This ensures that stable image feature points can be extracted under different lighting conditions.

[0021] Step 2: Construct a cross-modal attention feature fusion model to enhance the correlation between 3D point clouds and 2D visual features, and achieve the complementary advantages and collaborative perception of multi-modal data.

[0022] Step 2.1: The cross-modal attention fusion model uses the multi-head scaled dot-product attention mechanism to establish geometric consistency constraints on the spatial alignment module, and stably extracts the common features and complementary information of different modalities in complex environments. It maps them to a unified feature space through a learnable linear transformation:

[0023]

[0024] where is the actively queried feature, is the feature to be matched, is the actually transmitted information, is the point cloud feature matrix, is the image feature matrix, , , are learnable projection matrices.

[0025] The Hamming distance is used to measure the feature similarity:

[0026]

[0027]

[0028] where , are the descriptors of feature points A and B respectively, is the dynamic Hamming distance threshold, is the base threshold, is the threshold adjustment amplitude, is the image normalized contrast, and the matching pairs that retain are used as candidate corresponding points. The rough matching stage can quickly screen out potential matching relationships.

[0029] Step 2.2: Then, perform feature fine matching and construct an optimization objective function that includes geometric error and motion prior :

[0030] E ( T ) = ∑ [ ρ ( ‖ π ( T · p i ) − q i ‖ 2 ) ] + λ ‖ log ( T · T prev − 1 ) ‖ F 2

[0031]

[0032] where is the transformation matrix describing the projection relationship from the point cloud to the image, is the projected 2D image point, is the Huber robust kernel function, is the camera projection model, is the regularization coefficient, represents the transformation matrix of the previous frame, is the input error, is the threshold parameter. The objective function is iteratively solved by the Gauss-Newton algorithm until convergence.

[0033] Step 2.3: Perform a non-linear mapping on the optimized matching result obtained from the optimized objective function, and use the multi-head attention mechanism to capture the dependency relationships between features from multiple dimensions. The attention weights are calculated as follows:

[0034]

[0035] where is the association strength between each point of the point cloud and each region of the image, is the scaling factor, is the row-wise normalization process.

[0036] Adopt the multi-head scaled dot-product attention mechanism to expand the projected features into attention heads, and each attention head is calculated independently. The expansion calculation of the multi-head attention is as follows:

[0037]

[0038]

[0039] where is the output of the single-head attention, is the output projection matrix. Design the cross-modal attention weight mechanism to calculate as follows:

[0040]

[0041] Finally, output the environmental state representation :

[0042]

[0043] where is the fused feature output by the cross-modal attention, is the feed-forward neural network, is the layer normalization process.

[0044] Step 3: Construct an LSTM-PPO hybrid decision network and train to obtain the optimized result.

[0045] Step 3.1: The LSTM-PPO decision network architecture is based on the LSTM network and the PPO algorithm to achieve precise control of the movement of large bulk material equipment. The network model consists of an input layer, an LSTM layer, a fully connected layer, an Actor (strategy) network, and a Critic (value) network.

[0046] The input layer receives multi-source data after feature extraction and matching, including environmental state representation, equipment operation status and material flow data; the LSTM layer uses the gating mechanism to capture the time series dependencies in the data and explore the long-term and short-term changes in the equipment motion state; the fully connected layer integrates the features of the LSTM layer output; the Actor generates the equipment motion control strategy based on the integrated features, and the Critic evaluates the value of the strategy and provides feedback for optimization.

[0047] The number of hidden units in the LSTM layer is 128, and the number of layers is 2. The bidirectional structure enhances the perception of historical and future information, ensuring that the network can fully learn the dynamic changes in the equipment movement process. Output short-term memory as follows:

[0048]

[0049] in is the unit state, For output gating, is the activation function.

[0050] Step 3.2: Fusion of LSTM encoded temporal features , construct an Actor-Critic framework, in which the Actor network is responsible for strategy generation and the Critic network is responsible for value evaluation. The PPO algorithm is integrated into the decision network, and the iterative update of the control strategy is achieved by alternately optimizing the strategy network and the value network.

[0051] Actor output equipment continuous control action :

[0052] a t = tanh [ W a ⋅ Re LU ( W 1 ⋅ h t + b 1 ) + b a ]

[0053] in , is the fully connected layer parameter, is the activation function. , is the output layer parameter, Constrain the action to [ − 1 , 1 ] , and then linearly mapped to the actual control range.

[0054] The Critic predicts the long-term operating benefits of the equipment's current state, guides the direction of policy optimization, and estimates the discounted cumulative reward starting from the state at time as follows: It is calculated as follows:

[0055]

[0056] where , are the parameters of the fully connected layer, , are the parameters of the output layer.

[0057] Step 3.3: Adopt an end-to-end training method based on the PPO algorithm, set the training batch size to 64, and the number of optimization iterations to 10. Using the cumulative discounted return as the optimization goal can limit the amplitude of policy updates and avoid policy collapse caused by overly large single updates. The advantage function is calculated as follows:

[0058]

[0059]

[0060] where the discount factor is set to 0.95, the smoothing coefficient is set to 0.9, is the immediate reward, is the temporal difference error at time

[0061] L CLIP ( θ ) = E t { min [ r t ( θ ) A t , clip ( r t ( θ ) , 1 − ε , 1 + ε ) A t ] }

[0062]

[0063] where is the ratio of the new and old policies, is the clipping threshold set to 0.2, is a random policy. The training objective function and the total loss function are calculated as follows:

[0064]

[0065] where the Critic loss weight is set to 0.5, is the policy entropy weight set to 0.02, is the value loss function, is the policy entropy.

[0066] Monitor the changes of the loss function and performance metrics during the network training process in real time, adjust the parameters in a timely manner to ensure the network generalization ability, and obtain the final optimized control strategy.

[0067] The beneficial effects of the present invention are as follows:

[0068] 1. The motion control optimization method for large bulk material equipment first adopts a collaborative extraction method based on curvature features and improved ORB algorithm for multi-source data, solves the problems of inconsistent feature scales and low matching accuracy between point cloud and image data, can improve the pose estimation accuracy of large bulk material equipment, and reduce the control deviation caused by sensor data mismatch.

[0069] 2. By constructing a cross-modal attention feature fusion model, solve the problem of weak correlation in cross-modal feature fusion, realize high-precision alignment and complementary enhancement of multi-modal data, and can improve the target recognition and tracking ability of large bulk material equipment in complex environments.

[0070] 3. Adopt an LSTM-PPO hybrid decision-making network, use the long-time sequence dependence modeling ability of LSTM to extract state features, and combine the policy gradient optimization of the PPO algorithm to achieve adaptive decision-making in dynamic environments, solve the problem of gradient disappearance in long-time sequence dependence modeling, and enable large bulk material equipment to have better environmental adaptability and control accuracy.

[0071] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0072] Figure 1 is a schematic flow chart of the motion control optimization for large bulk material equipment of the present invention;

[0073] Figure 2 is the structural diagram of the cross-modal attention fusion model. Detailed Embodiments

[0074] The following will specifically describe the detailed embodiments of the present invention with reference to the accompanying drawings.

[0075] A motion control optimization method for large bulk material equipment based on deep learning, please refer to the attached Figure 1 , including the following steps:

[0076] Step 1: Collect the data of the stockyard environment through multi-source sensors, extract the robust features of the point cloud and the image respectively based on the curvature features and the improved ORB algorithm, and use spatio-temporal alignment and normalization processing to realize the collaborative preprocessing of multi-source data, providing highly consistent input for cross-modal fusion.

[0077] Step 2: Construct a cross-modal attention fusion model. Through feature matching optimization and dynamic weight allocation under geometric constraints, enhance the semantic correlation between point cloud and visual features, and output a multi-modal joint feature representation with complementary advantages.

[0078] Step 3: Construct an LSTM-PPO hybrid decision network. Utilize the long-term temporal dependence modeling ability of LSTM to extract state features, and combine the policy gradient optimization of the PPO algorithm to achieve adaptive decision-making in a dynamic environment, obtaining the final optimized policy.

[0079] Implementation Step 1: Conduct data collection and preprocessing through multi-source sensors, and extract point cloud and image features based on curvature features and improved ORB algorithms respectively.

[0080] Step 1.1: The lidar scans the stockyard environment to obtain real-time three-dimensional point cloud data of the stockpile, including the shape, height, and surface contour information of the stockpile. The RGB industrial camera collects the surface image data of the stockpile, extracts the material color distribution, texture features, and particle size information to assist in identifying the material type and stacking state. The inertial sensor monitors the motion state of the equipment in real time, including spatial position, acceleration, and angular velocity, which is used to compensate for the measurement errors caused by equipment vibration. The laser material flow meter monitors the material flow data in real time to ensure the real-time and accuracy of the material flow data.

[0081] Step 1.2: Adopt a curvature-based feature point detection method for the point cloud data obtained by the lidar. For any point in the point cloud, construct its local neighborhood point set , and characterize the local geometric features by calculating the covariance matrix of this neighborhood point:

[0082]

[0083]

[0084] where is a certain point in the neighborhood, is the centroid of the neighborhood point set, is the transpose. Obtain the eigenvalues according to the covariance matrix, and then calculate the curvature of point .

[0085]

[0086]

[0087] Retain the points that satisfy , is the threshold, is the local point cloud density, and screen out the key features of the stockpile contour.

[0088] Step 1.3: Extract visual features from the images collected by the RGB industrial camera using an improved ORB feature detector. Use the FAST feature detection algorithm with an adaptive threshold to calculate the response value of the feature points:

[0089]

[0090] where represents the detection window centered at , is the pixel intensity value, is the central pixel intensity value. At the same time, dynamically control the number of feature points according to the overall contrast of the image:

[0091] N f = N max [ 1 − exp ( − k φ ) ]

[0092] where is the maximum number of feature points, represents the image contrast metric value, is the adjustment coefficient. This ensures that stable image feature points can be extracted under different lighting conditions.

[0093] Implementation step 2: Based on the features extracted in implementation step 1, refer to Appendix Figure 2 , construct a cross-modal attention fusion model, and enhance the semantic relevance between the point cloud and visual features through feature matching optimization and dynamic weight allocation, and output a multi-modal environmental state representation with complementary advantages.

[0094] Step 2.1: For the cross-modal attention fusion model, by adopting the multi-head scaled dot-product attention mechanism, establish geometric consistency constraints on the spatial alignment module, and stably extract the common features and complementary information across modalities in a complex environment. Use a learnable linear transformation to map it to a unified feature space:

[0095]

[0096] where is the actively queried feature, is the feature to be matched, is the actually transmitted information, is the point cloud feature matrix, is the image feature matrix, , , are learnable projection matrices.

[0097] Use the Hamming distance to measure the feature similarity:

[0098]

[0099]

[0100] Among them 、 are the descriptors of feature points A and B respectively, is the dynamic Hamming distance threshold, is the base threshold, is the threshold adjustment amplitude, is the image normalized contrast, and the retained matching pairs are used as candidate corresponding points. In the rough matching stage, potential matching relationships can be quickly screened out.

[0101] Step 2.2: Then perform feature fine matching and construct an optimization objective function that includes geometric error and motion prior :

[0102] E ( T ) = ∑ [ ρ ( ‖ π ( T · p i ) − q i ‖ 2 ) ] + λ ‖ log ( T · T prev − 1 ) ‖ F 2

[0103]

[0104] Among them is the transformation matrix describing the projection relationship from the point cloud to the image, is the projected 2D image point, is the Huber robust kernel function, is the camera projection model, is the regularization coefficient, represents the transformation matrix of the previous frame, is the input error, is the threshold parameter. The objective function is iteratively solved by the Gauss-Newton algorithm until convergence.

[0105] Step 2.3: Non-linearly map the optimized matching results obtained from the optimization objective function, and use the multi-head attention mechanism to capture the dependency relationships between features from multiple dimensions. The attention weight is calculated as follows:

[0106]

[0107] Among them is the association strength between each point of the point cloud and each region of the image, is the scaling factor, is the row-wise normalization process.

[0108] Adopt the multi-head scaled dot-product attention mechanism to expand the projected features into attention heads, and each attention head is calculated independently. The multi-head attention expansion calculation is as follows:

[0109]

[0110]

[0111] in is the single-head attention output, is the output projection matrix. The cross-modal attention weight mechanism is designed to be calculated as follows:

[0112]

[0113] Finally, the environmental status representation is output :

[0114]

[0115] in is the fusion feature of the cross-modal attention output, is a feed-forward neural network, Normalize the layer.

[0116] Implementation step 3: Construct an LSTM-PPO hybrid decision network, input the fusion features generated in implementation step 2, use the long-term dependency modeling capability of LSTM to extract state features, and combine the policy gradient optimization of the PPO algorithm to achieve adaptive decision-making in a dynamic environment to obtain the final optimization strategy.

[0117] Step 3.1: The LSTM-PPO decision network architecture is based on the LSTM network and the PPO algorithm to achieve precise control of the movement of large bulk material equipment. The network model consists of an input layer, an LSTM layer, a fully connected layer, an Actor network, and a Critic network.

[0118] The input layer receives multi-source data after feature extraction and matching, including environmental state representation, equipment operation status and material flow data; the LSTM layer uses the gating mechanism to capture the time series dependencies in the data and explore the long-term and short-term changes in the equipment motion state; the fully connected layer integrates the features of the LSTM layer output; the Actor generates the equipment motion control strategy based on the integrated features, and the Critic evaluates the value of the strategy and provides feedback for optimization.

[0119] The number of hidden units in the LSTM layer is 128, and the number of layers is 2. The bidirectional structure enhances the perception of historical and future information, ensuring that the network can fully learn the dynamic changes in the equipment movement process. Output short-term memory as follows:

[0120]

[0121] in is the unit state, is the output gating, is the activation function.

[0122] Step 3.2: Fuse the temporal features encoded by LSTM , construct the Actor-Critic framework, where the Actor network is responsible for policy generation, the Critic network conducts value evaluation, integrate the PPO algorithm into the decision-making network, and realize the iterative update of the control strategy by alternately optimizing the policy network and the value network.

[0123] The Actor outputs continuous control actions for the equipment :

[0124] a t = tanh [ W a ⋅ Re LU ( W 1 ⋅ h t + b 1 ) + b a ]

[0125] where , are the parameters of the fully connected layer, is the activation function. , are the parameters of the output layer, Constrain the action to [ − 1 , 1 ] , and then linearly map it to the actual control range.

[0126] The Critic predicts the long-term operating benefit of the current state of the equipment, guides the direction of policy optimization, and estimates the discounted cumulative reward starting from the state at time as follows: The discounted cumulative reward is calculated as follows:

[0127]

[0128] where , are the parameters of the fully connected layer, , are the parameters of the output layer.

[0129] Step 3.3: Adopt an end-to-end training method based on the PPO algorithm, set the training batch size to 64, and the number of optimization iterations to 10. Using the cumulative discounted return as the optimization goal can limit the amplitude of policy updates and avoid policy collapse caused by overly large single updates. The advantage function is calculated as follows:

[0130]

[0131]

[0132] r t = [ 0 . 6 α + 0 . 7 β ]

[0133] where the discount factor is set to 0.95, and the smoothing coefficient is set to 0.9, is the immediate reward, is the time required to reach the control target, is the spatial offset between the target position and the actual position, is the temporal difference error at time, and set the objective function:

[0134] L CLIP ( θ ) = E t { min [ r t ( θ ) A t , clip ( r t ( θ ) , 1 − ε , 1 + ε ) A t ] }

[0135]

[0136] where is the ratio of the new and old policies, is the clipping threshold set to 0.2, is a random policy. Train the objective function, and the total loss function is calculated as follows:

[0137]

[0138] where the Critic loss weight is set to 0.5, is the policy entropy weight set to 0.02, is the value loss function, is the policy entropy.

[0139] By monitoring the changes in the loss function and performance metrics during the network training process, adjust the parameters in a timely manner to ensure the convergence of the network training and obtain the final optimized control strategy.

Claims

1. An optimization method for the motion control of large bulk material equipment based on deep learning, characterized in that, It includes the following steps: Step 1: Collect the stockyard environment data through multi-source sensors, extract the robust features of the point cloud and the image based on the curvature feature and the improved ORB algorithm respectively, and use spatio-temporal alignment and normalization processing to achieve the collaborative preprocessing of multi-source data; Step 2: Construct a cross-modal attention fusion model, and output the multi-modal joint feature expression with complementary advantages through feature matching optimization and dynamic weight allocation under geometric constraints. It specifically includes the following steps: Step 2.1: In the cross-modal attention fusion model, by adopting the multi-head scaled dot-product attention mechanism, establish geometric consistency constraints on the spatial alignment module, stably extract the common features and complementary information across modalities in complex environments, and map them to a unified feature space by using a learnable linear transformation: Among them is the feature for active query is the feature to be matched is the actually transmitted information is the point cloud feature matrix is the image feature matrix and and are learnable projection matrices, using Hamming distance to measure feature similarity Among them and are the descriptors of feature points A and B respectively, is the dynamic Hamming distance threshold, is the basic threshold, is the threshold adjustment amplitude, is the image normalized contrast, and the matching pairs that remain are used as candidate corresponding points; Step 2.2: Then, perform feature fine matching to construct an optimization objective function that includes geometric error and motion prior : where is the transformation matrix describing the projection relationship from the point cloud to the image, is the projected 2D image point, is the Huber robust kernel function, is the camera projection model, is the regularization coefficient, represents the transformation matrix of the previous frame, is the input error, is the threshold parameter, and the objective function is iteratively solved by the Gauss-Newton algorithm until convergence; Step 2.3: Perform a non-linear mapping on the optimized matching results obtained from the optimized objective function, and use the multi-head attention mechanism to capture the dependency relationships between features from multiple dimensions. The attention weights are calculated as follows: wherein is the association strength between each point of the point cloud and each region of the image, is the scaling factor, is the row-wise normalization process; Using the multi-head scaled dot-product attention mechanism, the projected features are extended to attention heads, and each attention head is calculated independently. The extended calculation of multi-head attention is as follows: Among them is the single-head attention output, is the output projection matrix, and the cross-modal attention weight mechanism is designed and calculated as follows: Final output environmental state characterization : Among them is the fused feature of the cross-modal attention output, is the feed-forward neural network, is the layer normalization process; Step 3: Construct an LSTM-PPO hybrid decision network, use the long-time sequence dependence modeling ability of LSTM to extract state features, and combine the policy gradient optimization of the PPO algorithm to achieve adaptive decision-making in a dynamic environment, and obtain the final optimized strategy. It specifically includes the following steps: Step 3.1: The LSTM-PPO decision network architecture is based on the LSTM network and the PPO algorithm to achieve precise control of the movement of large-scale bulk material equipment. The network model consists of an input layer, an LSTM layer, a fully connected layer, an Actor network, and a Critic network; The input layer receives the multi-source data after feature extraction and matching, including the environmental state representation, the equipment operation state, and the material flow data; the number of hidden units in the LSTM layer is 128, and the number of layers is 2; Step 3.2: Integrate the temporal features encoded by LSTM , construct an Actor-Critic framework, and realize the iterative update of the control strategy by alternately optimizing the policy network and the value network. The Actor outputs continuous control actions for the equipment :[[]]END]] Among them , are the parameters of the fully connected layer, is the activation function, , are the parameters of the output layer, constrain the action to , and then linearly map it to the actual control range; The Critic estimates the discounted cumulative reward starting from the state at time as follows: It is calculated as follows: Among them , are the parameters of the fully connected layer, , are the parameters of the output layer; Step 3.3: Adopt an end-to-end training method based on the PPO algorithm, set the training batch size to 64, the number of optimization iterations to 10, use the cumulative discounted return as the optimization goal, limit the amplitude of policy updates, and avoid policy collapse caused by overly large single updates. The advantage function is calculated as follows: Among them, the discount factor is set to 0.95, and the smoothing coefficient is set to 0.

9. is the immediate reward, is the time required to achieve the control target, is the spatial offset between the target position and the actual position, is the temporal difference error at time, and set the objective function: where is the ratio of the new and old strategies, the pruning threshold is set to 0.2, is a random strategy, and the training objective function and total loss function are calculated as follows: Among them, the weight of the Critic loss is set to 0.5, the weight of the policy entropy is set to 0.02, is the value loss function, is the policy entropy.

2. The motion control optimization method for large bulk material equipment based on deep learning according to claim 1, characterized in that The step of collecting the stockyard environment data through multi-source sensors in Step 1, extracting the robust features of the point cloud and the image based on the curvature feature and the improved ORB algorithm respectively, and using spatio-temporal alignment and normalization processing to achieve the collaborative preprocessing of multi-source data, the specific steps are as follows: Step 1.1: Use a lidar to obtain the three-dimensional point cloud data of the stockpile in real time, including the shape, height, and surface contour information of the stockpile. An RGB industrial camera collects the surface image data of the stockpile, extracts the material color distribution, texture features, and particle size information to assist in identifying the material type and stacking state. An inertial sensor obtains the equipment operation state in real time, including the spatial position, acceleration, and angular velocity. A laser material flow meter obtains the material flow data in real time; Step 1.2: Adopt a curvature-based feature point detection method for the point cloud data obtained by the lidar. For any point in the point cloud , construct its local neighborhood point set , and characterize the local geometric features by calculating the covariance matrix of the neighborhood points : Among them is a certain point within the neighborhood is the centroid of the neighborhood point set is the transpose, and eigenvalues are obtained from the covariance matrix , then the curvature of point is calculated ; Retain the points that satisfy . is the threshold value, is the local point cloud density, and the key features of the stockpile contour are screened; Step 1.3: Use an improved ORB feature detector to extract visual features from the images collected by the RGB industrial camera, and use the FAST feature detection algorithm with an adaptive threshold to calculate the response value of the feature points: Among them represents a detection window centered on The pixel intensity value is the pixel intensity value, and the center pixel intensity value is , and at the same time, the number of feature points is dynamically controlled according to the overall contrast of the image is the center pixel intensity value, and at the same time, the number of feature points is dynamically controlled according to the overall contrast of the image quantity: wherein is the maximum feature point, represents the image contrast metric value, is the adjustment coefficient.

Citation Information

Patent Citations

  • Industrial process optimization method based on deep reinforcement learning

    CN116842856A

  • Industrial automatic intelligent manufacturing management system based on block chain technology

    CN118627841A

  • Cloud edge computing task scheduling method based on reinforcement learning

    CN118740835A

  • Predictive Model Data Stream Prioritization

    US20230123322A1

  • Three-dimensional target detection method based on multimodal fusion and depth attention mechanism

    US20250037299A1

Cited By

  • Bulk cargo storage yard automatic unloading control method and device

    CN121074865A

  • Robot obstacle avoidance method and system based on three-dimensional point cloud deep learning

    CN121500981A