Multi-roller cooperative belt deviation rectification control method and system

Through the multi-agent deep reinforcement learning model and the space-time graph attention mechanism, combined with visual edge detection and action reconstruction model, the response hysteresis and instability of belt deviation correction control in long distances and complex environments is solved, and an efficient, stable and safe belt conveying system is achieved.

CN120156828AActive Publication Date: 2025-06-17UNIV OF SCI & TECH BEIJING

Patent Information

Application Number
CN202510643888.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing belt deviation correction control technology has problems such as response hysteresis, accuracy misalignment and system instability in long-distance conveying and complex environments. Especially in multi-roll collaborative control, there is a lack of effective collaborative control strategies and safety constraint mechanisms.

Method used

A multi-agent deep reinforcement learning model is adopted to combine the space-time graph attention mechanism to build a collaborative bias correction control strategy, generate safe actions through visual edge detection and offline training, and introduce action reconstruction models in actual deployment to ensure the safety and physical feasibility of actions.

Benefits of technology

It improves the operating efficiency and stability of the belt conveyor system in long distances and complex environments, significantly improves the system's adaptability and safety, and avoids the risk of belt damage caused by wrong decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120156828A_ABST
    Figure CN120156828A_ABST
Patent Text Reader

Abstract

The invention provides a multi-roller cooperative belt deviation rectification control method and system, and the method comprises the steps: comparing a current belt edge with a belt edge set value, obtaining a deviation distance # imgabs1 # and a deviation speed # imgabs2 # of a pixel correlation point # imgabs0 # at the intersection of each deviation rectification roller and a standard edge line, and constructing a state vector # imgabs4 # in combination with a deviation rectification roller angle # imgabs3 #; # imgabs5 outputs an action through a strategy network of a reinforcement learning model, combines the action and a state vector into a state action pair, inputs an action reconstruction model to reconstruct the action, and outputs a safety action for subsequent motor control; the state action pair also inputs the action value network and the action cost network of the reinforcement learning model, and evaluates the action output by the strategy network. The belt deviation rectifying device can perform deviation rectifying control on the belt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of belt deviation correction control, and particularly to a multi-roller collaborative belt deviation correction control method and system. Background Art

[0002] Belt conveyor systems are key material transportation equipment in mines, ports, logistics, and manufacturing. However, during their operation, belt deviation often occurs due to complex structures, changing environments, and long-term operation. Such deviation can lead to equipment wear, material leakage, energy waste, and associated downtime failures, seriously affecting production efficiency and equipment lifespan. Therefore, precise deviation correction control is crucial for ensuring production safety and economy. Currently, in industrial sites, the treatment of belt deviation mainly relies on installing deviation correction devices, such as deviation correction rollers, and adjusting the lateral friction force of the belt or the roller angle of these devices manually or automatically to gradually guide the belt back to the normal trajectory. Due to its high execution efficiency and low transformation cost, this method has become a common solution in engineering applications.

[0003] Since belt deviation is a progressive process, its local deviation will gradually spread through the overall transmission effect of the belt, ultimately leading to global operation anomalies. Although the current technical practice of using a single deviation correction roller for a specific area is relatively common, it has obvious limitations - especially in long-distance transportation scenarios, the effective range of a single deviation correction roller usually does not exceed 20 meters. When the transportation distance exceeds 50 meters, multiple sets of deviation correction units need to be configured. However, the parallel deployment of multiple rollers will cause new technical bottlenecks. Synergistic problems such as the control timing matching and force coupling between the deviation correction rollers require the establishment of a real-time communication mechanism. Otherwise, a deviation correction torque counteracting phenomenon may occur. Such non-coordinated actions not only cannot correct the deviation but will instead increase the risk of system oscillation, making the complexity of deviation correction control increase exponentially. This technical dilemma not only reveals the essential difference between single-point control and global optimization but also confirms the inevitable trend of the deviation correction system to transform towards distributed intelligent control.

[0004] There are still many problems in this transformation requirement in the traditional control architecture. For example, although the conventional PID control algorithm can achieve the basic deviation correction function, its fixed parameter settings are difficult to adapt to the dynamic changes of belt tension; the mechanical linkage device based on preset thresholds can complete simple coordination but cannot analyze the non-linear coupling relationship between multiple rollers. More critically, the traditional system lacks the ability to predict the deviation propagation path and often misses the best regulation window period when obvious deviation is sensed. These inherent defects lead to the dual dilemmas of system response lag and deviation correction accuracy inaccuracy in the existing solutions when dealing with long-distance transportation and variable load conditions - neither can effectively intercept the deviation during the diffusion stage nor maintain trajectory stability during dynamic changes. This structural lack of control ability is precisely the root cause for most industrial scenarios to be unable to get rid of manual real-time monitoring and compensation operations.

[0005] On the other hand, adopting intelligent control strategies also faces problems such as difficulties in obtaining multi-roller collaborative training data and security issues during model deployment. First of all, constructing an efficient intelligent control method usually relies on a large amount of high-quality data for training and optimization. However, in the actual industrial environment, the complex operating conditions and the diversity of equipment types make it extremely difficult to obtain comprehensive and rich training data. In addition, the online learning and data acquisition processes may interfere with normal production operations, which further limits the effective application of traditional data-driven methods. At the same time, in the actual system deployment, the design of the deviation correction control strategy must ensure the safety of the belt conveyor system to avoid equipment damage, belt tearing or overall unstable operation caused by excessive control or frequent adjustments. Therefore, the control system not only needs to achieve the effect of deviation correction, but also must strictly follow the physical limitations of the equipment and have good robustness to adapt to the changing industrial environment. However, many intelligent control strategies have deficiencies in the modeling of safety and robustness, which may lead to failures and safety hazards in actual operation. Summary of the Invention

[0006] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a multi-roller collaborative belt deviation correction control method and system, and the technical solutions are as follows:

[0007] On the one hand, a multi-roller collaborative belt deviation correction control method is provided, and the method includes:

[0008] S1. Obtain the current belt edge through visual edge detection technology;

[0009] S2. Compare the current belt edge with the set value of the belt edge to obtain the pixel correlation points at the intersection of each deviation correction roller and the standard edge line of the belt edge deviation value, including the offset distance and the offset speed , combined with the deviation correction roller angle , construct a state vector , where i is the deviation correction roller number and n is the total number of deviation correction rollers;

[0010] S3. The state vector passes through the policy network of the multi-agent deep reinforcement learning model completed by offline training to output an action, combines the action with the state vector to form a state-action pair, inputs the action reconstruction model completed by offline training, reconstructs the action, and outputs a safe action for subsequent motor control;

[0011] S4. The state-action pair will also be input into the action-value network and action-cost network of the multi-agent deep reinforcement learning model that has completed offline training to evaluate the action output by the policy network, and online train the policy network, action-value network, and action-cost network;

[0012] S5. The output safe action is processed and converted into an electrical signal form and input into the servo motor driver, and the driver controls the rotation of each deviation correction roller servo motor to perform belt deviation correction control.

[0013] On the other hand, a multi-roller collaborative belt deviation correction control system is provided. The system includes:

[0014] A detection module, configured to detect the current belt edge through visual edge detection technology;

[0015] A construction module, configured to compare the current belt edge with the set value of the belt edge to obtain the pixel correlation points at the intersection of each deviation correction roller and the standard edge line of the belt edge deviation value, including the offset distance and the offset speed , and combine with the deviation correction roller angle to construct a state vector , where i is the deviation correction roller number and n is the total number of deviation correction rollers;

[0016] An output module, configured to output an action through the policy network of the multi-agent deep reinforcement learning model that has completed offline training for the state vector , combine the action with the state vector to form a state-action pair, input it into the action reconstruction model that has completed offline training to reconstruct the action, and output a safe action for subsequent motor control;

[0017] An evaluation module, configured to input the state-action pair into the action-value network and action-cost network of the multi-agent deep reinforcement learning model that has completed offline training to evaluate the action output by the policy network, and online train the policy network, action-value network, and action-cost network;

[0018] A deviation correction control module, configured to process the output safe action, convert it into an electrical signal form and input it into the servo motor driver, and the driver controls the rotation of each deviation correction roller servo motor to perform belt deviation correction control.

[0019] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory, and at least one instruction is stored in the memory. The at least one instruction is loaded and executed by the processor to implement the above multi-roller collaborative belt deviation correction control method.

[0020] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above multi-roller collaborative belt deviation correction control method.

[0021] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0022] 1) Aiming at the problem that multiple deviation correction rollers need to be adjusted simultaneously in a long-distance belt conveyor system, the present invention proposes a solution based on a multi-agent deep reinforcement learning model. By having multiple agents respectively learn the control strategies of their respective deviation correction rollers and relying on a linkage mechanism to ensure the coordinated actions among multiple deviation correction rollers, the overall operation efficiency and stability of the system are effectively improved. To capture the linkage relationship between different deviation correction rollers in time and space, the present invention constructs a spatio-temporal graph model. This model can effectively represent the relationships among various deviation correction rollers in the belt conveyor system, including their positions, states, and interactions with the belt. By introducing a spatio-temporal graph attention mechanism, the interactive influence among various deviation correction rollers can be dynamically weighted, ensuring that the information transmission and decision-making processes during the belt operation are effectively modeled. The spatio-temporal graph attention mechanism can intelligently evaluate the actions of different deviation correction rollers based on the current state and historical data, enabling each agent to not only consider its own state when making decisions but also effectively absorb the dynamic information from other agents. The introduction of this mechanism enables the model to ensure the consistency of time response while achieving spatial coordination, avoiding system instability caused by information lag or inconsistent decisions among the rollers. In case of emergencies, the system can quickly identify the deviation correction rollers that need to be adjusted and, through the cooperation among agents, quickly respond and make corresponding adjustments. This not only improves the operation efficiency of the long-distance belt conveyor system but also significantly enhances its adaptability and safety in complex working environments.

[0023] 2) To solve the problem of difficult acquisition of training data in the process of multi-agent deep reinforcement learning, the present invention adopts a method for enhancing multi-roller collaborative data. Offline data for multi-agent deep reinforcement learning with collaboration is generated through a diffusion model, making up for the deficiency of difficult acquisition of large-scale and high-quality training data in the actual environment. By enhancing the diversity and representativeness of the data, the generalization ability of the reinforcement learning model is greatly improved, enabling the model to better cope with complex operating environments.

[0024] 3) In the actual deployment of multi-roll collaborative control, the action strategies of each agent are crucial for the safety and stability of the system. Traditional methods usually rely only on simple constraints to limit the action range of the agent, and cannot effectively cope with the dynamic changes and potential risks in complex environments. The present invention further enhances the safety of the agent by introducing a safety constraint mechanism driven by deep learning. In the training stage of reinforcement learning, an action reconstruction model is used to achieve dual functions: on the one hand, the action reconstruction model is used to extract the distribution deviation degree of the agent's action strategy, and the KL divergence is calculated to quantify the difference between the action strategy and the prior distribution, so as to detect whether the action deviates from the normal range; on the other hand, the action reconstruction model also physically constrains the reconstructed action to ensure that the generated action conforms to the data distribution characteristics while following the predefined physical constraint conditions. When the action of the agent approaches or exceeds the predetermined safety boundary, the action reconstruction model can timely identify the distribution deviation degree of the action and impose physical constraints on the reconstructed action to prevent out-of-distribution actions that may lead to unsafe consequences. At the same time, the KL divergence and the physical constraint loss are calculated simultaneously during the training process to optimize the safety and physical feasibility of the model. This mechanism ensures that all agents execute actions within physical constraints and safety ranges, thereby significantly reducing the risk of belt damage caused by wrong decisions and enhancing the reliability and practical application value of the model. In addition, when it is detected that the action exceeds the normal distribution range, the system will replace the original action by sampling the nearest in-distribution action to further ensure the safety of the action. This innovative method not only ensures the safe operation of multi-agents in complex environments, but also provides strong support for their feasibility in actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 is a flowchart of a multi-roll collaborative belt deviation correction control method provided by an embodiment of the present invention;

[0027] Figure 2 is an overall block diagram of a multi-roll collaborative belt deviation correction control method provided by an embodiment of the present invention;

[0028] Figure 3 is a schematic diagram of the training process of a multi-roll collaborative belt deviation correction control method provided by an embodiment of the present invention;

[0029] Figure 4 is a structural block diagram of an agent deep reinforcement learning model provided by an embodiment of the present invention;

[0030] Figure 5 is the structural block diagram of the spatio-temporal graph attention mechanism provided by the embodiment of the present invention;

[0031] Figure 6 is the structural block diagram of each spatio-temporal graph attention layer provided by the embodiment of the present invention;

[0032] Figure 7 is the flowchart of the diffusion model data augmentation method based on collaborative reward drive provided by the embodiment of the present invention;

[0033] Figure 8 is the structural block diagram of the action reconstruction model provided by the embodiment of the present invention;

[0034] Figure 9 is the block diagram of a multi-roll collaborative belt deviation correction control system provided by the embodiment of the present invention;

[0035] Figure 10 is the schematic structural diagram of an electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0036] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0037] The embodiment of the present invention provides a multi-roll collaborative belt deviation correction control method, which can be implemented by an electronic device, and the electronic device can be a terminal or a server. Figure 1 The flowchart of the method is shown as follows, Figure 2 The overall block diagram of the method is shown as follows, and the processing flow may include the following steps:

[0038] S1. Obtain the current belt edge through visual edge detection technology;

[0039] The visual belt edge technology in the embodiment of the present invention can adopt existing technologies, which will not be elaborated here.

[0040] S2. Compare the current belt edge with the set value of the belt edge to obtain the pixel correlation points at the intersection of each deviation correction roll and the standard edge line of the belt edge deviation value, including the offset distance and the offset speed , and combine the deviation correction roll angle to construct a state vector , where i is the deviation correction roll number and n is the total number of deviation correction rolls;

[0041] S3. The state vector The policy network of the multi-agent deep reinforcement learning model completed through offline training outputs actions. The actions are combined with the state vector to form state-action pairs, which are input into the action reconstruction model completed through offline training to reconstruct the actions and output safe actions for subsequent motor control;

[0042] S4. The state-action pairs are also input into the action value network and action cost network of the multi-agent deep reinforcement learning model completed through offline training to evaluate the actions output by the policy network, and to online train the policy network, action value network, and action cost network;

[0043] S5. The output safe actions are processed and converted into the form of electrical signals and input into the servo motor driver. The driver controls the rotation of each deviation correction roller servo motor to perform belt deviation correction control (meanwhile, various state parameters of the motor are fed back to the motor and added to the state sequence of the motor. Subsequently, the edge of the belt is detected through vision to feedback the edge of the belt to achieve a closed loop).

[0044] Optionally, the offline training process of the multi-agent deep reinforcement learning model is as follows:

[0045] Collect multi-dimensional data during the normal operation of the belt by manually controlling the deviation correction roller motor, and record the real-time state characteristics of the th deviation correction roller, including the deviation correction roller angle and the motor rotation angle increment corresponding to controlling the action of the deviation correction roller . Based on the historical operation data, establish the standard edge position of the belt as the reference (as shown by the dotted line in the on-site belt multi-roller system Figure 3 ). At the standard edge position, define the pixel association points on the image for each deviation correction roller at the intersection of the standard edge line and the deviation correction roller . Real-time obtain the actual edge position of the belt through vision edge detection technology, and calculate the offset distance between the reference point and the actual edge of the belt at the same horizontal position and the offset speed . Construct a state vector containing , being the total number of deviation correction rollers. At the level of designing the reward function, define the immediate reward , where is the offset speed penalty coefficient. This design finally generates an offline training data set through non-linear weighted fusion of the position deviation and dynamic characteristics:

[0046] The data contains groups of complete interaction trajectories with a length of T, where Represents the multi-roller control action;

[0047] Then, offline training of the reinforcement learning model is carried out, including offline training of the policy network, action-value network, and action-cost network. The policy network outputs actions, the action-value network evaluates the long-term impact of the adjusted actions on the belt state, and affects the trend of the policy network to generate actions through the value output. The action-cost network evaluates the safety of actions and suppresses the generation of unsafe actions through the cost output;

[0048] For the requirement of multi-roller collaborative control of the belt conveying system, through the data augmentation method of the diffusion model driven by collaborative rewards, the exploration range of the multi-agent state-action space is effectively expanded, and the integrated data is put into the experience replay pool for the reinforcement learning model to train;

[0049] In addition, based on the action reconstruction model, a part of the model outputs the distribution deviation and action cost of the action for updating the action-cost network, and a part outputs the reconstructed action as the safe action of the actual output.

[0050] Optionally, the online training process of the multi-agent deep reinforcement learning model is as follows:

[0051] After the state-action pair is input into the action-value network and the action-cost network, the loss of the action-value network is: the long-term cumulative reward value output by the action-value network and the immediate reward after the action Calculate the loss. The loss of the action-cost network is: the distribution deviation output by the action reconstruction model and the cost value output by the action-cost network Calculate the loss, and use the weighted sum of the outputs of the two networks as the comprehensive reward , and organize it with the state-action pair into and send it into the experience replay pool, and optimize the policy network through mini-batch incremental training, positively reinforce the actions with high returns and low risks, and reduce the influence weight of high-risk actions to improve the adaptability and robustness of the system.

[0052] Optionally, as Figure 4 shown, the multi-agent deep reinforcement learning model consists of three networks: the action-value network, the action-cost network, and the policy network. All three networks take a [state, action] trajectory sequence as input, extract the corresponding feature vectors through the spatio-temporal graph attention mechanism, and finally output the corresponding targets through the fully connected layer;

[0053] Among them, the policy network , outputs the mean and variance of the Gaussian distribution, , and through action sampling Generate a random sampling action, which will be evaluated as the input of the action value network and the action cost network;

[0054] The action value network maps the feature vector extracted by the spatio-temporal graph attention mechanism to the long-term cumulative reward value , and updates the action value network based on the Bellman error loss , where represents the immediate reward obtained when taking action in state , is the discount factor, with a value between , which is used to balance the relative importance of the current reward and the future reward, represents the next state transferred to after executing action , is the value estimation of the target network for the next state and all selectable actions , represents the maximum value among all selectable actions in the next state , which reflects the highest expected reward that can be brought by choosing the optimal action in the future state;

[0055] The action cost network learns the potential risks of actions based on the historical trajectory and outputs a high cost value for dangerous actions . By taking the difference between the distribution deviation output by the action reconstruction model and the output by the action cost network as the loss value of the network to update the network;

[0056] Then, based on the evaluation results of the two paths, the policy network calculates the comprehensive reward through weighted calculation, which is used to dynamically adjust the action generation policy. The overall optimization objective is defined as , where is the adjustable risk coefficient, is the entropy regularization coefficient with a value of 0.01 - 0.1. The policy entropy is used to measure the randomness of the policy. A high entropy value indicates that the policy is more random, which can avoid safety risks while increasing the action reward. During the training process, the state-action trajectories randomly sampled from the experience replay pool are constructed into spatio-temporal graph data, which drives the three networks to alternately update parameters to form collaborative optimization. Among them, the policy network updates the parameters by calculating the policy gradient through Monte Carlo.

[0057] Optionally, as Figure 5As shown, the spatio-temporal graph attention mechanism focuses on establishing an adaptive modeling method that can synchronously capture spatial correlation and temporal dynamics. It is composed of multiple spatio-temporal graph attention layers connected by residual connections, and the hidden feature vectors are continuously passed to the deep network;

[0058] Each spatio-temporal graph attention layer, as Figure 6 shown, first models the augmented data of the multi-roll system as a graph structure , where the node represents the attribute vector of the th deviation correction roll, including the angle , associated point , the offset distance between the associated point and the standard edge line and the offset speed . To effectively extract features, all attribute vectors are normalized to the range [0, 1]. These features are all stored in the form of time series, , and the edge represents the relative distance between each deviation correction roll;

[0059] Then, the spatial adjacency matrix is calculated. The calculation process is as follows:

[0060] Through the edge E in the graph structure, the geometric association between discrete nodes is transformed into a matrix form. Taking the belt axis as a one-dimensional coordinate axis, the installation positions of deviation correction roll nodes are extracted. A non-linear attenuation function based on the Gaussian kernel is used to map the physical distance to the adjacency weight in the interval [0, 1]):

[0061]

[0062] where is the adjacency weight between the th deviation correction roll and the th deviation correction roll, is the installation position of the th deviation correction roll and the th deviation correction roll, is the standard deviation of the Gaussian kernel, which controls the neighborhood attenuation rate. To balance the calculation efficiency and accuracy, a distance threshold is introduced to perform a hard truncation on the spatial adjacency matrix. When , is set to 0, and only the adjacency relationships within the effective range are retained, finally generating a highly sparse spatial adjacency matrix ;

[0063] Then, the spatial attention mechanism is calculated. The dynamic calculation of node influence is realized through the cross-dimensional interaction of the triple weight matrix. The calculation process is as follows:

[0064] Utilize time compression weights Perform temporal feature aggregation on the input and then establish cross-dimensional association rules through channel-time interaction weights to capture the mutual influence between the feature channels and the time dimension. Then, with the help of spatial projection weights generate the original attention matrix between nodes , and the formula is: where is a learnable weight vector used to adjust the scaling of the original attention matrix is the Sigmoid activation function is the bias term;

[0065] After the original attention matrix is normalized by row-wise Softmax, the dynamic attention weights conforming to the probability distribution characteristics are output , and each of its elements clearly represents the influence intensity of the deviation rectifying roller on the deviation rectifying roller , and the formula is:

[0066]

[0067] Fuse the dynamic attention weights with the spatial adjacency matrix through the Hadamard product to obtain the enhanced adjacency matrix : , and then construct the spatial attention hidden features. In the r-th spatio-temporal graph attention layer, the hidden features output by the spatial attention are calculated from the hidden features of the previous spatio-temporal graph attention layer, expressed as:

[0068] where the hidden features of the first spatio-temporal graph attention layer are calculated from the node attribute features: ; then calculate the time attention mechanism, and the calculation process is:

[0069] The time attention mechanism captures the dynamic dependencies in the time dimension through parametric modeling, and its core operation is formalized as:

[0070]

[0071] where is a learnable weight vector used to adjust the scaling of the attention matrix is the weight matrix for compressing the node dimension to a scalar is the weight matrix for constructing the association rules between the feature channels and the nodes The weight vector for extracting the feature combination across time steps is the bias term, used to adjust the output of the model;

[0072] The obtained original attention matrix is normalized to generate a probability distribution weight matrix: , and the weight matrix dynamically adjusts the input time-series features through matrix multiplication, enabling the model to strengthen the correlation intensity of key time segments:

[0073]

[0074] Then, the spatio-temporal features are fused through a spatio-temporal feature fusion module, and the process is as follows:

[0075] A graph convolution operation is designed based on the spectral graph theory in the spatial dimension, and conventional convolution is used to extract features in the time dimension. The formula for the spatial graph convolution is: , where is the representation of the feature vector after spatial graph convolution at the th time step, is the node attribute vector at the th time step, is the learnable polynomial coefficient, is the th-order Chebyshev polynomial, recursively defined as , with the initial term , , represents the scaled Laplacian matrix, where is the graph Laplacian matrix, is 's largest eigenvalue, is the identity matrix, is the order of the polynomial expansion, controlling the neighborhood range of node feature aggregation, is the spatial attention weight matrix, adjusting the adjacency relationship through the Hadamard product , and this formula combines polynomial expansion and spatial attention mechanism. Through the th-order polynomial , the node feature aggregation is restricted within the to th-order neighbor range, avoiding the computational burden of directly decomposing the large-scale graph Laplacian matrix. At the same time, the spatial attention matrix is introduced to enhance the dynamic spatial relationship modeling, adjusting the adjacency relationship through the Hadamard product, enabling the feature aggregation process to adapt to the dynamically changing spatial associations;

[0076] The time convolution part performs a convolution operation on the time-series feature , defined as:

[0077]

[0078] wherein is the output feature vector after temporal convolution; represents the temporal convolution kernel parameters, represents the temporal feature of the -th spatio-temporal graph attention layer at the channel number of and time step of The feature vectors at are extracted, and the features at different time steps are summed over all channels and time offsets to aggregate the feature information from different channels and time steps to capture and strengthen the dynamic information, is the corresponding bias term,

[0079] After concatenation in the channel dimension, cross-channel information integration is achieved through convolution:

[0080]

[0081] wherein represents the concatenation operation in the channel dimension, represents the fusion convolution kernel for the linear combination of cross-channel information, is the bias term of the fusion layer, is the standard convolution operation;

[0082] Finally, multiple spatio-temporal graph attention layers optimize the feature extraction ability through residual connections, enhancing the depth and expressive ability of the model.

[0083] Optionally, the augmented data of the multi-roller system is obtained by augmenting the real sampled data through a diffusion model data augmentation method based on collaborative reward driving. The augmentation method guides the data generation process through a collaborative reward mechanism, making the generated data not only diverse but also fully reflect the collaborative characteristics of the multi-roller system, improving the learning efficiency and generalization ability of the multi-roller deviation correction strategy (to solve the problems of scarce training data and lack of collaboration in traditional augmentation methods in multi-agent deep reinforcement learning), as Figure 7 shown, including:

[0084] First, the real sampled multi-agent trajectory data is preprocessed to construct a trajectory sequence in a unified format , including the state vector of each deviation correction roller and the action and the global reward for this section of the trajectory , representing the cumulative reward value obtained for each step of this trajectory. The trajectory sequence is expressed as:

[0085]

[0086] where, represents the number of agents, is the number of time steps;

[0087] The following collaborative reward mechanism is designed to guide the generated data to be collaborative:

[0088] 1) Action coordination reward, aiming to suppress the conflicts between the actions of adjacent deviation rectifying rollers. The action coordination reward is defined as the sum of the squares of the differences in the actions of all adjacent deviation rectifying rollers:

[0089]

[0090] where, is the weight coefficient of the action coordination constraint, and respectively represent the actions of the th and the th deviation rectifying rollers;

[0091] 2) Geometric smoothness reward, aiming to eliminate the sudden change in the curvature of the belt edge. The geometric smoothness reward is defined as the sum of the squares of the second-order differences in the edge positions of the deviation rectifying rollers:

[0092] where, is the weight coefficient of the geometric smoothness constraint, is the reference point at the pixel point of the actual edge of the belt at the same horizontal position value;

[0093] 3) Global cooperation reward, aiming to balance the overall offset and local deviation. The global cooperation reward is defined as the weighted sum of the average value and the maximum deviation value of the horizontal pixel position of the associated point and deviation:

[0094] where, is the weight coefficient of global collaborative optimization, is the weight coefficient of local deviation, is an adjustable coefficient;

[0095] During the training process of the conditional diffusion model, a noise addition function is constructed , and its specific implementation is as follows: In the initial stage of generating data, the model generates a blurred sample by adding noise to the real data. Suppose there is a real trajectory data , which is the starting point of generation. During the noise addition process, the trajectory state is gradually contaminated by noise over time, and this process is modeled in the following form:

[0096]

[0097] In this formula, is the trajectory state generated at the -th step, is the trajectory state generated at the -th step, controls the influence degree of the real trajectory in each step update, and is the added noise, and the scaling factors and ensure the balance between the noise and the real trajectory in the update;

[0098] Based on the said noise addition function , the joint optimization of data reconstruction and collaborative guidance is realized through a dual loss mechanism, where the data distribution reconstruction loss is used to ensure the consistency of the generated data and the real data in distribution, and its definition is:

[0099]

[0100] where, is the original trajectory data, is the noise trajectory data at the -th step in the diffusion process, is the noise sample, is the number of time steps;

[0101] The multi-roll collaboration loss is used to guide the generated data to be collaborative, and its definition is:

[0102] where, is the global value function, equal to the cumulative value of the composite reward , and ;

[0103] The total objective function combines these two losses:

[0104]

[0105] where, is the weight coefficient of the collaboration loss;

[0106] In the denoising sampling stage, the goal of the model is to gradually remove noise from the noise-corrupted trajectory to restore or generate clear trajectory data. The formula is:

[0107] In this formula, is the trajectory state generated after denoising, controls the weight of the noise, reflecting the degree of influence of the noise on the trajectory update, is the previous steps of cumulative product, denoted as , the predicted noise is calculated based on the current trajectory state and the number of steps , is the standard deviation, controlling the amplitude of the added noise, is the noise sampled from the standard normal distribution , the adjustment coefficient controls the degree of influence of the reward gradient on the trajectory update, is the collaborative reward with respect to the current trajectory state gradient.

[0108] In the actual deployment of multi-roll collaborative control, the action strategies of each agent are crucial for the safety and stability of the system. To cope with the dynamic changes and potential risks in the complex environment and ensure the physical feasibility and safety of the agent actions, the embodiment of the present invention designs an action reconstruction model.

[0109] Optionally, as shown in Figure 8 , the action reconstruction model includes a dual-branch encoder and a decoder. The core task of the encoder is to extract the latent features of the input data and parse the dynamic constraint parameters, while the decoder reconstructs the action based on the latent variables generated by the encoder and generates the nearest neighbor safe action within the constraint range through physical rules containing constraint parameters;

[0110] Among them, in the encoder, the state-action pairs sampled from the offline data are used as inputs. After dual-branch processing, the main branch passes through to extract the mean and variance of the latent features, describing the probability distribution of the latent space , is the input, is the encoder; then through latent space sampling, the reparameterization trick is used to generate the latent variable from the distribution defined by the mean and variance:

[0111]

[0112] The constraint branch of the encoder generates dynamic safety parameters through a multi-layer perceptron (MLP), including a margin coefficient , an offset influence factor , a non-linear exponent and a tolerance coefficient ;

[0113] The decoder is responsible for reconstructing the latent variable , generating a safe action, and constraining the action through the dynamic safety parameters, finally generating a physically compliant reconstructed action. The safe physical rules are defined as follows:

[0114] First, the dynamic corner constraint is defined as , where , is the maximum corner of the motor. This constraint means that when a large offset of the belt near the deviation correction roller is detected, the system will relax the corner limit of this deviation correction roller and allow a larger deviation correction action. Second, the rate constraint is defined as , where , is the maximum angular velocity of the motor, is the belt speed vector at the position of each idler. This constraint means that when the overall speed of the belt exceeds the safety threshold , the upper limit of the action rate decreases exponentially. Finally, the coordination constraint is defined as and . This constraint ensures that the ratio of the adjustment amplitudes of adjacent deviation correction rollers does not exceed the ratio of the belt offset amounts with a specific coefficient, achieving the optimal synthesis of deviation correction forces while avoiding mechanical interference;

[0115] The overall optimization objective loss of the action reconstruction model consists of KL divergence, reconstruction loss, and constraint loss;

[0116] Among them, KL divergence directly measures the difference between the decoder output action and the true action to ensure that the model can accurately restore the input data, the latent distribution loss, which calculates the distribution deviation between the latent distribution and the prior distribution through KL divergence. The calculation formula is , where the prior distribution uses the mean and covariance matrix based on offline data to define, . In this way, the generation of the latent space will be closer to the distribution of offline data;

[0117] The reconstruction loss ;

[0118] The constraint loss term quantifies the degree of violation of the safety boundary:

[0119] The KL divergence, reconstruction loss, and constraint loss are weighted and combined into a total distribution constraint loss, and the parameters of the encoder and decoder are optimized with this as the objective. Through this mechanism, while the model improves the data reconstruction accuracy, it also maintains the distribution rationality of the latent space and the physical feasibility of the actions;

[0120] The action reconstruction model determines whether the input action exceeds the normal range through a bimodal index, and defines ; The detection logic uses a dynamic threshold mechanism for judgment: when Score > or there is a single physical constraint violation degree > , it is determined as an abnormal action, where and are preset thresholds. At this time, the safety correction mechanism is triggered to generate a compliant action .

[0121] As Figure 9 shown, the embodiment of the present invention also provides a multi-roller collaborative belt deviation correction control system, and the system includes:

[0122] A detection module 910, configured to detect the current belt edge through visual edge detection technology;

[0123] A construction module 920, configured to compare the current belt edge with the set value of the belt edge to obtain the pixel association points at the intersection of each deviation correction roller and the standard edge line of the belt edge deviation value, including the offset distance and the offset speed , and combine the deviation correction roller angle to construct a state vector , where i is the deviation correction roller number and n is the total number of deviation correction rollers;

[0124] An output module 930, configured to output an action through the policy network of the multi-agent deep reinforcement learning model completed by offline training for the state vector , combine the action with the state vector to form a state-action pair, input the state-action pair into the action reconstruction model completed by offline training, reconstruct the action, and output a safe action for subsequent motor control;

[0125] An evaluation module 940, configured to input the state-action pair into the action value network and action cost network of the multi-agent deep reinforcement learning model completed by offline training, evaluate the action output by the policy network, and online train the policy network, action value network, and action cost network;

[0126] The deviation correction control module 950 is used to process the output safety actions and convert them into electrical signals for input to the servo motor driver. The driver controls the rotation of each deviation correction roller servo motor to perform belt deviation correction control.

[0127] A multi-roller collaborative belt deviation correction control system provided by an embodiment of the present invention has a functional structure corresponding to a multi-roller collaborative belt deviation correction control method provided by an embodiment of the present invention, which will not be elaborated herein.

[0128] Figure 10 FIG. 7 is a schematic structural diagram of an electronic device 1000 provided by an embodiment of the present invention. The electronic device 1000 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 1001 and one or more memories 1002. Among them, at least one instruction is stored in the memory 1002, and the at least one instruction is loaded and executed by the processor 1001 to implement the steps of the above multi-roller collaborative belt deviation correction control method.

[0129] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions. The above instructions can be executed by a processor in a terminal to complete the above multi-roller collaborative belt deviation correction control method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0130] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0131] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-roller coordinated belt deviation correction control method, characterized in that: The method comprises: S1, obtaining the current belt edge through visual edge detection technology; S2, compare the current belt edge with the belt edge setting value to obtain the pixel correlation point where each deviation correction roller intersects with the standard edge line Belt edge deviation values, including offset distance and displacement speed , combined with the angle of the correction roller , construct the state vector , i is the number of the correction roller, n is the total number of correction rollers; S3, the state vector The policy network of the multi-agent deep reinforcement learning model that has been trained offline outputs an action, merges the action with the state vector to form a state-action pair, inputs the action reconstruction model that has been trained offline, reconstructs the action, and outputs a safe action for subsequent motor control; S4. The state-action pair will also input the action value network and action cost network of the multi-agent deep reinforcement learning model that have been trained offline, evaluate the actions output by the strategy network, and train the strategy network, action value network, and action cost network online; S5. The output safety action is processed and converted into an electrical signal and input into the servo motor driver. The driver controls the rotation of the servo motors of each deviation correction roller to perform belt deviation correction control.

2. The method according to claim 1, characterized in that The offline training process of the multi-agent deep reinforcement learning model is as follows: By manually controlling the deviation correction roller motor, multi-dimensional data of the belt during normal operation is collected and recorded. The real-time status characteristics of the deflection roller, including the deflection roller angle , Control the motor rotation angle increment corresponding to the deviation correction roller action , establish the standard edge position of the belt based on historical operating data As a reference, at the standard edge position, the intersection of the standard edge line and the correction roller is used to define the pixel association point on the image for each correction roller. , obtain the actual belt edge position in real time through visual edge detection technology , by calculating the reference point The actual edge of the belt at the same horizontal position Offset distance and displacement speed , build contains The state vector of is the total number of deviation correction rollers, the reward function design level, and defines the instant reward ,in To offset the speed penalty coefficient, this design uses nonlinear weighted fusion of position deviation and dynamic characteristics to finally generate an offline training data set: ; Data contains The complete interaction trajectory of group length T, where Represents multi-roller control action; Then, the reinforcement learning model is trained offline, including offline training of the policy network, action value network and action cost network. The policy network outputs work, the action value network evaluates the long-term impact of the adjustment action on the belt state, and influences the trend of the policy network to generate actions through value output. The action cost network evaluates the safety of the action and suppresses the generation of unsafe actions through cost output. In response to the needs of multi-roller coordinated control of belt conveyor systems, a data augmentation method based on a diffusion model driven by collaborative rewards is used to effectively expand the exploration range of the multi-agent state-action space. The integrated data is put into the experience replay pool for training of the reinforcement learning model. In addition, based on the action reconstruction model, part of the model outputs the distribution deviation and action cost of the action, which are used to update the action cost network, and part of the model outputs the reconstructed action as the actual output safety action.

3. The method according to claim 2, characterized in that The online training process of the multi-agent deep reinforcement learning model is as follows: After the state-action pair is input into the action value network and the action cost network, the loss of the action value network is: the long-term cumulative reward value output by the action value network With immediate rewards after the action Calculate the loss. The loss of the action cost network is: the distribution deviation of the action reconstruction model output and the cost value of the action cost network output Calculate the loss and take the weighted sum of the outputs of the two networks as the combined reward , and state-action pairs are organized into The data is sent into the experience replay pool in the form of data from a dataset, and the strategy network is optimized through small batch incremental training to positively reinforce high-return and low-risk actions and reduce the impact weight of high-risk actions, thereby improving the adaptability and robustness of the system.

4. The method according to claim 3, characterized in that The multi-agent deep reinforcement learning model consists of three networks: an action value network, an action cost network, and a strategy network. The three networks all take a [state, action] trajectory sequence as input, extract corresponding feature vectors through the spatiotemporal graph attention mechanism, and finally output corresponding targets through the fully connected layer. The policy network , output Gaussian distribution mean and variance , , and through action sampling Generate randomly sampled actions, which will be evaluated as input to the action value network and the action cost network; Action value network, which maps the feature vector extracted by the spatiotemporal graph attention mechanism into a long-term cumulative reward value , and based on the Bellman error loss Update the action-value network, where Represents in state Take action Instant rewards for is the discount factor, with a value of Between, used to balance the relative importance of current rewards and future rewards, Indicates execution of an action Then transfer to the next state, is the target network's next state and all optional actions The estimated value of Indicates that in the next state All optional actions The maximum value in reflects the highest expected reward that can be brought by choosing the best action in the future state; The action cost network learns the potential risks of actions based on historical trajectories and outputs high cost values ​​for dangerous actions. , by combining the distribution deviation of the action reconstruction model output and the output of the action cost network Make a difference and update the network as the loss value of the network; Then the policy network obtains the comprehensive reward based on the two evaluation results through weighted calculation. , which is used to dynamically adjust the action generation strategy. The overall optimization goal is defined as ,in is an adjustable risk factor, is the entropy regularization coefficient with a value of 0.01-0.1, and the strategy entropy It is used to measure the randomness of the strategy. A high entropy value indicates a more random strategy, which can improve the action reward while avoiding safety risks. During the training process, the state action trajectory randomly sampled by the experience replay pool is constructed as spatiotemporal graph data, which drives the three networks to alternately update parameters to form a collaborative optimization. Among them, the policy network calculates the policy gradient through Monte Carlo. Implement parameter update.

5. The method according to claim 4, characterized in that The core of the spatiotemporal graph attention mechanism is to establish an adaptive modeling method that can simultaneously capture spatial correlation and temporal dynamics. It consists of multiple spatiotemporal graph attention layers connected by residual connections, and the hidden feature vectors are continuously passed to the deep network. Each spatiotemporal graph attention layer first models the multi-roll system augmented data as a graph structure ,node Indicates The attribute vector of the correction roller, including the angle of the correction roller , association points , the offset distance between the associated point and the standard edge line and displacement speed ,To effectively extract features, all attribute vectors are standardized to the range of [0, 1]. These features are stored in the form of time series. ,side Indicates the relative distance between each deviation correction roller; Then calculate the spatial adjacency matrix, the calculation process is: Through the edge E in the graph structure, the geometric relationship between discrete nodes is converted into a matrix form, and the belt axis is used as the one-dimensional coordinate axis to extract Installation position of the guiding roller node , a nonlinear attenuation function based on a Gaussian kernel is used to map the physical distance to an adjacency weight in the interval [0, 1]): ; in For the correction roller and guide roller The adjacency weight between It is the correction roller and guide roller Installation location, is the standard deviation of the Gaussian kernel, which controls the neighborhood attenuation rate. In order to balance the computational efficiency and accuracy, a distance threshold is introduced. The spatial adjacency matrix is ​​hard truncated when hour, Set to 0, only retain the adjacency relationship within the effective range, and finally generate a highly sparse spatial adjacency matrix ; Then the spatial attention mechanism is calculated to dynamically calculate the node influence through the cross-dimensional interaction of the triple weight matrix. The calculation process is: Using time compression weights For input Aggregate temporal features and then use channel-time interaction weights Establish cross-dimensional association rules to capture the mutual influence between feature channels and time dimensions, and then use spatial projection weights Generate the raw attention matrix between nodes , the formula is: ,in is a learnable weight vector used to adjust the scaling of the original attention matrix, is the Sigmoid activation function, is the bias term; The original attention matrix is ​​processed by directional Softmax normalization to output the dynamic attention weights that conform to the probability distribution characteristics , each element of which Clearly characterize the correction roller in the current space-time state For correcting roller The influence strength is: ; The dynamic attention weights are converted into With the spatial adjacency matrix Fusion, get the enhanced adjacency matrix : , and then construct the spatial attention hidden features. In the rth spatiotemporal graph attention layer, the hidden features of the spatial attention output are calculated from the hidden features of the previous spatiotemporal graph attention layer, expressed as: ; The hidden features of the first spatiotemporal graph attention layer are calculated by node attribute features: ; Then calculate the time attention mechanism, the calculation process is: The temporal attention mechanism captures dynamic dependencies in the temporal dimension through parameterized modeling. Its core operation is formalized as follows: ; in is a learnable weight vector used to adjust the scaling of the attention matrix, The weight matrix used to compress the node dimensions to a scalar, The weight matrix used to construct the association rules between feature channels and nodes, The weight vector used to extract the feature combination across time steps, is a bias term used to adjust the output of the model; The resulting raw attention matrix The probability distribution weight matrix is ​​generated after normalization: , the weight matrix dynamically adjusts the input time series features through matrix multiplication, so that the model can strengthen the correlation strength of key time segments: ; Then the spatiotemporal features are fused through the spatiotemporal feature fusion module. The process is as follows: In the spatial dimension, the graph convolution operation is designed based on the spectral graph theory, and in the temporal dimension, conventional convolution is used to extract features. The formula for the spatial graph convolution is: ,in is the feature vector after spatial graph convolution The representation of time steps, It is The node attribute vector of time steps, are the learnable polynomial coefficients, For the The Chebyshev polynomial of order is recursively defined as , initial term , , represents the scaled Laplacian matrix, where is the graph Laplacian matrix, for The maximum eigenvalue of is the identity matrix, is the order of the polynomial expansion, which controls the neighborhood range of node feature aggregation. is the spatial attention weight matrix, through the Hadamard product Adjust the adjacency relationship. This formula combines polynomial expansion and spatial attention mechanism. Polynomial Limit node feature aggregation to arrive In the range of order neighbors, we avoid the computational burden of directly decomposing the Laplacian matrix of a large-scale graph and introduce the spatial attention matrix Enhance dynamic spatial relationship modeling, adjust the adjacency relationship through Hadamard product, so that the feature aggregation process can adapt to the dynamically changing spatial association; The temporal convolution part is used for time series features Perform a convolution operation, defined as: ; in is the output feature vector after time convolution; represents the temporal convolution kernel parameters, Indicates The temporal features of the spatiotemporal graph attention layer have a channel number of , the time step is The feature vector at , extracts the features at different time steps, sums all channels and time offsets, aggregates the feature information from different channels and time steps to capture and enhance dynamic information, is the corresponding bias term, is a Sigmoid function. This design enables the model to capture the local dynamic pattern of each belt deviation correction roller in the time dimension, complementing the spectral domain operation of spatial graph convolution to jointly build the ability of spatiotemporal feature extraction; After concatenation in the channel dimension, Convolution integrates information across channels: ; in represents the concatenation operation in the channel dimension, Represents the fused convolution kernel, which is used for the linear combination of cross-channel information. is the bias term of the fusion layer, It is a standard convolution operation; Finally, multiple spatiotemporal graph attention layers optimize feature extraction capabilities through residual connections, enhancing the depth and expressiveness of the model.

6. The method according to claim 5, characterized in that The multi-roller system augmented data is obtained by augmenting the real sampled data through a diffusion model data augmentation method driven by collaborative rewards. The augmentation method guides the data generation process through a collaborative reward mechanism, so that the generated data is not only diverse, but also fully reflects the collaborative characteristics of the multi-roller system, and improves the learning efficiency and generalization ability of the multi-roller correction strategy, including: First, the real sampled multi-agent trajectory data Perform preprocessing to construct a trajectory sequence in a unified format , including the state vector of each correction roller ,action And the global reward for this trajectory , represents the cumulative reward value for each step of this trajectory, and the trajectory sequence is expressed as: ; in, represents the number of agents, is the number of time steps; The following collaborative reward mechanism is designed to make data generation collaborative: 1) Motion coordination reward, which aims to suppress the conflict between adjacent deviation correction rollers. Defined as the sum of the squares of the differences in the movements of all adjacent guide rollers: ; in, is the weight coefficient of the action coordination constraint, and Respectively represent and The action of the deviation correction roller; 2) Geometric smoothness bonus, which aims to eliminate sudden changes in the curvature of the belt edge. Geometric smoothness bonus Defined as the sum of squares of the second-order differences of the edge positions of the guide roller: ; in, is the weight coefficient of the geometric smoothness constraint, It is the reference point The pixel points of the actual edge of the belt at the same horizontal position value; 3) Global synergy reward, which aims to balance the overall deviation and local deviation. Defined as the horizontal pixel position of the associated point and The weighted sum of the average and maximum deviation values: ; in, is the weight coefficient of global collaborative optimization, is the weight coefficient of local deviation, is the adjustable coefficient; During the training of the conditional diffusion model, a noise adding function is constructed , which is implemented as follows: In the initial stage of generating data, the model generates a fuzzy sample by adding noise to the real data. Assuming there is a real trajectory data , which is the starting point of generation. During the noise addition process, the trajectory state is gradually contaminated by noise over time. This process is modeled as follows: ; In this formula, It is in The trajectory state generated by the step, It is in The trajectory state generated by the step, Controls the influence of the true trajectory in each update step, and is the added noise, the scaling factor and Ensure that noise and true trajectories are balanced in updates; Based on the noise adding function , the joint optimization of data reconstruction and collaborative guidance is achieved through a dual loss mechanism, where the data distribution reconstruction loss It is used to ensure the consistency of the distribution of generated data and real data, which is defined as: ; in, is the original trajectory data, In the diffusion process Step noise trajectory data, is a noise sample, is the number of time steps; Multi-roller synergy loss The data used to guide the generation of data is collaborative and is defined as: ; in, is the global value function, equal to the compound reward The cumulative value of ; Overall objective function Combining these two losses: ; in, is the weight coefficient of collaborative loss; In the denoising sampling stage, the model aims to extract the trajectories contaminated by noise. The noise is gradually removed in order to restore or generate clear trajectory data. The formula is: ; In this formula, is the trajectory state generated after denoising, The weight of the control noise reflects the influence of the noise on the trajectory update. It is before Steps The cumulative product of , the predicted noise According to the current track status and number of steps Calculated, is the standard deviation, which controls the amplitude of added noise, is the noise sampled from a standard normal distribution , adjustment coefficient Controls the influence of reward gradient on trajectory update, It is a collaborative reward About the current track status gradient.

7. The method according to claim 1, characterized in that The action reconstruction model includes a dual-branch encoder and a decoder. The core task of the encoder is to extract the potential characteristics of the input data and parse the dynamic constraint parameters, while the decoder reconstructs the action based on the latent variables generated by the encoder and generates safe actions within the constraint range through physical rules containing constraint parameters. In the encoder, the state-action pairs sampled from the offline data are used as input, and after dual-branch processing, the main branch passes through Extracting the mean of latent features and variance , describing the latent space The probability distribution of For input, is the encoder; then, by sampling from the latent space, a reparameterization technique is used to generate latent variables from a distribution defined by mean and variance : ; The constraint branch of the encoder generates dynamic safety parameters, including margin coefficients, through a multi-layer perceptron MLP. , offset impact factor , nonlinear index and tolerance factor ; The decoder is responsible for converting the latent variables Reconstruct and generate safe actions, constrain the actions through dynamic safety parameters, and finally generate physically compliant reconstruction actions. Define the safe physical rules as follows: First, the dynamic corner constraint is defined as ,in , is the maximum rotation angle of the motor; secondly, the rate constraint is defined as ,in , is the maximum angular velocity of the motor, is the belt velocity vector at each roller position; finally, the coordination constraint is defined as and ; The overall optimization target loss of the action reconstruction model consists of KL divergence, reconstruction loss and constraint loss; Among them, KL divergence directly measures the difference between the decoder output action and the real action to ensure that the model can accurately restore the input data, potential distribution loss, and calculate the distribution deviation between the potential distribution and the prior distribution through KL divergence. The calculation formula is , where the prior distribution Use the mean based on offline data and the covariance matrix To define, ; Reconstruction loss ; The constraint loss term quantifies the degree of violation of the safety margin: ; The KL divergence, reconstruction loss and constraint loss are weighted together to form a total distribution constraint loss, and the parameters of the encoder and decoder are optimized based on this goal. Through this mechanism, the model improves the data reconstruction accuracy while maintaining the distribution rationality of the latent space and the physical feasibility of the action; The action reconstruction model determines whether the input action exceeds the normal range through the bimodal index and defines ; The detection logic uses a dynamic threshold mechanism to make judgments: when Score > Or there is a single physical constraint violation> When , it is judged as abnormal action, where and is a preset threshold, at which point the security correction mechanism is triggered to generate compliance actions .

8. A multi-roller coordinated belt deviation correction control system, characterized in that: The system comprises: A detection module, used to detect the current belt edge through visual edge detection technology; A construction module is used to compare the current belt edge with the belt edge setting value to obtain the pixel correlation point where each deviation correction roller intersects with the standard edge line Belt edge deviation values, including offset distance and displacement speed , combined with the angle of the correction roller , construct the state vector , i is the number of the correction roller, n is the total number of correction rollers; Output module for the state vector The policy network of the multi-agent deep reinforcement learning model that has been trained offline outputs an action, merges the action with the state vector to form a state-action pair, inputs the action reconstruction model that has been trained offline, reconstructs the action, and outputs a safe action for subsequent motor control; An evaluation module, for inputting the state-action pair into the action value network and the action cost network of the multi-agent deep reinforcement learning model that have been trained offline, evaluating the actions output by the strategy network, and training the strategy network, the action value network, and the action cost network online; The deviation correction control module is used to process the output safety action, convert it into an electrical signal and input it into the servo motor driver. The driver controls the rotation of the servo motors of each deviation correction roller to perform belt deviation correction control.

9. An electronic device, comprising a processor and a memory, wherein at least one instruction is stored in the memory, wherein: The at least one instruction is loaded and executed by the processor to implement the multi-roller cooperative belt deviation correction control method as described in any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the multi-roller cooperative belt deviation correction control method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Closed-loop protection system and method for belt conveyor

    CN118992455A

  • Automatic balance adjusting method for deviation rectifying device of belt conveyor

    CN119389707A

  • Self-adaptive variable-frequency speed regulation device based on scraper conveyor and control method

    CN119568680A

  • Automatic control method and system for multi-mode belt conveyor

    CN119620608A

  • Automatic deviation correction control method and system for chemical belt

    CN119821932A

Cited By

  • Real-time deviation rectifying method, device, equipment, medium and product for end slope mining cave track

    CN121596754A

  • Power grid project multi-specialty collaborative design method based on data fusion

    CN122288647A