Maintenance method and system based on collaborative operation of bulk cargo terminal production equipment

By combining reinforcement learning algorithms and model predictive control frameworks, the optimal maintenance plan is generated and dynamically adjusted, solving the problems of over-maintenance and missed fault detection in the maintenance of production equipment in bulk cargo terminals. This achieves a balance between equipment reliability and operational continuity, reduces maintenance costs, and improves production efficiency.

CN121120033AActive Publication Date: 2025-12-12CCCC FIRST HARBOR ENGINEERING CO LTD +2
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511677237.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2025-12-12
Estimated Expiration
2045-11-17

AI Technical Summary

Technical Problem

In the maintenance of production equipment at bulk cargo terminals, the existing technology leads to over-maintenance and increased operation and maintenance costs due to the regular inspection mode. Post-maintenance may result in missed faults, and it is difficult to achieve dynamic optimal maintenance in multi-equipment collaborative operation scenarios, which affects the continuity of operations and economic losses.

Method used

A reinforcement learning algorithm is used to balance maintenance costs and equipment reliability. Combined with the health dynamics model of production equipment and the constraints of work continuity, the optimal maintenance plan is generated through rolling optimization and dynamically adjusted through a model predictive control framework to update the maintenance strategy in real time.

Benefits of technology

It enables precise maintenance planning in multi-machine collaborative operation scenarios, reduces redundant maintenance investment and fault repair costs, improves equipment availability, ensures continuous and stable operation of the work chain, reduces operating costs and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120033A_ABST
    Figure CN121120033A_ABST
Patent Text Reader

Abstract

The invention discloses a maintenance method and system based on collaborative operation of bulk cargo wharf production equipment, and belongs to the technical field of intelligent operation and maintenance of port bulk cargo wharf production equipment. The method comprises the following steps: acquiring production equipment state data and operation data of a bulk cargo terminal; taking the state data and the operation data as input, and adopting a reinforcement learning algorithm to balance the maintenance cost and the equipment reliability to generate an initial maintenance strategy; inputting the initial maintenance strategy into a model prediction control framework, and obtaining an optimal maintenance plan through rolling optimization by taking a production equipment health dynamics model, an operation continuity constraint and a resource availability constraint as boundary conditions; and executing maintenance work according to the optimal maintenance plan, and feeding back an actual maintenance result to a training module of a reinforcement learning algorithm so as to continuously update an initial maintenance strategy. According to the maintenance method and system based on collaborative operation of the bulk cargo wharf production equipment, the operation and maintenance cost of the full life cycle of the bulk cargo wharf can be reduced, and the availability rate of the equipment can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent operation and maintenance technology of production equipment in port bulk cargo terminals, and particularly relates to a maintenance method and system based on collaborative operation of production equipment in bulk cargo terminals. Background Technology

[0002] Bulk cargo terminal operations rely on a variety of large equipment, such as tippers, belt conveyors, stacker-reclaimers, and ship loaders, operating in series and parallel. A single equipment failure can trigger a system-wide shutdown, resulting in significant economic losses. Currently, mainstream maintenance strategies in the industry have significant flaws: Regular maintenance plans are based on fixed cycles, forcing some equipment to shut down for maintenance even when it is in good health, increasing maintenance costs and the risk of operational interruptions, and easily leading to over-maintenance; reactive maintenance only addresses equipment failures after they occur, leaving room for missed fault detection, and sudden failures can disrupt operational continuity and amplify losses. While some technologies have attempted to introduce single-state monitoring or machine learning algorithms for fault prediction, they have not established a global decision-making loop among constraints on equipment reliability, maintenance costs, and operational continuity, making them unsuitable for continuous multi-equipment collaborative operation scenarios and difficult to achieve dynamically optimal maintenance. Summary of the Invention

[0003] In view of the shortcomings of the related technologies, the purpose of this invention is to provide a maintenance method and system based on the collaborative operation of production equipment in bulk cargo terminals, so as to solve the problems mentioned in the background technology.

[0004] To achieve the above objectives, the present invention provides the following technical solution: A maintenance method based on collaborative operation of production equipment in a bulk cargo terminal includes the following steps: S1. Install sensors on the production equipment at the bulk cargo terminal to obtain real-time status data of the production equipment; S2. Collect and store the operation data corresponding to the production equipment within the same operation cycle; S3. Using status data and operational data as input, a reinforcement learning algorithm is used to weigh maintenance costs against equipment reliability and generate an initial maintenance strategy. S4. Input the initial maintenance strategy into the model predictive control framework, and use the production equipment health dynamics model, operation continuity constraints and resource availability constraints as boundary conditions to obtain the optimal maintenance plan through rolling optimization. S5. Execute maintenance operations according to the optimal maintenance plan and feed the actual maintenance results back to the training module of the reinforcement learning algorithm to continuously update the initial maintenance strategy.

[0005] In some embodiments, in step S3, the reinforcement learning algorithm is a proximal policy optimization algorithm, and the reward function of the reinforcement learning algorithm is:

[0006] in, For the cost of a single maintenance, , These are the weighting coefficients. Downtime due to maintenance This is a reliability penalty item.

[0007] In some embodiments, before the state data is input into the reinforcement learning algorithm, it undergoes drift compensation, missing value imputation, and feature extraction processing in sequence to obtain the health indicators of the production equipment. Health indicators Used to determine the health or failure status of production equipment, among which... For time; when When triggered ,in, This is the failure threshold for production equipment.

[0008] In some embodiments, the objective function of the model predictive control framework is:

[0009] in, As a penalty weight, The initial maintenance strategy for the output of the reinforcement learning algorithm is the first... The corresponding action of the step After the model predictive control framework optimizes the initial maintenance strategy, the first... Step-by-step decision-making action, For the first The maintenance cost of the step The time domain length optimized for rolling.

[0010] In some embodiments, the time-domain length of the rolling optimization The rolling step size ranges from 6 to 48 hours. The interval is 10 to 60 minutes to accommodate the dynamic changes in the health status of bulk cargo terminal production equipment and the need for operational continuity.

[0011] In some embodiments, the status data of the production equipment includes at least one of vibration, temperature, current, acoustic signals, and visual images during the operation of the production equipment.

[0012] In some embodiments, the operational data includes at least one of the following: production equipment load, operational flow rate, operational instruction sequence, berth utilization rate, and yard utilization rate. The operational data is collected through the control system of the docking bulk cargo terminal.

[0013] A maintenance system based on collaborative operation of production equipment in a bulk cargo terminal, applied to the aforementioned maintenance method based on collaborative operation of production equipment in a bulk cargo terminal, the maintenance system comprising: The data acquisition module is used to collect status data and operation data of production equipment. The data processing and transmission module is used to preprocess and transmit status data and operation data; The intelligent decision-making module is used to run reinforcement learning algorithms and model predictive control optimization logic to generate initial maintenance strategies and optimal maintenance plans. The execution feedback module is used to issue maintenance instructions corresponding to the optimal maintenance plan, receive the actual maintenance results, and feed the actual maintenance results back to the intelligent decision-making module.

[0014] In some embodiments, the data acquisition module collects status data through sensors adapted to different types of production equipment at the bulk cargo terminal, and collects operational data by interfacing with the control system of the bulk cargo terminal.

[0015] In some embodiments, the data processing and transmission module is configured as an edge computing unit that is compatible with the data acquisition module and the intelligent decision-making module, supporting drift compensation and feature extraction; the intelligent decision-making module is configured as a computing node with multi-device collaborative computing capabilities; and the execution feedback module is configured as an instruction interaction unit that is compatible with the control interfaces of different production equipment.

[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. The maintenance method and system based on collaborative operation of production equipment in bulk cargo terminals provided by this invention are adapted to the multi-machine collaborative continuous operation scenario of bulk cargo terminals. With the help of dual data acquisition and dynamic optimization logic, maintenance plans can be accurately formulated, reducing redundant maintenance investment and the cost of repairing sudden failures, effectively controlling the operation and maintenance expenses of equipment throughout its entire life cycle, and alleviating the pressure on the overall operating cost of the terminal.

[0017] 2. The maintenance method and system based on collaborative operation of production equipment in bulk cargo terminals provided by the present invention rely on equipment health indicator monitoring and reliability trade-off mechanism to reduce maintenance downtime and failure risk, significantly improve the availability of production equipment, ensure the continuous and stable operation of the terminal operation chain, and further improve the overall operation efficiency and production benefits of bulk cargo terminals. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1This is a flowchart illustrating an embodiment of the maintenance method and system based on collaborative operation of production equipment at a bulk cargo terminal according to the present invention. Figure 2 This is a schematic diagram of the system architecture of an embodiment of the maintenance method and system based on the collaborative operation of production equipment in a bulk cargo terminal according to the present invention; Figure 3 This is a flowchart illustrating the coupling optimization of reinforcement learning algorithm and model predictive control in an embodiment of the maintenance method and system based on collaborative operation of production equipment in a bulk cargo terminal according to the present invention. Figure 4 This is a dynamic maintenance decision timing diagram illustrating an embodiment of the maintenance method and system based on collaborative operation of production equipment at a bulk cargo terminal according to the present invention. Detailed Implementation

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] In the description of this invention, it should be understood that the terms "center", "lateral", "longitudinal", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0021] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0022] Example 1: See appendix Figures 1 to 4 This paper presents an illustrative embodiment of the maintenance method based on the collaborative operation of production equipment in a bulk cargo terminal, as proposed in this invention. The maintenance method based on the collaborative operation of production equipment in a bulk cargo terminal includes the following steps: S1. Install sensors on the production equipment at the bulk cargo terminal to obtain real-time status data of the production equipment; S2. Collect and store the operation data corresponding to the production equipment within the same operation cycle; S3. Using status data and operational data as input, a reinforcement learning algorithm is used to weigh maintenance costs against equipment reliability and generate an initial maintenance strategy. S4. Input the initial maintenance strategy into the model predictive control framework, and use the production equipment health dynamics model, operation continuity constraints and resource availability constraints as boundary conditions to obtain the optimal maintenance plan through rolling optimization. S5. Execute maintenance operations according to the optimal maintenance plan and feed the actual maintenance results back to the training module of the reinforcement learning algorithm to continuously update the initial maintenance strategy.

[0023] In step S1, the sensing device is an intelligent sensing device used for sensing and measuring vibration, temperature, current, acoustics, vision, etc. Each intelligent sensing device is installed on the production equipment such as ship loaders, belt conveyors, unloading trolleys, stacker-reclaimers, and tippers at the bulk cargo terminal as required, so as to obtain the status data of each production equipment in real time. Each intelligent sensing device samples according to a preset sampling frequency, for example, the sampling frequency can be set to ≥1kHz.

[0024] The status data of the production equipment includes at least one of the following: vibration, temperature, current, acoustic signals, and visual images during the operation of the production equipment.

[0025] In step S2, the operational data includes at least one of the following: production equipment load, operational flow rate, operational instruction sequence, berth utilization rate, and yard utilization rate. This operational data is collected through the control system of the bulk cargo terminal. The control system of the bulk cargo terminal can be a programmable logic controller (PLC) or a distributed control system (DCS), collecting operational data from either the PLC or the DCS.

[0026] In step S3, before the state data is input into the reinforcement learning algorithm, the data is preprocessed. Specifically, the state data is sequentially processed by drift compensation, missing value imputation, and feature extraction to obtain the health indicators of the production equipment. Health indicators Used to determine the health or failure status of production equipment, among which... For time; when When triggered ,in, This is the failure threshold for production equipment.

[0027] In this embodiment, the reinforcement learning algorithm in step S3 is the Proximal Policy Optimization (PPO) algorithm. Step S3 specifically involves training the policy network and initially maintaining the policy output based on the Proximal Policy Optimization algorithm. The state data obtained in step S1 (health indicators obtained after preprocessing) is used as the basis for this process. Using the work data collected in step S2 as input, the Proximal Policy Optimization (PPO) algorithm is used to balance maintenance costs and equipment reliability to generate an initial maintenance strategy.

[0028] The specific training process is as follows: S31. Definition of State Space S: The state space contains the core influencing factors for bulk cargo terminal maintenance decisions, specifically: production equipment health indicators. (Obtained from state data through drift compensation, missing value imputation and feature extraction), operation queue length, remaining operation time, real-time weather conditions, berth utilization rate and yard utilization rate, and each state parameter is input into the network after being standardized by 0-1.

[0029] S32. Action Space A is defined as the set of optional operations for maintenance decisions, specifically including four types of actions: immediate maintenance, delayed maintenance, and maintenance of the action space. Maintenance, switching to standby machines, and adjusting the operation rate—each action corresponds to a discrete decision instruction, which can be directly mapped to the operation signals of subsequent execution modules.

[0030] S33. Reward function R design: The reward function is used to quantify the benefit-cost trade-off of maintenance decisions. The formula for the reward function of the reinforcement learning algorithm is:

[0031] in, For the cost of a single maintenance, , These are the weighting coefficients. Downtime due to maintenance This is a reliability penalty item.

[0032] In this embodiment, , It was obtained through optimization using a grid search method based on historical operation and maintenance data of bulk cargo terminals. For example, , .when When triggered ,in, The production equipment failure threshold is determined by combining the equipment manufacturer's recommended value with port failure statistics, and the penalty intensity varies accordingly. and The difference grows exponentially, specifically, ,in The penalty coefficient is... This is to strengthen the avoidance of equipment failure risks.

[0033] S34. Training details of the Proximal Policy Optimization (PPO) algorithm: Specifically, the PPO policy is trained using an Actor-Critic dual-network architecture, where the Actor network outputs the action probability distribution and the Critic network evaluates the state value; during training, the experience replay pool size is set to ≥10. 6 It is used to store historical state-action-reward-next state (s, a, r, s') samples to improve the utilization of training data. When the network is iteratively updated, the clipped surrogate objective function is adopted, the clip parameter is set to 0.2, the initial learning rate is 3e-4, and it decays linearly with the training rounds until the policy converges. The convergence criterion can be that the average reward fluctuation of 500 consecutive iterations is <5%.

[0034] In this embodiment, the model predictive control rolling optimization in step S4 is a fine-tuning of the initial maintenance strategy based on PPO. The initial maintenance strategy output in step S3 is input into the model predictive control (MPC) framework, and combined with the production equipment health dynamics model and operational constraints, the optimal maintenance plan is obtained through online rolling optimization. The specific process is as follows: S41. Initial maintenance plan input: The maintenance action sequence output by the PPO algorithm (e.g., "delay maintenance by 30 minutes" or "adjust work rate to 80%)) is used as the initial reference plan for model predictive control (MPC) to ensure that the optimization direction of model predictive control (MPC) is consistent with the global strategy objective.

[0035] S42. Establish the dynamic model and construct the mapping relationship between the health status of production equipment and maintenance actions and operational disturbances. The model expression is as follows:

[0036] in, For the device health status vector (including (Parameters such as wear of key components and equipment operating time). To maintain the action vector (corresponding to discrete actions in action space A, input in One-Hot encoded form), Disturbances to operational needs (such as fluctuations in operational flow or sudden adjustments to operational instructions). It is a nonlinear state transition function, which is obtained by training with historical production equipment health data and operation data of bulk cargo terminal.

[0037] S43. Optimization Objectives and Constraints: With the goal of minimizing maintenance costs and policy deviations in the future time domain, the objective function of the Model Predictive Control (MPC) framework is:

[0038] in, Specifically, for the penalty weight, The weight for policy bias penalty, and , The initial maintenance strategy for the output of the reinforcement learning algorithm is the first... The corresponding action of the step After the model predictive control framework optimizes the initial maintenance strategy, the first... Step-by-step decision-making action, For the first The maintenance costs of the step (including labor costs, spare parts costs, and downtime loss amortization), The time domain length optimized for rolling.

[0039] Constraints are imposed by inequalities. The characteristics include operational continuity constraints and resource availability constraints. The operational continuity constraint is that the duration of a single downtime caused by maintenance is no greater than the duration of the work interval (calculated from the work plan, usually no greater than 60 minutes); the resource availability constraint is that the spare parts inventory required for maintenance is ≥1 (obtained in real time through the dock spare parts management system) and the number of maintenance personnel is ≥2 (matching the dock maintenance team configuration).

[0040] S44. Online Solving and Parameter Setting: Specifically, a quadratic programming (QP) solver is used to solve the above-mentioned constrained optimization problem online, with the solution cycle consistent with the rolling step size; the time domain length of the rolling optimization... The rolling step size ranges from 6 to 48 hours. The time interval is set from 10 minutes to 60 minutes to adapt to the dynamic changes in the health status of bulk cargo terminal production equipment and the requirements for operational continuity. In this embodiment, a rolling time domain is set. Rolling step size Real-time status data is re-collected every 30 minutes. (Due to factors such as operational disturbances, etc.), the model is updated and the solution is re-solved to achieve dynamic optimization.

[0041] During continuous loading and unloading operations at bulk cargo terminals, dynamic optimization of maintenance timing and minimization of overall maintenance costs are achieved through multi-source real-time data-driven approaches, coupled with reinforcement learning and model predictive control.

[0042] In step S5, continuously updating the initial maintenance strategy specifically involves dynamically updating the parameters of the Proximal Policy Optimization (PPO) algorithm based on the actual maintenance results, thereby achieving adaptive evolution of the strategy. The specific process is as follows: S51. Maintenance work order generation and execution: Based on the maintenance timing in the optimal maintenance plan output in step S4 (such as "execute maintenance after the 3rd rolling step (90 minutes later)"), the system automatically generates a standardized maintenance work order (including equipment number, maintenance type, required spare parts, and executor), and sends it to the bulk cargo terminal computerized maintenance management system (CMMS) through the interface, which is then executed by the CMMS scheduling maintenance personnel.

[0043] S52. Feedback Data Collection and Transmission: After maintenance is completed, actual feedback data is collected through the CMMS system, including: actual maintenance time, actual maintenance cost, and equipment fault repair status (e.g., ...). The data includes the recovery value and the duration of job interruption during maintenance. The above data is packaged in the format of "timestamp + device ID + data type" and sent back to the training module of the reinforcement learning algorithm, namely the PPO policy training server.

[0044] S53, PPO strategy network update, transforms the actual feedback data returned into new experience samples, supplements them to the experience replay pool, and triggers incremental updates of the PPO network: every 1000 new samples accumulated, the Actor-Critic network is iterated and trained once, and the network parameters are updated so that the strategy gradually adapts to the actual operation and maintenance scenarios of bulk cargo terminals (such as fluctuations in operational demand caused by seasonal changes and changes in health status caused by equipment aging), and realizes the closed-loop evolution of maintenance strategies.

[0045] See appendix Figure 3 This paper demonstrates a device maintenance decision-making process based on the coupling of reinforcement learning (specifically PPO) and model predictive control (MPC), which is divided into three core stages. The first stage is the offline training stage, corresponding to step S3. Its core is to use the PPO algorithm to train an initial maintenance strategy by defining the state space S, action space A, and reward function R. This allows the algorithm to learn the basic mapping relationship between maintenance decisions and cost-reliability from historical data / theoretical scenarios, providing an initial reference strategy for subsequent online optimization.

[0046] The second stage is the online coupling stage, corresponding to step S4. Its core is the introduction of Model Predictive Control (MPC), which combines three inputs: the initial maintenance strategy from the offline training stage, real-world scenario constraints (including equipment dynamics models, job continuity constraints, and resource constraints), and maintenance cost targets. Through rolling optimization (i.e., re-optimizing at fixed intervals), the initial maintenance strategy is refined into an optimal maintenance plan that better suits the current situation. This stage adapts the general strategies learned offline to real-time equipment status, job requirements, and resource conditions, improving the accuracy of decision-making.

[0047] The third stage is the rolling execution and closed-loop stage, corresponding to step S5. Its core logic is to issue work orders and execute maintenance actions according to the optimal maintenance plan; then collect actual data (such as maintenance duration, cost, equipment health recovery status, and job interruption duration), and calculate the actual reward (using the reward function R from the offline stage to evaluate the actual decision-making effect); finally, store the "actual state-action-reward" data in the experience replay pool. This stage provides real-world scenario samples for the PPO in the offline training stage, allowing the PPO to iteratively update its strategy based on actual execution results, forming a closed loop of "decision-execution-feedback-re-decision," continuously optimizing the maintenance logic.

[0048] In the above illustrative embodiments, the maintenance method based on collaborative operation of production equipment in bulk cargo terminals is a maintenance method that can integrate real-time production equipment status data, operation data, reinforcement learning algorithms and model predictive control. It can minimize the total life cycle maintenance cost under the hard constraint of equipment reliability, and dynamically adjust the maintenance window under the requirement of continuous operation in bulk cargo terminals to avoid downtime conflicts. It also supports multi-machine collaborative scenarios and ensures operation continuity through task rescheduling and resource scheduling.

[0049] Example 2: See appendix Figures 1 to 4 This paper presents an illustrative embodiment of the maintenance system based on the collaborative operation of bulk cargo terminal production equipment proposed in this invention, applied to the maintenance method based on the collaborative operation of bulk cargo terminal production equipment in Embodiment 1. The maintenance system includes: The data acquisition module is used to collect status data and operation data of production equipment. The data processing and transmission module is used to preprocess and transmit status data and operation data; The intelligent decision-making module is used to run reinforcement learning algorithms and model predictive control optimization logic to generate initial maintenance strategies and optimal maintenance plans. The execution feedback module is used to issue maintenance instructions corresponding to the optimal maintenance plan, receive the actual maintenance results, and feed the actual maintenance results back to the intelligent decision-making module.

[0050] The data acquisition module collects status data through sensors adapted to different types of production equipment at the bulk cargo terminal, and collects operational data by interfacing with the terminal's control system. In this embodiment, the sensors are intelligent sensors, specifically including a triaxial vibration accelerometer, an acoustic emission sensor, a sound generator, a temperature sensor, a tension sensor, a fiber optic FBG, a current sensor, a vision camera, a laser displacement sensor, a pressure sensor, a speed sensor, an tilt sensor, an oil particle counter, and an ultrasonic thickness gauge. Each intelligent sensor is used to collect equipment status data. See Appendix. Figure 2The production equipment and sensing devices constitute the field equipment sensing layer in this embodiment.

[0051] The data processing and transmission module is configured as an edge computing unit that communicates with and is compatible with the data acquisition module and the intelligent decision-making module. It supports drift compensation and feature extraction, specifically serving as an edge computing gateway for data preprocessing and feature extraction. (See appendix) Figure 2 The edge computing gateway constitutes the edge aggregation layer in this embodiment.

[0052] The edge aggregation layer is a crucial intermediary between the field device sensing layer and the cloud-based optimization and decision-making layer. Its role is to perform preliminary processing, storage, and protocol adaptation of the raw data collected from the field, providing support for subsequent cloud-based decision-making. The edge computing gateway is the hardware carrier of the edge aggregation layer, providing local computing power support so that data processing, storage, and protocol conversion can be completed at the edge close to the device, reducing latency and bandwidth consumption during data transmission to the cloud. Data preprocessing includes drift compensation, missing value imputation, and feature extraction. Drift compensation corrects zero-point drift issues caused by long-term sensor operation, ensuring data measurement accuracy; missing value imputation fills in missing values ​​using algorithms when data acquisition is interrupted or lost, ensuring data integrity; feature extraction extracts key features from the raw data (such as frequency characteristics of vibration data and peak characteristics of current data) for subsequent device status analysis (such as health indicators). It provides the foundation for computing. In addition, it features a lightweight cache with 7-day rolling storage. Specifically, it temporarily stores the processed data of the last 7 days at the edge, with older data gradually overwritten by newer data—this is rolling storage. This supports rapid local backtracking of recent data (such as troubleshooting) while avoiding excessive resource consumption on edge devices due to long-term storage; simultaneously, it can temporarily cache data during network outages and upload it to the cloud once the network is restored. The protocol is converted to Modbus-TCP. OPC-UA MQTT can solve the problem of inconsistent communication protocols among field devices, edge gateways, and cloud systems. Field devices often use Modbus-TCP (a traditional industrial bus protocol) and OPC-UA (a cross-platform industrial data transmission standard), while cloud / IoT platforms mostly use MQTT (a lightweight IoT publish-subscribe protocol). Through protocol conversion, systems at different levels (field devices → edge → cloud) can communicate with each other, ensuring smooth data transmission.

[0053] The intelligent decision-making module is configured as a computing node with multi-device collaborative computing capabilities, specifically a cloud server, used to run the Proximal Policy Optimization (PPO) algorithm and Model Predictive Control (MPC) framework optimization.

[0054] The execution feedback module is configured as an instruction interaction unit adapted to the control interface of different production equipment, specifically a Computerized Maintenance Management System (CMMS) interface, used to issue maintenance work orders and receive feedback.

[0055] See appendix Figure 2 The cloud-based optimization decision layer in this embodiment comprises Proximal Policy Optimization (PPO), Model Predictive Control (MPC), and Computerized Maintenance Management System (CMMS) interfaces. Through the collaboration of these interfaces, intelligent maintenance decision-making and closed-loop optimization are achieved. Specifically, the policy network of the Proximal Policy Optimization (PPO) algorithm employs an Actor-Critic dual-network architecture for training the PPO policy. The Actor network outputs the action probability distribution, while the Critic network evaluates the state value. The experience replay pool size is ≥10. 6 It is used to store historical data of historical state-action-reward-next state (s, a, r, s'), which is used to train the PPO algorithm, allowing the model to learn from past maintenance decisions and generate an initial maintenance strategy.

[0056] In model predictive control (MPC) framework optimization, equipment health dynamics model Describes the health status of the equipment Maintenance actions and operational disturbance The dynamic evolution relationship between them, such as how the health status of the equipment changes after maintenance actions are performed, and how fluctuations in work flow affect equipment pressure.

[0057] The QP solver uses the quadratic programming (QP) algorithm to perform online rolling optimization of the objective function of MPC under the premise of satisfying constraints such as job continuity and resource availability, and refines the initial maintenance strategy generated by PPO into an optimal maintenance plan that is more in line with the actual scenario.

[0058] The Computerized Maintenance Management System (CMMS) interface enables a closed loop for work order issuance and feedback. Work order issuance involves translating the optimal maintenance plan optimized by the MPC (Master Maintenance Planning) into specific work orders (e.g., "Maintenance required for a certain piece of equipment in 30 minutes, 2 personnel, 1 spare part"), and issuing them to the field maintenance team. The feedback loop collects the results of actual maintenance (e.g., actual maintenance duration, cost, equipment health recovery level, and maintenance-induced downtime), and feeds this data back to the PPO (Property Management Point) algorithm module. This allows the PPO to update its strategy based on the actual results, forming a closed loop of "decision-making → execution → feedback → re-decision-making" to continuously optimize maintenance strategies.

[0059] In this embodiment, the edge computing gateway communicates with the cloud server via 5G or industrial Wi-Fi with low latency, with a communication latency of <50ms.

[0060] Furthermore, the maintenance method of the present invention can also be implemented by software, or by a combination of software and hardware. Additionally, the program for executing the maintenance method of the present invention can be stored on various computer-readable media and loaded into, for example, a CPU for execution when needed. There are no particular limitations on the computer-readable media; for example, HDDs and CDs can be used. ROM, CD Optical discs such as R, MO, MD, and DVD, IC cards, floppy disks, and semiconductor memories such as MROM, EPROM, EEPROM, and flash memory ROM.

[0061] In the above illustrative embodiment, the maintenance system based on the collaborative operation of production equipment in a bulk cargo terminal acquires status data in real time by deploying intelligent sensing devices on the production equipment, and simultaneously collects operational data. Based on these two types of data, a reinforcement learning algorithm is used to make a global trade-off between maintenance costs and equipment reliability to generate an initial maintenance strategy. Then, model predictive control is introduced to continuously optimize the maintenance plan under the constraint of operational continuity, thereby achieving dynamic optimization of maintenance timing and minimization of overall costs. This overcomes the problems of over-maintenance, missed fault detection, and operational interruption caused by traditional periodic inspections or post-maintenance. It is suitable for continuous operation scenarios with multiple machines working together, significantly reducing the total life-cycle operation and maintenance costs of bulk cargo terminals and improving the availability of production equipment.

[0062] The following example illustrates the relationship between changes in equipment health and the selection of maintenance strategies at different times, using the dynamic maintenance decision-making sequence of the bulk cargo terminal production equipment as an example. (See Appendix) Figure 4 This is an example diagram illustrating the timeline of dynamic maintenance decisions. In this diagram, the coordinate axes include a horizontal axis and a vertical axis. The horizontal axis represents time (0-24 hours), indicating the timeline of maintenance decisions within a day, while the vertical axis implicitly represents equipment health indicators. (Values ​​range from 0-100, with 100 representing complete health and 0 representing complete failure). The core markings include an orange warning line, a red warning line, and a blue broken line. The orange warning line "Warning: 80" indicates that when the equipment health drops below 80, it enters a warning state requiring attention. The red warning line "Failure: 30" indicates that when the equipment health drops below 30, it is approaching or has failed, requiring immediate intervention. The blue broken line represents the dynamic decline of equipment health, starting from near 100 (high initial health) and gradually decreasing over time, reflecting the natural decline in equipment health during operation or wear and tear caused by workload. The colored areas in the graph represent maintenance decisions at different times, and the corresponding maintenance strategies for each time period are marked below the horizontal axis. The colored areas visually show the time window in which the strategies take effect. The system is divided into several zones: Green Zone (2-hour delay): Approximately 6-8 hours later, the decision is to delay maintenance by 2 hours. At this time, the equipment health is just near the warning line, so immediate repair is not recommended. Observation is conducted for 2 hours to avoid unnecessary downtime affecting operations. Yellow Zone (4-hour advance switch to standby): Approximately 10-12 hours later, the decision is to switch to standby 4 hours in advance. It is anticipated that the health of the main equipment will decline rapidly, so the standby equipment will take over the operation in advance to ensure the continuity of the process. Purple Zone (standby): Approximately 12-16 hours later, the decision is to activate the standby equipment. The health of the main equipment continues to decline, and the standby equipment will take over the production tasks to buy time for the main equipment repair. Red Zone (immediate repair): After 16 hours, the decision is to repair immediately. The health of the main equipment is close to the failure line, and emergency repairs must be carried out to restore the equipment and avoid complete failure.

[0063] The graph, through a health degradation line and a time-segmented maintenance strategy, illustrates the core of dynamic maintenance: instead of waiting for equipment failure before repair, it involves real-time monitoring of health indicators. By selecting strategies such as "delayed maintenance, switching to standby, and immediate repair" at different health stages, the reliability of equipment can be guaranteed while ensuring operational continuity, reducing unnecessary downtime, and achieving a balance between "cost, reliability, and operational efficiency" for bulk cargo terminal production equipment.

[0064] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0065] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A maintenance method based on collaborative operation of production equipment in a bulk cargo terminal, characterized in that, Includes the following steps: S1. Install sensing devices on the production equipment at the bulk cargo terminal to obtain the status data of the production equipment in real time; S2. Within the same work cycle, collect and store the work data corresponding to the production equipment; S3. Using the aforementioned status data and operational data as input, a reinforcement learning algorithm is employed to weigh maintenance costs against equipment reliability and generate an initial maintenance strategy. S4. Input the initial maintenance strategy into the model predictive control framework, and use the production equipment health dynamics model, operation continuity constraints and resource availability constraints as boundary conditions to obtain the optimal maintenance plan through rolling optimization; S5. Perform maintenance operations according to the optimal maintenance plan, and feed the actual maintenance results back to the training module of the reinforcement learning algorithm to continuously update the initial maintenance strategy.

2. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to claim 1, characterized in that, In step S3, the reinforcement learning algorithm is a proximal policy optimization algorithm, and the reward function of the reinforcement learning algorithm is: in, For the cost of a single maintenance, , These are the weighting coefficients. Downtime due to maintenance This is a reliability penalty item.

3. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to claim 2, characterized in that, Before being input into the reinforcement learning algorithm, the state data undergoes drift compensation, missing value imputation, and feature extraction processes in sequence to obtain the health indicators of the production equipment. Health indicators Used to determine the health or failure status of production equipment, among which... For time; when When triggered ,in, This is the failure threshold for production equipment.

4. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to claim 1, characterized in that, The objective function of the model predictive control framework is: in, As a penalty weight, The initial maintenance strategy for the output of the reinforcement learning algorithm is the first... The corresponding action of the step After the model predictive control framework optimizes the initial maintenance strategy, the first... Step-by-step decision-making action, For the first The maintenance cost of the step The time domain length optimized for rolling.

5. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to claim 4, characterized in that, The time domain length of the rolling optimization The rolling step size ranges from 6 hours to 48 hours. The interval is 10 to 60 minutes to accommodate the dynamic changes in the health status of bulk cargo terminal production equipment and the need for operational continuity.

6. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to any one of claims 1-5, characterized in that, The status data of the production equipment includes at least one of the following: vibration, temperature, current, acoustic signals, and visual images during the operation of the production equipment.

7. The maintenance method based on collaborative operation of production equipment in a bulk cargo terminal according to any one of claims 1-5, characterized in that, The operational data includes at least one of the following: load of the production equipment, operational flow rate, operational instruction sequence, berth utilization rate, and yard utilization rate. The operational data is collected through the control system of the bulk cargo terminal.

8. A maintenance system based on collaborative operation of production equipment in a bulk cargo terminal, applied to the maintenance method based on collaborative operation of production equipment in a bulk cargo terminal as described in any one of claims 1-7, characterized in that, The maintenance system includes: The data acquisition module is used to collect status data and operation data of the production equipment; A data processing and transmission module is used to preprocess and transmit the status data and job data; The intelligent decision-making module is used to run reinforcement learning algorithms and model predictive control optimization logic to generate initial maintenance strategies and optimal maintenance plans. The execution feedback module is used to issue maintenance instructions corresponding to the optimal maintenance plan, receive actual maintenance results, and feed the actual maintenance results back to the intelligent decision-making module.

9. The maintenance system based on collaborative operation of production equipment in a bulk cargo terminal according to claim 8, characterized in that, The data acquisition module collects status data through sensing devices adapted to different types of production equipment at the bulk cargo terminal, and collects operational data by connecting to the control system of the bulk cargo terminal.

10. The maintenance system based on collaborative operation of production equipment in a bulk cargo terminal according to claim 8, characterized in that, The data processing and transmission module is configured as an edge computing unit that is compatible with the data acquisition module and the intelligent decision-making module, and supports drift compensation and feature extraction. The intelligent decision-making module is configured as a computing node with multi-device collaborative computing capabilities; the execution feedback module is configured as an instruction interaction unit adapted to the control interfaces of different production equipment.

Citation Information

Patent Citations

  • Port comprehensive energy system and bulk cargo wharf distribution cooperative scheduling method

    CN117669924A

  • Comprehensive energy system model prediction control method based on deep reinforcement learning

    CN119443674A

  • A dynamic optimization method for industrial automation system based on 5G private network

    CN119743772A

  • Wharf operation resource dynamic allocation method based on multi-objective optimization algorithm

    CN120355146A

  • Wind turbine generator predictive maintenance decision-making method and system applying reinforcement learning

    CN120563110A