Screw ship unloader work order processing method and system based on reinforcement learning

By combining the reinforcement learning model with heuristic search path planning, the operating efficiency and safety issues of the screw unloader in complex environments were solved, and an efficient and safe unmanned operation process was achieved.

CN120725633AActive Publication Date: 2025-09-30SHANGHAI SHIDONGKOU NO 2 POWER PLANT HUANENG INTERNATIONAL POWER CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511256988.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-09-30
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing screw ship unloader systems are difficult to achieve efficient and safe unmanned operations when faced with complex unstructured environments, and lack self-learning and adaptive optimization capabilities, resulting in limited improvements in operating efficiency.

Method used

A reinforcement learning model based on the Actor-Critic architecture is combined with heuristic search path planning to generate a multi-objective optimization operation strategy. Through multi-source information fusion and path optimization, the operation path and action sequence of the screw ship unloader are planned to ensure autonomous intelligent control of the system in complex environments.

Benefits of technology

It significantly improves the operating efficiency, safety and adaptability of the screw ship unloader under unmanned conditions, realizes the automation of the entire process from task reception to precise execution, and reduces operating and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725633A_ABST
    Figure CN120725633A_ABST
Patent Text Reader

Abstract

The invention discloses a spiral ship unloader work order processing method and system based on reinforcement learning, and the method comprises the steps: receiving an external work order of a spiral ship unloader, and generating an executable task instruction; real-time environment information of the working environment of the spiral ship unloader is collected; based on a reinforcement learning model of an Actor-Critic architecture, generating a job strategy according to the task instruction and the real-time environment information; and according to the operation strategy, a heuristic search algorithm is adopted, path optimization, obstacle avoidance and efficiency optimization are taken as priorities, an operation path and an action sequence of the screw ship unloader are planned, and the screw ship unloader is controlled to act according to the operation path and the action sequence. Through combination of reinforcement learning intelligent decision and heuristic search path planning, multi-target collaborative optimization of the whole operation process of the unattended screw ship unloader is realized, and the operation efficiency, the safety and the adaptive capacity are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of screw ship unloader control, and in particular to a method and system for processing a screw ship unloader operation work order based on reinforcement learning. Background Art

[0002] In modern port logistics systems, the efficiency and level of automation in bulk cargo loading and unloading are directly related to the throughput capacity and operating costs of the entire port. As key equipment for bulk cargo unloading operations, the traditional operating mode of screw unloaders relies heavily on the operator's experience for manual control. This mode not only requires extremely high technical proficiency from the operator, requiring them to remain focused at all times to cope with the complex cabin environment and irregular cargo stacking patterns, but also leads to bottlenecks in manual operation efficiency that are difficult to overcome. At the same time, the cabin environment is complex and visibility may be poor. Long-term, high-intensity manual operations are also accompanied by high safety risks, which can easily lead to equipment collisions or personnel safety accidents. This has become a pain point that the industry urgently needs to address.

[0003] With the continuous development of industrial automation and artificial intelligence technologies, the trend towards unmanned and intelligent port equipment is becoming increasingly evident. To improve operational safety and efficiency, some attempts have emerged on the market to introduce automated screw unloader systems. However, these existing systems generally have certain limitations. First, their data sources are often relatively limited, typically relying only on preset programs or limited sensor information, making it difficult to fully and accurately perceive and respond to the dynamic and changing real-world operating environment. Second, their core decision-making algorithms are mostly based on simple pre-programmed rules or traditional control logic, lacking self-learning and adaptive optimization capabilities. When faced with unprecedented cargo stacking configurations, new obstacle layouts, or unexpected situations, these systems often perform poorly, lack flexibility, and may even cause operational interruptions.

[0004] Specifically, the operational planning modules of existing systems lack flexibility, often employing static or computationally limited path planning methods. This makes it difficult to calculate the optimal solution in real time, balancing the shortest path, obstacle avoidance, safety, and energy efficiency in complex, unstructured environments. The decision-making process is often fragmented, failing to integrate environmental perception, intelligent decision-making, motion planning, and safety control into a seamless, integrated loop. This results in limited improvements in overall operational efficiency and fails to fully meet the stringent requirements of modern ports for continuous, efficient, and unmanned operations. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a method and system for processing work orders for screw ship unloader operations based on reinforcement learning. By combining reinforcement learning intelligent decision-making with heuristic search path planning, multi-objective collaborative optimization of the entire process of unmanned screw ship unloader operations is achieved, significantly improving operational efficiency, safety and adaptability.

[0006] To solve the above technical problems, a first aspect of an embodiment of the present invention provides a method for processing a screw ship unloader operation work order based on reinforcement learning, comprising the following steps: Receive external work orders for screw unloaders and generate executable task instructions; Collecting real-time environmental information of the screw ship unloader operating environment; A reinforcement learning model based on an Actor-Critic architecture generates an operation strategy according to the task instructions and the real-time environment information; According to the operation strategy, a heuristic search algorithm is adopted, with path optimization, obstacle avoidance and efficiency optimization as priorities, to plan the operation path and action sequence of the screw ship unloader, and control the screw ship unloader to perform actions according to the operation path and action sequence.

[0007] Furthermore, the reinforcement learning model based on the Actor-Critic architecture generates an operation strategy according to the task instructions and the real-time environment information, including: Fusing the task instructions and the real-time environment information to construct a state vector representing the current state of the system; Inputting the state vector into the policy network of the reinforcement learning model, and obtaining a multi-dimensional action vector through forward propagation calculation of the policy network; The operation strategy is obtained based on the multi-dimensional action vector.

[0008] Furthermore, the task instruction and the real-time environment information are fused and processed to construct a state vector representing the current state of the system, including: Performing structured analysis on the external work order, extracting key parameters such as the target location, cargo type, and work priority, and generating standardized task instruction data; Synchronously collecting the real-time environmental information, which includes cargo pile outlines, obstacle coordinate sets, and environmental temperature and humidity data; aligning the task instruction data with the real-time environment information in time and space so that all data correspond to the system state at the same decision moment; Combining the task instruction data and the real-time environment information into a preliminary feature vector by using a weighted splicing method based on multi-source information fusion; Performing dimension reduction and normalization processing on the preliminary feature vector to obtain the state vector.

[0009] Furthermore, the step of inputting the state vector into the policy network of the reinforcement learning model and obtaining a multi-dimensional action vector through forward propagation calculation of the policy network includes: Inputting the state vector into the input layer of the policy network and generating a standardized state vector through a standardized preprocessing process; Inputting the normalized state vector into the hidden layer, generating deep control features through feature extraction and conversion process; The deep control features are input into the output layer, and an initial multi-dimensional action vector is generated through the action generation process. Each dimensional element of the initial multi-dimensional action vector corresponds to the horizontal movement speed instruction, vertical movement speed instruction, rotation angle instruction and screw mechanism start and stop control signal required to control the screw ship unloader; Inputting the initial multi-dimensional motion vector into an output constraint process to generate a constrained motion vector, wherein the values ​​of each dimension of the constrained motion vector are all within the effective control range of the actuator; The constrained motion vector is input into the actuator adaptation process to generate a final multi-dimensional motion vector, which fully matches the physical limit and response characteristics of the servo motor and the screw mechanism.

[0010] Furthermore, according to the operation strategy, a heuristic search algorithm is used to plan the operation path and action sequence of the screw ship unloader with path optimization, obstacle avoidance and efficiency optimization as priorities, including: Determine the starting operation point and target operation point of the screw ship unloader based on the operation strategy, obtain the static obstacle map and dynamic obstacle prediction information in the current environment, and generate environmental obstacle information; Constructing a composite cost function based on the environmental obstacle information, wherein the composite cost function takes minimizing the total path length as a primary goal, taking safe obstacle avoidance as a hard constraint, and minimizing the operation energy consumption as an optimization goal; A heuristic search algorithm is used to gradually expand the path nodes starting from the starting operation point, the comprehensive cost of each path node is calculated using the composite cost function, the node with the lowest comprehensive cost is selected as the current optimal path point, and a preliminary path is finally generated; Performing smooth optimization on the preliminary path to eliminate unnecessary turning points and speed mutation points in the path, thereby obtaining a continuous and smooth operation path that meets the kinematic constraints of the screw ship unloader; Discretizing the continuous smooth working path into a number of path points, determining motion control instructions for each path point, wherein the motion control instructions include movement speed, steering angle, and screw mechanism operation state, and generating an action sequence that fully matches the working path; Timestamp information is added based on the action sequence, and a complete operation path and action sequence including timestamp information is finally output.

[0011] Furthermore, constructing a composite cost function based on the environmental obstacle information includes: Determine the density and distribution characteristics of obstacles in the current working environment based on the environmental obstacle information, and generate environmental characteristic parameters; Configuring dynamic weight factors based on the environmental characteristic parameters to generate a composite cost function framework with adaptive weight distribution; Integrating the energy consumption assessment model into the composite cost function framework to generate a multi-objective optimization function including the energy consumption prediction dimension; The dynamic weight factor is used to balance the priority relationship among the three optimization objectives of path length, safety constraint and energy consumption prediction to generate the composite cost function.

[0012] Furthermore, the smoothing optimization process of the preliminary path includes: Applying a Bezier curve algorithm to smooth the preliminary path to generate a preliminary smooth path; Performing kinematic feasibility verification on the preliminary smooth path based on the kinematic constraints of the screw ship unloader, and generating a verified path; Performing curvature continuity processing on the verified path to generate a curvature continuous path; Adjusting the moving speed according to the curvature change characteristics of the curvature continuous path to generate a speed optimized path; The speed optimization path is verified to be second-order continuous and differentiable, and the continuous and smooth operation path is finally generated.

[0013] Furthermore, the adding of timestamp information based on the action sequence includes: Based on the dynamic model of the screw unloader, an initial timestamp is assigned to each path point to generate an initial time series. Establishing a time-space mapping relationship for the initial time series to generate a time coordination sequence; Introducing a buffer time mechanism at key action transition points in the time coordination sequence to generate an action sequence with buffer time; Performing time synchronization optimization on the action sequence with buffer time to generate a time synchronization action instruction sequence; The time synchronization action instruction sequence is subjected to final timing verification, and a complete operation path and action sequence with time synchronization marks is output.

[0014] Accordingly, a second aspect of an embodiment of the present invention provides a system for processing a screw ship unloader work order based on reinforcement learning, which processes an external work order based on the above-mentioned method for processing a screw ship unloader work order based on reinforcement learning, including: A work order receiving module is used to receive external work orders for the screw unloader and generate executable task instructions; An information collection module, which is used to collect real-time environmental information of the screw ship unloader's operating environment; A strategy generation module, which is used to generate an operation strategy according to the task instructions and the real-time environment information based on the reinforcement learning model of the Actor-Critic architecture; An action planning module is used to plan the operation path and action sequence of the screw ship unloader according to the operation strategy, using a heuristic search algorithm, with path optimization, obstacle avoidance and efficiency optimization as priorities, and control the screw ship unloader to perform actions according to the operation path and action sequence.

[0015] Accordingly, a third aspect of an embodiment of the present invention provides an electronic device comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the above-mentioned reinforcement learning-based screw unloader operation work order processing method.

[0016] Accordingly, a fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned reinforcement learning-based screw unloader operation work order processing method.

[0017] The above technical solutions of the embodiments of the present invention have the following beneficial technical effects: 1. By introducing a reinforcement learning model based on an actor-critic architecture, the system can autonomously generate optimal operating strategies based on real-time environmental information and task instructions. This model, with its continuous learning and optimization capabilities, can effectively cope with complex, unstructured environments such as irregular cargo stacks and variable obstacle positions within the ship's hold. This overcomes the rigidity and lack of flexibility of traditional pre-programmed systems, significantly improving the screw unloader's intelligent decision-making and environmental adaptability in unmanned operation. 2. By fusing multi-source information to construct a precise state representation and employing a processing flow that incorporates output constraints and actuator adaptation, this ensures that the action instructions output by the reinforcement learning strategy network both meet the optimization objectives and can be safely and accurately executed by the physical equipment. Furthermore, the combination of heuristic search and refined path planning (including smoothing optimization and timing allocation) generates spatial and temporal sequences that fully account for the equipment's dynamic constraints, thereby significantly ensuring the optimality of the operation path, the accuracy of action execution, and the safety and reliability of the entire operation process. 3. Multi-objective optimization, such as path length, obstacle avoidance safety, and operation energy consumption, is integrated into a composite cost function and adaptively balanced through a dynamic weight mechanism. This ensures that the planned operation path and action sequence are not only efficient in space and time, but also economical in energy consumption. This achieves full process automation from task reception to precise execution, reducing efficiency losses and energy waste caused by manual operation uncertainty and suboptimal decision-making, thereby significantly improving overall operation efficiency and reducing comprehensive operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a method for processing a screw ship unloader operation order based on reinforcement learning provided by an embodiment of the present invention; Figure 2 This is a module block diagram of a screw ship unloader operation work order processing system based on reinforcement learning provided by an embodiment of the present invention.

[0019] Reference numerals: 1. Work order receiving module, 2. Information collection module, 3. Strategy generation module, 4. Action planning module. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present invention.

[0021] Please refer to Figure 1 A first aspect of an embodiment of the present invention provides a method for processing a screw ship unloader operation work order based on reinforcement learning, comprising the following steps: Step S100: receiving an external operation work order of a screw ship unloader and generating an executable task instruction.

[0022] It receives work instructions from external systems (such as port management platforms) through secure communication protocols, parses and verifies the received work order data, extracts core work parameters, including target work location, type of cargo to be unloaded, work priority and other key information, and then converts them into standardized, structured, executable task instructions that can be recognized and processed within the system, providing accurate and standardized input basis for subsequent decision-making processes.

[0023] Step S200: collecting real-time environmental information of the screw ship unloader operating environment.

[0024] Multi-source sensors deployed at the work site continuously collect environmental data relevant to the operation. This information includes cargo stack outlines captured by visual sensors, precise coordinates of static and dynamic obstacles acquired by ranging radar, and parameters such as temperature and humidity acquired by environmental sensors. This raw data undergoes a series of preprocessing steps, including filtering, denoising, time synchronization, and coordinate normalization, ultimately forming a complete and consistent real-time environmental information dataset describing the current state of the work environment.

[0025] Step S300: Based on the reinforcement learning model of the Actor-Critic architecture, an operation strategy is generated according to the task instructions and real-time environment information.

[0026] First, the standardized task instructions generated above are combined with preprocessed real-time environmental information to construct a high-dimensional state vector that comprehensively represents the current state of the system. This state vector is then input into a policy network trained using a deep deterministic policy gradient algorithm. After forward propagation, the network outputs a multidimensional continuous action vector, each dimension of which precisely corresponds to the basic instructions used to control the screw unloader's actuators, such as horizontal and vertical movement speeds, rotation angles, and the start and stop status of the screw mechanism. By evaluating and guiding decisions through the value network, an optimal operating strategy is ultimately generated that takes into account both immediate operational benefits and long-term operational safety and efficiency.

[0027] Step S400: Based on the operation strategy, a heuristic search algorithm is used to prioritize path optimization, obstacle avoidance, and efficiency optimization to plan the operation path and action sequence of the screw ship unloader, and the screw ship unloader is controlled to operate according to the operation path and action sequence.

[0028] Based on the operation strategy, a heuristic search algorithm is employed, prioritizing path optimization, obstacle avoidance, and efficiency. The operation path and motion sequence of the screw unloader are planned, and the unloader is controlled to execute according to the operation path and motion sequence. This process first determines the start and target positions of the operation based on the operation strategy, and integrates real-time environmental information to construct an environmental map that includes all obstacles. Subsequently, a heuristic search algorithm is employed, using a composite cost function that incorporates multiple optimization objectives, such as path length, safe obstacle avoidance distance, and estimated energy consumption, to plan a preliminary collision-free path within the configured search space. This path is further smoothed and optimized to eliminate unnecessary turns and sudden speed changes, and rigorously verified to comply with the kinematic constraints of the equipment. Finally, the optimized path is discretized into a series of time-stamped path points, and precise motion control parameters are calculated for each point. This generates a detailed motion sequence that is synchronized in time and space and can directly drive the actuators to complete the unloading operation.

[0029] Through the sequential and tightly coupled technical steps described above, a complete closed-loop process from intelligent decision-making to precise execution has been established. The technical effect is that the system achieves autonomous intelligent control of the entire process of unmanned screw unloader operation in the complex and dynamic port operating environment, significantly improving the automation level of the operation process, the intelligent level of decision-making and planning, the ability to adapt to unstructured environments, and the safety and economic efficiency of the overall operation.

[0030] Specifically, the reinforcement learning model based on the Actor-Critic architecture in step S300 generates an operation strategy based on task instructions and real-time environmental information, including: Step S310 : The task instruction and the real-time environment information are integrated and processed to construct a state vector representing the current state of the system.

[0031] The task instructions and real-time environmental information are fused to construct a state vector representing the current state of the system. This process first normalizes the parsed structured task instruction data to ensure that its numerical range matches the real-time environmental information. Simultaneously, the real-time environmental information collected by sensors undergoes preprocessing, including denoising and coordinate unification, to form a standardized environmental dataset. Subsequently, a spatiotemporal alignment mechanism is employed to ensure that the task instruction requirements and the environmental perception data are fully synchronized in terms of timestamps and spatial reference frames. Furthermore, a multi-source information fusion algorithm based on dynamic weighting factors is used to weightedly concatenate and combine environmental information, including the task target location, cargo type, and priority instructions, with the cargo pile outline, obstacle coordinates, and environmental parameters, to generate a high-dimensional preliminary feature vector. Finally, this feature vector undergoes dimensionality reduction and normalization to remove redundant information and unify its dimensions. Ultimately, a fixed-dimensional state vector is obtained that unambiguously and comprehensively represents the overall operational state of the screw unloader system at the current moment.

[0032] In step S320, the state vector is input into the policy network of the reinforcement learning model, and a multi-dimensional action vector is obtained through forward propagation calculation of the policy network.

[0033] The state vector is input into the policy network of the reinforcement learning model, and a multi-dimensional action vector is obtained through the forward propagation calculation of the policy network. This step takes the state vector constructed above as input and passes it to the policy network using the deep deterministic policy gradient algorithm. The network input layer first performs standardization preprocessing on the state vector to eliminate dimensional differences. The data is then transformed nonlinearly through multiple hidden layers. Each hidden layer is composed of fully connected neurons and uses activation functions to enhance its ability to extract and understand complex state feature relationships. The processed features are finally passed to the output layer. The output layer generates a multi-dimensional continuous action vector based on the continuous control requirements of the screw unloader. Each dimensional element of the vector precisely corresponds to a basic control instruction, including horizontal movement speed, vertical movement speed, rotation angle, and start and stop control signals of the screw mechanism.

[0034] Step S330: Obtaining an operation strategy based on the multi-dimensional action vector.

[0035] The operation strategy is obtained based on the multi-dimensional action vector. This step performs subsequent processing on the initial multi-dimensional action vector output by the strategy network to form an operation strategy that can directly guide planning. First, the output constraint processing is performed on the numerical values ​​of each dimension of the action vector to ensure that all control instruction values ​​are within the effective operating range of each actuator to prevent the instruction from exceeding the limit. Then, the actuator adaptation processing is performed. According to the specific physical response characteristics, delay characteristics and physical limits of the servo motor and screw mechanism, the instruction value is scaled and fine-tuned to ensure that the final action vector generated can be accurately and smoothly executed by the physical system. At this point, the final multi-dimensional action vector constitutes a complete, reliable and executable operation strategy.

[0036] Through the above steps S310 to S330, the transformation from multi-source heterogeneous information to executable control strategies is completed. The technical effect is that, through a rigorous set of data fusion, network reasoning and output adaptation processes, it is ensured that the reinforcement learning decision results can accurately reflect the operation intention and environmental status, while strictly meeting the physical constraints of the actuator, thereby providing an optimized and feasible intelligent decision-making basis for subsequent path planning and action execution, significantly improving the overall decision-making reliability, control accuracy and operation safety of the system.

[0037] Furthermore, in step S310, the task instructions and the real-time environment information are integrated and processed to construct a state vector representing the current state of the system, including: Step S311 , perform structured analysis on the external work order, extract key parameters such as the work target location, cargo type, and work priority, and generate standardized task instruction data.

[0038] External work orders are structured and parsed to extract key parameters such as the target location, cargo type, and job priority, generating standardized task instruction data. This process first parses and analyzes the received raw work order data to identify and extract key operational parameters, including the three-dimensional coordinates of the target space and unloading point, the specific type and attributes of the bulk cargo, and the job priority level. These heterogeneous raw parameters are then cleaned and standardized. For example, the coordinate system is converted to the unloader's base coordinate system, the cargo type is mapped to a predefined category code, and the priority is converted into a numerical quantification indicator. Ultimately, standardized task instruction data with a well-organized structure and standardized numerical specifications is generated, allowing for direct use by subsequent algorithms.

[0039] Step S312: synchronously collect real-time environmental information, which includes cargo pile outline, obstacle coordinate set, and environmental temperature and humidity data.

[0040] Real-time environmental information is collected simultaneously, including the cargo pile outline, obstacle coordinates, and ambient temperature and humidity data. This process coordinates multiple sensing devices deployed at the work site to simultaneously collect multi-dimensional environmental data. Visual sensors (such as 3D cameras or lidar) scan and capture three-dimensional point cloud data of the cargo surface within the hold. After processing, accurate geometric information about the cargo pile outline is obtained. Millimeter-wave radar continuously surveys the work area, outputting the coordinates of static obstacles (such as bulkheads and pillars) and dynamic obstacles (such as personnel and other equipment) relative to the unloader. High-precision digital temperature and humidity sensors monitor ambient temperature and humidity parameters in real time. All of this data is assigned a unified timestamp, laying the foundation for subsequent fusion processing.

[0041] Step S313 , aligning the task instruction data with the real-time environment information in time and space, so that all data correspond to the system state at the same decision moment.

[0042] The task instruction data is aligned with the real-time environmental information in time and space, so that all data corresponds to the system state at the same decision moment. This process is a key preprocessing step before data fusion. Time and space alignment involves two levels: time alignment ensures that all data samples used for decision-making have the same timestamp, reflecting a snapshot of the system at the same moment. This is usually achieved through data caching and synchronous triggering mechanisms; spatial alignment unifies all data into a common coordinate system (usually with the screw unloader itself as a reference). This ensures that the target position in the task instruction, the obstacle coordinates in the environmental information, and the cargo outline have a consistent spatial reference system, enabling accurate relative relationship calculations.

[0043] In step S314 , the task instruction data and the real-time environment information are combined into a preliminary feature vector using a weighted splicing method based on multi-source information fusion.

[0044] A weighted splicing method based on multi-source information fusion is used to combine task instruction data and real-time environmental information into a preliminary feature vector. This process combines standardized task instruction data that has undergone spatiotemporal alignment with preprocessed real-time environmental information at the feature level. Fusion is not a simple data stacking process, but rather assigns dynamic weighting factors to different types of data and then performs weighted splicing to form a high-dimensional preliminary feature vector. The weighting factors are not fixed but are adaptively adjusted based on the core objectives of the current operation phase and the complexity of the environment. For example, the weight of obstacle coordinate information is increased in areas with dense obstacles, and the weight of target position information is increased when approaching the target point, thereby highlighting the most critical information features at the moment.

[0045] Step S315 , performing dimension reduction and normalization processing on the preliminary feature vector to obtain a state vector.

[0046] The initial feature vectors are then subjected to dimensionality reduction and normalization to obtain a state vector. This process then optimizes the fused high-dimensional initial feature vectors to construct the final state vector. Dimensionality reduction (such as principal component analysis (PCA) or other feature selection methods) is used to remove redundant information and highly linearly correlated features from the feature vectors, reducing the data dimension while preserving the salient features of the original data. This improves the computational efficiency of the subsequent reinforcement learning model and mitigates the curse of dimensionality. The reduced feature data is then normalized (such as using Min-Max scaling or Z-Score standardization) to map the values ​​of each dimension to a uniform range (such as [0, 1] or [-1, 1]). This eliminates the large numerical variations caused by different physical dimensions and ensures its suitability as an input to the neural network model. The final output is a state vector of moderate dimensionality, standardized values, and a clear and accurate representation of the overall state of the system at the moment of decision.

[0047] Through the above steps S311 to S315, a complete and efficient multi-source heterogeneous data fusion and state construction process is constructed. Through a series of rigorous data processing steps such as structured analysis, synchronous collection, spatiotemporal alignment, weighted fusion, and dimensionality reduction and normalization, the original work order information and environmental perception information with different meanings and dimensions are successfully converted into a high-quality state representation that can be effectively understood and processed by the reinforcement learning model, providing a solid data foundation for generating accurate and reliable operation strategies, and fundamentally improving the perception accuracy and decision effectiveness of the intelligent decision-making system.

[0048] Furthermore, in step S320, the state vector is input into the policy network of the reinforcement learning model, and a multi-dimensional action vector is obtained through forward propagation calculation of the policy network, including: Step S321: Input the state vector to the input layer of the policy network and generate a standardized state vector through a standardized preprocessing process.

[0049] The state vector is input to the input layer of the policy network, and a standardized state vector is generated through a standardization preprocessing process. This process receives the state vector generated by the previous data fusion step as the initial input of the policy network.

[0050] The input layer first performs standardization preprocessing on the state vector, usually using the Z-Score standardization method, that is, based on the mean and standard deviation of each feature dimension obtained from historical data statistics, the input data is centered and scaled to convert it into a distribution with a mean of zero and a standard deviation of one; this eliminates the adverse effects of differences in physical dimensions and numerical ranges of each feature dimension on network training and inference, improves the stability and convergence speed of network training, and ultimately generates a standardized state vector whose numerical distribution characteristics are more suitable for neural network processing.

[0051] Step S322: input the normalized state vector into the hidden layer, and generate deep control features through feature extraction and conversion process.

[0052] The normalized state vector is input into the hidden layer, where deep control features are generated through a feature extraction and transformation process. This process feeds the normalized state vector into the hidden layer portion of the policy network. The hidden layer typically consists of multiple fully connected layers, each of which performs a linear transformation on its input data (via a weight matrix and bias vector) and applies a nonlinear activation function (such as ReLU). Through this layer-by-layer process, the network is able to gradually extract and combine complex feature patterns in the input state, capturing the deep, nonlinear mapping relationship between state variables and optimal control actions. This transforms the original normalized state vector into a series of highly abstract, semantically rich deep control features that form a high-level representation of the final decision action.

[0053] In step S323, the deep control features are input into the output layer, and an initial multi-dimensional motion vector is generated through the motion generation process. The dimensional elements of the initial multi-dimensional motion vector correspond to the horizontal movement speed instruction, vertical movement speed instruction, rotation angle instruction and screw mechanism start and stop control signal required to control the screw ship unloader.

[0054] The deep control features are input into the output layer, where an initial multidimensional action vector is generated through an action generation process. Each dimensional element of the initial multidimensional action vector corresponds to the horizontal speed command, vertical speed command, rotation angle command, and screw mechanism start / stop control signal required to control the screw unloader. This process transmits the deep control features output by the hidden layer to the output layer of the policy network. The number of neurons in the output layer matches the dimensionality of the action space, as required by the continuous control task of the screw unloader. This layer performs a final linear or nonlinear transformation on the input features, directly generating a multidimensional continuous action vector.

[0055] Each dimension of this vector is predefined to correspond to a specific basic control command: for example, the first dimension output represents the command value for horizontal movement speed, the second dimension represents the command value for vertical movement speed, the third dimension represents the command value for the fuselage rotation angle, and the fourth dimension represents the state control signal for the start and stop of the screw mechanism (usually mapping a continuous value range to discrete start and stop states). This generates the initial multi-dimensional action vector inferred by the policy network.

[0056] In step S324 , the initial multi-dimensional motion vector is input into the output constraint process to generate a constrained motion vector, wherein the values ​​of each dimension of the constrained motion vector are all within the effective control range of the actuator.

[0057] The initial multi-dimensional action vector is input into the output constraint process to generate a constrained action vector. The values ​​of each dimension of the constrained action vector are all within the effective control range of the actuator, and hard constraints are imposed on the initial action vector output by the policy network.

[0058] For each dimension of the motion vector, valid upper and lower limits are set based on the performance specifications and operating limits of the corresponding physical actuator (such as a servo motor or screw drive). For example, the horizontal movement speed command must be limited to the maximum forward and reverse speeds that the motor can provide. This process uses a clipping function to force any values ​​in the initial motion vector that exceed the preset limits of their corresponding dimension to be within these limits, ensuring that all command values ​​are theoretically executable, thereby generating a safe constrained motion vector.

[0059] In step S325 , the constrained motion vector is input into the actuator adaptation process to generate a final multi-dimensional motion vector. The final multi-dimensional motion vector fully matches the physical limit and response characteristics of the servo motor and the screw mechanism.

[0060] The constrained motion vector is input into the actuator adaptation process to generate a final multidimensional motion vector that fully matches the physical limits and response characteristics of the servo motor and screw mechanism. Based on the output constraints, this process further considers the dynamic response characteristics and nonlinear factors of the actuator to fine-tune the instructions. This process fine-tunes and scales the constrained motion vector based on a pre-established actuator response model (which may take into account factors such as response delay, nonlinear dead zone, and minimum control resolution). The goal is to ensure that the generated final control instructions are not only executable by the physical system but also have a higher degree of match and predictability between the instructions and the actual response of the mechanism, thereby improving control accuracy and smoothness, ultimately outputting a final multidimensional motion vector that perfectly matches the real physical system.

[0061] The above steps complete the precise transformation from abstract state perception to specific, executable control instructions. Through the forward reasoning of the policy network combined with a rigorous output post-processing process, it ensures that the decision actions generated by the reinforcement learning agent not only reflect the intelligence based on environmental state optimization, but also strictly meet the physical constraints and actual dynamic response characteristics of all actuators, thereby achieving the best balance between theoretical optimization and practical feasibility, and greatly enhancing the reliability, safety and control performance of the entire intelligent decision-making system deployed in a real physical environment.

[0062] Specifically, in step S400, based on the operation strategy, a heuristic search algorithm is used to prioritize path optimization, obstacle avoidance, and efficiency optimization to plan the operation path and action sequence of the screw ship unloader, including: Step S410: determining the starting operation point and target operation point of the screw ship unloader based on the operation strategy, and obtaining the static obstacle map and dynamic obstacle prediction information in the current environment to generate environmental obstacle information.

[0063] Based on the operation strategy, the screw unloader's starting and target operating points are determined. A static obstacle map and dynamic obstacle prediction information are obtained for the current environment, generating comprehensive environmental obstacle information. The determination of the starting and target operating points is based on the task instructions and the current state of the equipment. Static obstacle maps are typically derived from pre-built environmental models or real-time sensor scan data. Dynamic obstacle prediction information relies on multi-frame sensor data fusion and motion trajectories to identify and estimate obstacle position trends over time, providing an accurate environmental perception foundation for path planning.

[0064] Step S420 , constructing a composite cost function based on the environmental obstacle information, wherein the composite cost function takes minimizing the total path length as the primary goal, taking safe obstacle avoidance as the hard constraint, and taking minimizing the operation energy consumption as the optimization goal.

[0065] Based on the acquired environmental obstacle information, a composite cost function is constructed. This function takes minimizing the total length of the path as the primary optimization goal, while using safe obstacle avoidance as a hard constraint to ensure that the path does not pass through any obstacle areas. On this basis, the minimum operating energy consumption is further incorporated into the optimization goal. By introducing an energy consumption model related to the motion state of the equipment, the system operating costs can be reduced while meeting operational efficiency, thereby improving overall economic efficiency.

[0066] In step S430, a heuristic search algorithm is used to gradually expand the path nodes starting from the starting operation point, the comprehensive cost of each path node is calculated through a composite cost function, and the node with the lowest comprehensive cost is selected as the current optimal path point, and finally a preliminary path is generated.

[0067] Starting from the starting point, a heuristic search algorithm is used to gradually expand the path nodes. A composite cost function is constructed to evaluate the comprehensive cost of each candidate node. The node with the lowest comprehensive cost is selected as the current expansion direction. Progress is made until the target point is reached, ultimately generating a preliminary path. This process takes into account multiple factors, including path length, safety, and energy consumption, to ensure that the generated path is both feasible and superior in complex operating environments.

[0068] Step S440: performing smooth optimization processing on the preliminary path, eliminating unnecessary turning points and speed mutation points in the path, and obtaining a continuous smooth operation path that meets the kinematic constraints of the screw ship unloader.

[0069] The preliminary path is smoothed and optimized by introducing methods such as spline curve interpolation or Bezier curve fitting to eliminate sharp turns and speed mutation points in the path, making the path continuous and smooth and complying with the kinematic constraints of the screw unloader, including the maximum turning radius and acceleration limit, thereby improving the stability and control accuracy of the equipment in actual operation.

[0070] In step S450, the continuous smooth operation path is discretized into a number of path points, and the motion control instructions for each path point are determined. The motion control instructions include the moving speed, steering angle, and the operating state of the screw mechanism, and an action sequence that fully matches the operation path is generated.

[0071] The optimized continuous and smooth operation path is discretized into several path points at certain time or distance intervals, and corresponding motion control instructions are generated for each path point, including movement speed, steering angle, and start, stop, and speed control instructions for the spiral mechanism, thus forming an action sequence that fully matches the path, ensuring that the equipment can perform unloading operations accurately and coherently.

[0072] Step S460 , adding timestamp information based on the action sequence, and finally outputting a complete operation path and action sequence including the timestamp information.

[0073] Based on the equipment's motion capabilities and operational requirements, corresponding timestamp information is assigned to each control instruction in the action sequence to form a complete operation plan with a temporal relationship. Ultimately, the path and action sequence with timestamps are output, providing a clear and schedulable time benchmark and action basis for subsequent execution control.

[0074] The above path and motion planning process significantly improves the operational safety, efficiency and economy of the screw unloader in complex port environments by integrating multi-objective optimization, dynamic obstacle prediction and kinematic constraint processing, while also enhancing the system's adaptability to dynamic environments and decision-making reliability.

[0075] Furthermore, constructing a composite cost function based on environmental obstacle information in step S420 includes: Step S421 : Determine the density and distribution characteristics of obstacles in the current working environment based on the environmental obstacle information, and generate environmental characteristic parameters.

[0076] Based on environmental obstacle information, the current working space is quantitatively analyzed. By calculating the number of obstacles per unit area and their spatial distribution characteristics, environmental characteristic parameters such as obstacle density and distribution characteristics are extracted. This process usually involves rasterizing environmental maps and probabilistic occupancy assessments to numerically characterize the environmental complexity and the difficulty of passage.

[0077] Step S422 : configuring dynamic weight factors based on environmental characteristic parameters to generate a composite cost function framework with adaptive weight distribution.

[0078] Weight factors are dynamically configured based on the acquired environmental characteristic parameters. For example, the weight of safe obstacle avoidance is increased in areas with dense obstacles, while more emphasis is placed on path length or energy consumption indicators in open areas. This constructs a composite cost function framework with environmental adaptability. This framework can flexibly adjust the contribution of different optimization objectives according to the real-time environmental status, enhancing the rationality of the system's decision-making in changing scenarios.

[0079] Step S423 , integrating the energy consumption evaluation model into the composite cost function framework to generate a multi-objective optimization function including the energy consumption prediction dimension.

[0080] The energy consumption assessment model constructed based on the equipment dynamics model and historical operation data is integrated into the composite cost function framework. This model can predict the energy consumption under different action sequences and path forms. While optimizing path length and safety, the composite cost function introduces energy consumption prediction as another important optimization dimension, realizing true multi-objective collaborative optimization.

[0081] Step S424 , balancing the priority relationship among the three optimization objectives of path length, safety constraint, and energy consumption prediction through a dynamic weight factor, and generating a composite cost function.

[0082] By using dynamic weight factors to adjust the trade-offs between the three objectives of shortest path length, safe obstacle avoidance constraints, and lowest energy consumption in real time, a composite cost function that comprehensively reflects actual operational needs is constructed. This function performs multi-dimensional evaluations of candidate nodes during each step of the path search process, thereby guiding the search algorithm to generate high-quality paths that are not only the shortest at the geometric level but also take into account safety and economy.

[0083] The adaptive composite cost function can dynamically coordinate multiple optimization objectives in a complex and changing operating environment, significantly improving the rationality, economy and reliability of path planning. At the same time, it enhances the screw unloader's ability to adapt to different working conditions, providing core decision-making support for efficient, safe and low-consumption autonomous operations.

[0084] Furthermore, the smoothing optimization process of the preliminary path in step S440 includes: Step S441 : Applying a Bezier curve algorithm to smooth the preliminary path to generate a preliminary smoothed path.

[0085] The preliminary path is smoothed by applying the Bezier curve algorithm. The specific process includes: first, identifying the key turning points in the path as the control points of the Bezier curve, and then inserting an appropriate number of intermediate control points between adjacent control points according to the complexity of the path and the required smoothness; by calculating the parametric equation of the Bezier curve, a smooth curve path is generated, ensuring the continuity of the first-order derivative of the path and eliminating the angle mutations and jagged fluctuations in the original path.

[0086] Step S442: performing kinematic feasibility verification on the preliminary smooth path based on the kinematic constraints of the screw ship unloader, and generating a verified path.

[0087] A detailed feasibility verification of the preliminary smooth path is conducted based on the kinematic constraints of the screw ship unloader: a kinematic model is established that includes parameters such as the equipment's minimum turning radius, maximum angular velocity, and maximum acceleration; the radius of curvature of each point on the path is calculated and compared with the equipment's minimum turning radius; path segments that do not meet the kinematic constraints are identified and adjusted by curve refitting or inserting transition segments; the adjusted path is verified to satisfy all kinematic constraints, generating a verified path that ensures the equipment can actually execute it.

[0088] Step S443 : performing curvature continuity processing on the verified path to generate a curvature continuous path.

[0089] The verified path is subjected to curvature continuity processing: curve types with continuously changing curvature, such as clothoid curves or quintic polynomial splines, are used; appropriate transition curve segments are inserted at the detected curvature mutation points to ensure that the curvature change of the entire path is continuous and smooth; numerical calculation methods are used to verify whether the processed path meets the curvature continuity requirements, and a path with continuous curvature characteristics is generated.

[0090] Step S444: adjusting the moving speed according to the curvature variation characteristics of the curvature continuous path to generate a speed optimized path.

[0091] Refined speed adjustments are made based on the curvature variation characteristics of curvature-continuous paths: a curvature-velocity mapping model is established that takes into account the dynamic characteristics of the equipment, including factors such as the maximum centripetal acceleration limit and the drive system performance constraints. The travel speed is reduced on path sections with greater curvature to ensure equipment stability and safety. The travel speed is increased on straight lines or large-radius curves with less curvature to optimize operational efficiency. The physical limitations of acceleration and jerk are also considered to generate a smoothly varying velocity profile to ensure smooth equipment movement.

[0092] In step S445 , a second-order continuous differentiability verification is performed on the speed optimization path to ultimately generate a continuous and smooth operation path.

[0093] The speed optimization path is rigorously verified to be second-order continuous and differentiable: the first-order derivative (velocity) and second-order derivative (acceleration) of the path position function are calculated using numerical differentiation methods; points or sections where the derivatives are discontinuous are detected and identified; these sections are further optimized, such as using higher-order spline interpolation or adjusting speed planning parameters; and finally, an operating path is generated that is continuous and smooth in terms of position, velocity, and acceleration, ensuring that the equipment can track and execute smoothly and accurately.

[0094] The path smoothing optimization process described above not only significantly improves the geometric quality and smoothness of the generated path, but also ensures that the path fully complies with the kinematic and dynamic characteristics of the screw ship unloader, enabling the equipment to complete its operations in a smoother and more efficient manner. It also effectively reduces the impact load and energy consumption of the mechanical system, thereby improving the operating efficiency, control accuracy, and long-term reliability of the entire operating system.

[0095] Furthermore, adding timestamp information based on the action sequence in step S460 includes: Step S461: assigning an initial timestamp to each path point based on the dynamic model of the screw ship unloader to generate an initial time series.

[0096] Based on the screw unloader's dynamic model, an initial timestamp is assigned to each path point to generate an initial time series. This process first calculates the minimum time required to complete each path segment while satisfying the equipment's dynamic constraints, based on the screw unloader's physical and dynamic characteristics. This includes key performance parameters such as the maximum acceleration, maximum deceleration, and maximum sustainable speed of each axis, combined with the spatial distances between adjacent path points in the planned path. Based on this, a theoretical earliest arrival time is calculated for each point in the path sequence, starting at time zero. This generates an initial time series containing the timestamps corresponding to each path point, ensuring the theoretical executable nature of the movement instructions.

[0097] Step S462: Establish a time-space mapping relationship for the initial time sequence to generate a time coordination sequence.

[0098] A time-space mapping relationship is established for the initial time series to generate a time-coordinated sequence. This process fine-tunes the initial time series. Its core is to establish a strict time-space mapping relationship to ensure that the ship unloader's position instructions at any moment accurately match its actual motion capabilities. The adjustment ensures that between any two adjacent timestamps, the ship unloader, moving at the calculated speed, will precisely cover the corresponding spatial distance within the specified time. This avoids any temporal disconnect between instructions and execution, and generates a fully coordinated and synchronized action sequence in time and space.

[0099] Step S463: introducing a buffer time mechanism at the key action transition points in the time coordination sequence to generate an action sequence with buffer time.

[0100] A buffer time mechanism is introduced at key action transition points in a time-coordinated sequence to generate action sequences with buffer time. This process identifies key action transition nodes in the time-coordinated sequence, such as the moment of change in movement direction, the moment of starting or stopping a screw mechanism, and the moment of significant change in speed.

[0101] A short time buffer can be inserted before these key nodes. This buffer time is used to allow the mechanical system to smoothly transition to a new motion state and absorb vibrations and shocks caused by system inertia or minor control errors, thereby significantly improving the smoothness of motion execution and equipment life.

[0102] Step S464: perform time synchronization optimization on the action sequence with buffer time to generate a time synchronization action instruction sequence.

[0103] Time synchronization optimization is performed on action sequences with buffered time to generate a time-synchronized action instruction sequence. This process, from a global system perspective, ensures that the action instructions of all actuators that need to work together (such as the trolley travel motor, the trolley travel motor, and the screw rotation motor) are completely synchronized in time. For example, the instruction for the screw to start rotating is precisely aligned with the instruction for the screw mechanism to move directly above the cargo, avoiding erroneous idling at the unloading point or failure to start after reaching the cargo position. By fine-tuning and aligning the triggering time of each independent action instruction, a time-synchronized action instruction sequence is generated with all sub-actions highly coordinated.

[0104] Step S465 , performing final timing verification on the time synchronization action instruction sequence, and outputting a complete operation path and action sequence with a time synchronization mark.

[0105] Final timing verification is performed on the time-synchronized action instruction sequence, outputting a complete operation path and action sequence with time synchronization markers. This process, as the final quality checkpoint for timing planning, comprehensively verifies the generated time-synchronized action instruction sequence logically and physically. Verification includes checking for time conflicts (e.g., two conflicting actions being scheduled at the same time), whether the timing of all actions strictly meets the dynamic subsystem constraints of all motion axes (e.g., whether acceleration exceeds limits), and whether the total operation time meets expected requirements. Only after all verifications are completed will a complete operation path and action sequence with precise time synchronization markers be output, ready for direct execution by the control system.

[0106] The above steps complete the entire process of assigning a precise time dimension to the spatial path action sequence. The technical effect is that, through a series of rigorous timing planning steps based on initial allocation of dynamic models, space-time mapping coordination, key point buffering, multi-mechanism synchronization, and final verification, the generated control instruction sequence not only has the optimal spatial path, but is also precise, smooth, coordinated, and absolutely feasible in the time dimension, thereby ensuring that the screw unloader can automatically operate in the most efficient, stable, and reliable manner in actual operation, maximizing operational efficiency, equipment safety, and control accuracy.

[0107] Accordingly, please refer to Figure 2 A second aspect of an embodiment of the present invention provides a system for processing a screw ship unloader work order based on reinforcement learning, which processes an external work order based on the above-mentioned screw ship unloader work order processing method based on reinforcement learning, including: The work order receiving module 1 is used to receive the external work order of the screw ship unloader and generate an executable task instruction; Information collection module 2, which is used to collect real-time environmental information of the screw ship unloader operating environment; Strategy generation module 3, which is used to generate operation strategies based on task instructions and real-time environment information using the reinforcement learning model based on the Actor-Critic architecture; The action planning module 4 is used to plan the operation path and action sequence of the screw ship unloader based on the operation strategy and adopt a heuristic search algorithm with path optimization, obstacle avoidance and efficiency optimization as priorities, and control the screw ship unloader to perform actions according to the operation path and action sequence.

[0108] Accordingly, a third aspect of an embodiment of the present invention provides an electronic device comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the above-mentioned reinforcement learning-based screw unloader operation work order processing method.

[0109] Accordingly, a fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned reinforcement learning-based screw unloader operation work order processing method.

[0110] The embodiment of the present invention aims to provide a method and system for processing a screw ship unloader operation work order based on reinforcement learning, which has the following effects: 1. By introducing a reinforcement learning model based on an actor-critic architecture, the system can autonomously generate optimal operating strategies based on real-time environmental information and task instructions. This model, with its continuous learning and optimization capabilities, can effectively cope with complex, unstructured environments such as irregular cargo stacks and variable obstacle positions within the ship's hold. This overcomes the rigidity and lack of flexibility of traditional pre-programmed systems, significantly improving the screw unloader's intelligent decision-making and environmental adaptability in unmanned operation. 2. By fusing multi-source information to construct a precise state representation and employing a processing flow that incorporates output constraints and actuator adaptation, this ensures that the action instructions output by the reinforcement learning strategy network both meet the optimization objectives and can be safely and accurately executed by the physical equipment. Furthermore, the combination of heuristic search and refined path planning (including smoothing optimization and timing allocation) generates spatial and temporal sequences that fully account for the equipment's dynamic constraints, thereby significantly ensuring the optimality of the operation path, the accuracy of action execution, and the safety and reliability of the entire operation process. 3. Multi-objective optimization, such as path length, obstacle avoidance safety, and operation energy consumption, is integrated into a composite cost function and adaptively balanced through a dynamic weight mechanism. This ensures that the planned operation path and action sequence are not only efficient in space and time, but also economical in energy consumption. This achieves full process automation from task reception to precise execution, reducing efficiency losses and energy waste caused by manual operation uncertainty and suboptimal decision-making, thereby significantly improving overall operation efficiency and reducing comprehensive operation and maintenance costs.

[0111] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0112] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0113] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for processing work orders for screw ship unloaders based on reinforcement learning, characterized in that: The steps include: Receive external work orders for screw unloaders and generate executable task instructions; Collecting real-time environmental information of the screw ship unloader operating environment; A reinforcement learning model based on an Actor-Critic architecture generates an operation strategy according to the task instructions and the real-time environment information; According to the operation strategy, a heuristic search algorithm is adopted, with path optimization, obstacle avoidance and efficiency optimization as priorities, to plan the operation path and action sequence of the screw ship unloader, and control the screw ship unloader to perform actions according to the operation path and action sequence.

2. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning model based on the Actor-Critic architecture generates an operation strategy according to the task instructions and the real-time environment information, including: Fusing the task instructions and the real-time environment information to construct a state vector representing the current state of the system; Inputting the state vector into the policy network of the reinforcement learning model, and obtaining a multi-dimensional action vector through forward propagation calculation of the policy network; The operation strategy is obtained based on the multi-dimensional action vector.

3. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 2, characterized in that: The fusing of the task instruction and the real-time environment information to construct a state vector representing the current state of the system includes: Performing structured analysis on the external work order, extracting key parameters such as the target location, cargo type, and work priority, and generating standardized task instruction data; Synchronously collecting the real-time environmental information, which includes cargo pile outlines, obstacle coordinate sets, and environmental temperature and humidity data; aligning the task instruction data with the real-time environment information in time and space so that all data correspond to the system state at the same decision moment; Combining the task instruction data and the real-time environment information into a preliminary feature vector by using a weighted splicing method based on multi-source information fusion; Performing dimension reduction and normalization processing on the preliminary feature vector to obtain the state vector.

4. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 2, characterized in that: Inputting the state vector into the policy network of the reinforcement learning model and obtaining a multi-dimensional action vector through forward propagation calculation of the policy network includes: Inputting the state vector into the input layer of the policy network and generating a standardized state vector through a standardized preprocessing process; Inputting the normalized state vector into the hidden layer, generating deep control features through feature extraction and conversion process; The deep control features are input into the output layer, and an initial multi-dimensional action vector is generated through the action generation process. Each dimensional element of the initial multi-dimensional action vector corresponds to the horizontal movement speed instruction, vertical movement speed instruction, rotation angle instruction and screw mechanism start and stop control signal required to control the screw ship unloader; Inputting the initial multi-dimensional motion vector into an output constraint process to generate a constrained motion vector, wherein the values ​​of each dimension of the constrained motion vector are all within the effective control range of the actuator; The constrained motion vector is input into the actuator adaptation process to generate a final multi-dimensional motion vector, which fully matches the physical limit and response characteristics of the servo motor and the screw mechanism.

5. The method for processing a screw ship unloader operation work order based on reinforcement learning according to any one of claims 1 to 4, characterized in that: According to the operation strategy, a heuristic search algorithm is used to plan the operation path and action sequence of the screw ship unloader with path optimization, obstacle avoidance and efficiency optimization as priorities, including: Determine the starting operation point and target operation point of the screw ship unloader based on the operation strategy, obtain the static obstacle map and dynamic obstacle prediction information in the current environment, and generate environmental obstacle information; Constructing a composite cost function based on the environmental obstacle information, wherein the composite cost function takes minimizing the total path length as a primary goal, taking safe obstacle avoidance as a hard constraint, and minimizing the operation energy consumption as an optimization goal; A heuristic search algorithm is used to gradually expand the path nodes starting from the starting operation point, the comprehensive cost of each path node is calculated using the composite cost function, the node with the lowest comprehensive cost is selected as the current optimal path point, and a preliminary path is finally generated; Performing smooth optimization on the preliminary path to eliminate unnecessary turning points and speed mutation points in the path, thereby obtaining a continuous and smooth operation path that meets the kinematic constraints of the screw ship unloader; Discretizing the continuous smooth working path into a number of path points, determining motion control instructions for each path point, wherein the motion control instructions include movement speed, steering angle, and screw mechanism operation state, and generating an action sequence that fully matches the working path; Timestamp information is added based on the action sequence, and a complete operation path and action sequence including timestamp information is finally output.

6. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 5, characterized in that: The constructing of a composite cost function based on the environmental obstacle information includes: Determine the density and distribution characteristics of obstacles in the current working environment based on the environmental obstacle information, and generate environmental characteristic parameters; Configuring dynamic weight factors based on the environmental characteristic parameters to generate a composite cost function framework with adaptive weight distribution; Integrating the energy consumption assessment model into the composite cost function framework to generate a multi-objective optimization function including the energy consumption prediction dimension; The dynamic weight factor is used to balance the priority relationship among the three optimization objectives of path length, safety constraint and energy consumption prediction to generate the composite cost function.

7. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 5, characterized in that: The smoothing optimization process of the preliminary path includes: Applying a Bezier curve algorithm to smooth the preliminary path to generate a preliminary smooth path; Performing kinematic feasibility verification on the preliminary smooth path based on the kinematic constraints of the screw ship unloader, and generating a verified path; Performing curvature continuity processing on the verified path to generate a curvature continuous path; Adjusting the moving speed according to the curvature change characteristics of the curvature continuous path to generate a speed optimized path; The speed optimization path is verified to be second-order continuous and differentiable, and the continuous and smooth operation path is finally generated.

8. The method for processing screw ship unloader work orders based on reinforcement learning according to claim 5, characterized in that: The adding of timestamp information based on the action sequence includes: Based on the dynamic model of the screw unloader, an initial timestamp is assigned to each path point to generate an initial time series. Establishing a time-space mapping relationship for the initial time series to generate a time coordination sequence; Introducing a buffer time mechanism at key action transition points in the time coordination sequence to generate an action sequence with buffer time; Performing time synchronization optimization on the action sequence with buffer time to generate a time synchronization action instruction sequence; The time synchronization action instruction sequence is subjected to final timing verification, and a complete operation path and action sequence with time synchronization marks is output.

9. A screw unloader work order processing system based on reinforcement learning, characterized in that: The external work order is processed based on the reinforcement learning-based screw ship unloader work order processing method according to any one of claims 1 to 8, comprising: A work order receiving module is used to receive external work orders for the screw unloader and generate executable task instructions; An information collection module, which is used to collect real-time environmental information of the screw ship unloader's operating environment; A strategy generation module, which is used to generate an operation strategy according to the task instructions and the real-time environment information based on the reinforcement learning model of the Actor-Critic architecture; An action planning module is used to plan the operation path and action sequence of the screw ship unloader according to the operation strategy, using a heuristic search algorithm, with path optimization, obstacle avoidance and efficiency optimization as priorities, and control the screw ship unloader to perform actions according to the operation path and action sequence.

10. An electronic device, characterized in that: include: at least one processor; And a memory connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the reinforcement learning-based screw unloader work order processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Container yard double-yard-bridge dynamic cooperative scheduling method

    CN110363380A

  • Salamander robot path tracking hierarchical control method based on reinforcement learning

    CN111552301A

  • Intelligent horizontal transportation system and method for full-automatic container loading and unloading wharf.

    CN113486293A

  • Multi-planning algorithm integrated unmanned driving track planning method and related device

    CN117007066A

  • Off-line reinforcement learning harbor equipment global resource allocation algorithm based on uncertainty perception and generalization enhancement

    CN119168312A