Safety behavior evaluation method and electronic device for remote operation of shore-based container crane

By combining maximum entropy inverse reinforcement learning and spatiotemporal attention mechanisms with a dynamic safety potential field model, the accuracy and reliability issues of remote operation behavior evaluation for quayside container cranes are solved, achieving objective and accurate evaluation of remote operation behavior and improving the practicality and credibility of the evaluation.

CN121823397BActive Publication Date: 2026-05-08SHANGHAI MARITIME UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI MARITIME UNIVERSITY
Filing Date
2026-03-13
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for evaluating the safety behavior of remote operation of quay container cranes rely on human experience, lack a detailed understanding of the spatiotemporal heterogeneity of the operation sequence, and the evaluation model is not sufficiently integrated with the safety boundary constraints of the physical world, resulting in insufficient accuracy and reliability of the evaluation.

Method used

The maximum entropy inverse reinforcement learning framework is used to automatically learn the safety intentions of experts. Combined with the spatiotemporal attention mechanism and dynamic safety potential field model, a multimodal time series dataset is constructed to quantify the spatial risk sensitivity and temporal criticality of operational behavior. The trajectory safety margin index is calculated in real time to generate a comprehensive evaluation result.

Benefits of technology

It enables objective and accurate evaluation of remote operation behavior, improves the accuracy and credibility of the evaluation, can identify high-risk critical stages in complex environments and integrate physical security constraints, thus improving the practicality and reliability of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121823397B_ABST
    Figure CN121823397B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of shore-based container crane remote operation safety behavior evaluation method and electronic equipment, comprising: the operation action sequence of the driver to be evaluated is obtained and the state trajectory of work environment, construct multimodal time series dataset;Safety behavior benchmark model obtained by pre-training is called, and the behavior norm index under the expert strategy of operation trajectory is calculated;Operation action sequence is encoded, and the spatial risk sensitivity and time criticality of different work stages are quantified;Dynamic safety potential field model is constructed, and the repulsive potential energy between spreader and hoisting weight and environmental obstacles is calculated in real time, and trajectory safety margin index is generated;Fusion behavior norm index, spatial risk sensitivity and time criticality, and trajectory safety margin index, generate the safety behavior comprehensive evaluation result of the driver to be evaluated.The present application can automatically learn expert safety intent, have the differentiated evaluation ability to the space-time key characteristics of operation behavior, can effectively fuse physical world safety constraint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart port and crane safety monitoring technology, and in particular to a method and electronic device for evaluating the safety behavior of remote operation of a quayside container crane. Background Technology

[0002] In the construction of smart ports, remote control of quay container cranes has become the mainstream operating mode. Scientific and accurate evaluation of the safety behavior of remote operators is a crucial link in ensuring the safety and efficiency of port operations. Currently, behavior evaluation methods based on deep reinforcement learning are considered an important technological direction in this field. However, existing methods of this type suffer from at least three interrelated technical shortcomings in practical applications, limiting their accuracy and practicality:

[0003] First, the construction of evaluation benchmarks heavily relies on human experience. Existing methods typically require engineers to pre-define the specific weights of various reward and punishment rules (such as collision penalties and speed rewards) to form the reward function for the agent's learning. Because crane operating environments are complex and variable, and safe operation logic involves the global optimization of long-term action strategies, this manually defined, static weight combination cannot accurately and completely depict the implicit and dynamic optimal safety decision-making logic followed by excellent operators. This leads to an inherent deviation between the intrinsic standards of the learned evaluation model and the actual optimal safety behavior.

[0004] Secondly, the evaluation process lacks a nuanced understanding of the spatiotemporal heterogeneity of operational sequences. A complete crane operation cycle comprises multiple stages (e.g., movement, container alignment, and crossing the ship's side), each with drastically different sensitivities to safety risks, and the risk distribution also varies across different areas within the workspace. Existing methods typically treat continuous operational sequences as homogeneous time steps, failing to identify and focus on critical temporal segments and highly sensitive spatial areas that have a decisive impact on overall safety. Consequently, it is difficult to thoroughly assess the appropriateness of operator behavior in specific high-risk contexts.

[0005] Furthermore, the evaluation model is not sufficiently integrated with the safety boundary constraints of the physical world. Purely data-driven models mainly rely on historical data for learning. When faced with extreme or occasional operating conditions not fully covered in the training data, they may make judgments that violate physical laws due to a lack of basic physical common sense (such as the impenetrability of objects and the relationship between speed and braking distance). This causes the evaluation recommendations to fail in boundary conditions, resulting in insufficient generalization ability and reliability. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, device, and electronic equipment for evaluating the safety behavior of remote operation of quay container cranes that can automatically learn expert safety intentions, have the ability to differentiate the spatiotemporal key characteristics of operational behavior, and effectively integrate physical world safety constraints, in order to address the above-mentioned technical problems.

[0007] This invention provides a method for evaluating the safety behavior of remote operation of a quay crane, the method comprising:

[0008] Obtain the sequence of driver actions and the corresponding work environment state trajectory to construct a multimodal time series dataset;

[0009] Based on the maximum entropy inverse reinforcement learning framework, the multimodal time series dataset is used to call the safety behavior benchmark model that has been pre-trained through the historical operation trajectory of skilled drivers, so as to calculate the behavioral normative index of the operation trajectory of the driver to be evaluated under the expert policy.

[0010] The spatiotemporal attention mechanism is used to encode the sequence of operational actions in order to extract and quantify the spatial risk sensitivity and temporal criticality corresponding to different operational stages;

[0011] A dynamic safety potential field model is constructed, and the repulsive potential energy between the lifting equipment and the load and environmental obstacles is calculated in real time based on the working environment state trajectory to generate a trajectory safety margin index.

[0012] By integrating the behavioral normative indicators, the spatial risk sensitivity and time criticality, and the trajectory safety margin indicators, a comprehensive evaluation result of the driver's safety behavior is generated.

[0013] In one embodiment, obtaining the sequence of actions of the driver to be evaluated and the corresponding work environment state trajectory includes:

[0014] The system synchronously collects operation sequence data including handle opening, button status, and duration of action, as well as operational environment status trajectory data including spreader spatial coordinates, lifting speed, trolley displacement, load swing angle, and real-time distance to the nearest obstacle.

[0015] In one embodiment, the training process of the safety behavior benchmark model includes:

[0016] The historical operating trajectories of skilled drivers are abstracted into Markov decision processes, and the underlying reward function of the safety behavior benchmark model is defined as a linear combination of state features, expressed as:

[0017]

[0018] in, For the feature weight vector, It is the feature mapping function for state-action pairs;

[0019] By using maximum likelihood estimation, the expected behavioral features generated by the model are made to align with the expected behavioral features of the historical operating trajectories of the skilled driver, thus obtaining the feature weight vector representing the expert's operating logic. .

[0020] In one embodiment, the calculation of the behavioral normative index of the driver's operation trajectory under the expert strategy is specifically as follows:

[0021] According to the feature weight vector Calculate the driving trajectory of the driver to be evaluated. Log-likelihood probability under expert policy distribution ), which serves as the behavioral normative indicator.

[0022] In one embodiment, encoding the sequence of operational actions using a spatiotemporal attention mechanism includes:

[0023] The sequence of operational actions is encoded using a bidirectional long short-term memory network to extract temporal features;

[0024] A spatiotemporal attention mechanism is coupled to the output layer of the bidirectional long short-term memory network to dynamically calculate the attention weights at different sampling times.

[0025] In one embodiment, the spatiotemporal attention mechanism is used, in the time dimension, to identify critical operational transients that contribute significantly to security risks; and to calculate the attention weights. The methods include:

[0026]

[0027]

[0028] in, Let be the hidden state at time t. , and For learnable parameters, For the corresponding to the first The unnormalized attention score at each moment.

[0029] In one embodiment, the extraction and quantification of spatial risk sensitivity and temporal criticality corresponding to different operational stages includes:

[0030] Based on the attention weight Generate a comprehensive vector that reflects the characteristics of key risk stages. ;

[0031] Based on the distribution of the comprehensive vector and the attention weights in the high-sensitivity phase, the time criticality is calculated, and the risk entropy distribution in different coordinate areas is calculated by dynamically dividing the work area into grids to quantify the spatial risk sensitivity.

[0032] In one embodiment, the construction of the dynamic safety potential field model, which calculates the repulsive potential energy between the lifting device and the suspended load and environmental obstacles in real time based on the working environment state trajectory, includes:

[0033] An exponentially decreasing repulsive field is established with the center of the obstacle as the potential source, the repulsive potential energy being... Represented as:

[0034]

[0035] in, This refers to the real-time distance between the spreader and the obstacle. The normal velocity component of the spreading device as it approaches the obstacle. , , This is the scaling factor. The maximum influence distance of the potential field;

[0036] The trajectory safety margin index is obtained by integrating the repulsive potential energy along the driving trajectory of the driver to be evaluated. Or its function transformation.

[0037] In one embodiment, the process of integrating the behavioral normative indicators, the spatial risk sensitivity and temporal criticality, and the trajectory safety margin indicators to generate a comprehensive safety behavior evaluation result includes:

[0038] Regarding the aforementioned behavioral normative indicators The key stage stability index calculated based on the aforementioned spatial risk sensitivity and time criticality. and the physical safety redundancy index calculated based on the trajectory safety margin index. Perform weighted fusion;

[0039] The comprehensive evaluation results of the safety behaviors It is generated by the following formula:

[0040]

[0041] in, , , The preset fusion weight coefficients are used, and Norm() is the normalization function.

[0042] The present invention also provides a remote operation safety behavior evaluation device for quay cranes, the device comprising:

[0043] The multimodal time series dataset construction module is used to obtain the operation action sequence of the driver to be evaluated and the corresponding work environment state trajectory in order to construct a multimodal time series dataset.

[0044] The behavioral normative index calculation module is used to call the safety behavior benchmark model, which is pre-trained through the historical operation trajectory of skilled drivers, based on the maximum entropy inverse reinforcement learning framework and the multimodal time series dataset, to calculate the behavioral normative index of the operation trajectory of the driver to be evaluated under the expert strategy.

[0045] The operation action sequence encoding module is used to encode the operation action sequence using a spatiotemporal attention mechanism in order to extract and quantify the spatial risk sensitivity and temporal criticality corresponding to different operation stages.

[0046] The trajectory safety margin index generation module is used to construct a dynamic safety potential field model and calculate the repulsive potential energy between the lifting equipment and the load and environmental obstacles in real time based on the working environment state trajectory to generate the trajectory safety margin index.

[0047] The comprehensive safety behavior evaluation result generation module is used to integrate the behavioral normative indicators, the spatial risk sensitivity and time criticality, and the trajectory safety margin indicators to generate a comprehensive safety behavior evaluation result for the driver to be evaluated.

[0048] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the remote operation safety behavior evaluation method for quay container cranes as described above.

[0049] The aforementioned method and electronic equipment for evaluating the safety behavior of remote operation of quay container cranes firstly, by invoking a safety behavior benchmark model based on a maximum entropy inverse reinforcement learning framework, can automatically learn and model the implicit, dynamic global safety decision-making logic from the historical operation trajectories of skilled drivers. This eliminates the reliance on manually preset reward function weights, allowing the generated behavioral normative indicators to objectively reflect the degree of consistency between operation and expert strategy, thus solving the technical problem of the subjective nature of the evaluation benchmark. Secondly, by introducing a spatiotemporal attention mechanism to encode the sequence of operational actions, it automatically identifies and quantifies the spatial risk sensitivity and temporal criticality of different stages in the operation process. This enables differentiated focusing and evaluation of high-risk critical stages and low-risk routine stages in long-term crane operation processes, allowing the evaluation to deeply identify risk behavior patterns in specific operational contexts. Finally, by constructing a dynamic safety potential field model and calculating trajectory safety margin indicators, it embeds fundamental physical safety constraints such as insurmountable limits and distance-speed correlations into the evaluation system in the form of calculable potential energy. This enhances the model's generalization ability and reliability when facing boundaries or extreme conditions not fully covered by training data, compensating for the lack of physical common sense in purely data-driven models. Behavioral normative indicators ensure the objectivity of the evaluation benchmark and the consistency of experts, the spatiotemporal attention mechanism gives the evaluation a fine perception of the heterogeneity of the operation process, and the dynamic safety potential field provides the underlying physical safety guarantee. The three ultimately work together through a fusion mechanism to form a comprehensive evaluation scheme that can automatically learn expert intentions, has the ability to differentiate key stages, and integrates physical safety constraints, thereby improving the accuracy, credibility, and practical value of remote operation safety behavior evaluation as a whole. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating a method for evaluating the safety behavior of remote operation of a quayside container crane, as shown in one embodiment.

[0052] Figure 2 A schematic diagram of a remote operation safety behavior evaluation device for a quayside container crane according to one embodiment;

[0053] Figure 3 This is an internal structural diagram of an electronic device according to one embodiment. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] The following is combined Figures 1-3 This invention describes a method and electronic device for evaluating the safe operation behavior of a quayside container crane during remote operation.

[0056] like Figure 1 As shown, in one embodiment, a method for evaluating the safe operation behavior of a quayside container crane remotely includes the following steps:

[0057] Step S110: Obtain the sequence of the driver's actions to be evaluated and the corresponding work environment state trajectory to construct a multimodal time series dataset.

[0058] The system synchronously collects operation action sequences, including handle opening degree, button status and action duration, and operation environment status trajectory, including the spatial coordinates of the lifting device, lifting speed, trolley displacement, load swing angle and real-time distance to the nearest obstacle, through industrial bus and sensing system. After Kalman filtering and noise reduction, it is constructed into a standardized temporal state space, thus providing an accurate, synchronous and low-noise multimodal data foundation for subsequent evaluation.

[0059] Step S120: Based on the maximum entropy inverse reinforcement learning framework, the safety behavior benchmark model, which has been pre-trained using the multimodal time series dataset and the historical operation trajectory of a skilled driver, is invoked to calculate the behavioral normativity index of the driver's operation trajectory under the expert policy.

[0060] The training process for this safety behavior benchmark model includes:

[0061] The historical operating trajectories of skilled drivers are abstracted into Markov decision processes, and the underlying reward function of the safety behavior benchmark model is defined as a linear combination of state features, expressed as:

[0062]

[0063] in, For the feature weight vector, The feature mapping function for the state-action pair includes features such as path curvature, energy loss, swing angle change rate, and obstacle avoidance minimum gap. Maximum likelihood estimation is used to maximize the log-likelihood of the expert trajectory. Using gradient Optimization is performed to make the expected behavioral features generated by the model consistent with the expected behavioral features of the historical operation trajectories of the skilled driver, and the feature weight vector representing the expert's operation logic is obtained through inversion. This process is based on the principle of maximum entropy, and the probability of the expert trajectory appearing is... ,in For the allocation function. During the calling phase, based on the feature weight vector. Calculate the driving trajectory of the driver to be evaluated. Log-likelihood probability under expert policy distribution This process uses reverse learning to automatically deduce the optimal safety logic from expert data, constructing an objective behavioral benchmark that does not rely on subjective human settings. This effectively overcomes the shortcomings of traditional methods, such as the strong subjectivity in the design of reward and punishment functions and the inability to accurately reflect the complex decision-making logic of experts.

[0064] Step S130: Using a spatiotemporal attention mechanism, the sequence of operational actions is encoded to extract and quantify the spatial risk sensitivity and temporal criticality corresponding to different operational stages.

[0065] Specifically, a bidirectional long short-term memory network is used to encode the sequence of operation actions, extract temporal features, and obtain the hidden state at time t. A spatiotemporal attention mechanism is coupled to the output layer of a bidirectional long short-term memory network to dynamically calculate attention weights at different sampling times. In the time dimension, the spatiotemporal attention mechanism is used to identify critical operational transients (such as those involving container handling or passing over ship rails) that contribute significantly to safety risks; and to calculate attention weights. The methods include:

[0066] First, calculate the unnormalized fraction. ,in, Let be the hidden state at time t. , and These are learnable parameters; then the normalized weights are obtained through the Softmax function. ,in Let be the unnormalized attention score corresponding to the j-th time.

[0067] Then, based on the attention weights Generate a comprehensive vector that reflects the characteristics of key risk stages. Based on the distribution of the comprehensive vector and the attention weights during the highly sensitive phase, the time criticality is calculated. Furthermore, by dynamically dividing the work area into grids, the risk entropy distribution within different coordinate regions is calculated to quantify spatial risk sensitivity. This process utilizes an attention mechanism to non-linearly weight the operation sequence, enabling the model to focus on high-risk moments and sensitive spatial regions. This addresses the shortcomings of existing methods in perceiving key spatiotemporal features of long program sequences and identifying deep-seated potential risk behaviors.

[0068] Step S140: Construct a dynamic safety potential field model and calculate the repulsive potential energy between the lifting device and the load and environmental obstacles in real time based on the working environment state trajectory to generate a trajectory safety margin index.

[0069] Specifically, an exponentially decreasing repulsive field is established with the center of the obstacle as the potential source, and the repulsive potential energy... Represented as:

[0070]

[0071] in, This refers to the real-time distance between the spreader and the obstacle. The normal velocity component of the spreading device as it approaches the obstacle. , , This is the scaling factor. The maximum influence distance of the potential field is used; a risk correction coefficient based on the prediction velocity is introduced. The coverage and intensity gradient of the repulsive field are dynamically adjusted based on the real-time velocity vector of the spreader; the trajectory safety margin index is obtained by integrating the repulsive potential energy along the operating trajectory of the driver to be evaluated. or its function transformation (e.g.) This process abstracts the collision risk in the physical world into potential energy, introducing dynamic boundary constraints that conform to actual physical laws into the data-driven evaluation model, thereby enhancing the model's generalization ability and safety assurance when facing extreme or accidental conditions.

[0072] Step S150: Integrate behavioral normative indicators, spatial risk sensitivity and time criticality, and trajectory safety margin indicators to generate a comprehensive evaluation result of the driver's safety behavior to be evaluated.

[0073] behavioral normative indicators Stability indicators for key stages calculated based on spatial risk sensitivity and time criticality (Measured by high attention weight) The smoothness of the movement within the time period, and the physical safety redundancy index calculated based on the trajectory safety margin index. Weighted fusion; comprehensive evaluation results of safe behaviors It is generated by the following formula: ,in, , , The preset fusion weight coefficients are used, and Norm() is the normalization function. This process integrates quantitative indicators from three dimensions—logical standardization, key stage stability, and physical safety redundancy—to generate a comprehensive, multi-dimensional evaluation score and report. This achieves an automated assessment of the safe behavior of remotely operated drivers that is objective, accurate, and well-interpretable, thereby improving the scientific nature and efficiency of safety management.

[0074] It should be added that, during the training process of the safety behavior benchmark model, the feature mapping function of the state-action pair... Specifically, features include path curvature, energy loss, swing angle change rate, and minimum obstacle avoidance clearance, which together constitute the underlying semantics describing operational safety and efficiency. When using a spatiotemporal attention mechanism for encoding, in addition to calculating attention weights in the time dimension to identify critical operational transients, the method also quantifies spatial risk sensitivity by dynamically gridding the work area and calculating the risk entropy distribution within different coordinate regions, thus enabling targeted assessment of sensitive spaces such as "narrow passages" and "multi-obstacle areas." Regarding stability indicators for key stages... The calculation, in its specific technical essence, is to measure the attention weights. Over a longer period, the smoothness or variance of the driver's action sequence is used to reflect the driver's operational stability during risk-sensitive phases. Furthermore, the functional modules involved in this method can be comprised of a dedicated evaluation device, which includes: an interface module for data acquisition and input, connected to the remote control console PLC bus and sensors; and a module for storing computer instructions and expert weights. The system includes a memory for real-time data; a core processor for running the maximum entropy inverse reinforcement learning algorithm, attention network, and potential field model calculations; and a communication component for transmitting evaluation results. All hardware components are interconnected via an internal bus system, and the processor executes programs stored in memory to access expert weights. It processes the real-time input data and applies it according to the comprehensive evaluation formula. The calculations are performed, and finally, after the completion of a single work cycle, a comprehensive safety behavior evaluation score and a detailed report including the location of high-risk moments are output.

[0075] The following describes the remote operation safety behavior evaluation device for quay container cranes provided by the present invention. The remote operation safety behavior evaluation device for quay container cranes described below can be referred to in correspondence with the remote operation safety behavior evaluation method for crane operators described above.

[0076] like Figure 2 As shown, in one embodiment, a remote operation safety behavior evaluation device for a quayside container crane includes a multimodal time-series dataset construction module 210, a behavior normative index calculation module 220, an operation action sequence encoding module 230, a trajectory safety margin index generation module 240, and a comprehensive safety behavior evaluation result generation module 250.

[0077] The multimodal time series dataset construction module 210 is used to obtain the operation action sequence of the driver to be evaluated and the corresponding work environment state trajectory in order to construct a multimodal time series dataset.

[0078] The behavioral normative index calculation module 220 is used to calculate the behavioral normative index of the driver's operation trajectory under the expert policy by calling the safety behavior benchmark model pre-trained by the historical operation trajectory of skilled drivers based on the maximum entropy inverse reinforcement learning framework and using a multimodal time series dataset.

[0079] The operation action sequence encoding module 230 is used to encode the operation action sequence using a spatiotemporal attention mechanism in order to extract and quantify the spatial risk sensitivity and temporal criticality corresponding to different operation stages.

[0080] The trajectory safety margin index generation module 240 is used to construct a dynamic safety potential field model and calculate the repulsive potential energy between the lifting device and the load and environmental obstacles in real time based on the working environment state trajectory, so as to generate the trajectory safety margin index.

[0081] The comprehensive safety behavior evaluation result generation module 250 is used to integrate behavioral normative indicators, spatial risk sensitivity and time criticality, and trajectory safety margin indicators to generate a comprehensive safety behavior evaluation result for the driver to be evaluated.

[0082] Figure 3 This example illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal. Its internal structure diagram can be as follows: Figure 3 As shown, the electronic device includes a processor, memory, and network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the remote operation safety behavior evaluation method for quay container cranes according to any of the above embodiments.

[0083] Those skilled in the art will understand that Figure 3The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device to which the present invention is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0084] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the remote operation safety behavior evaluation method for quay container cranes according to any of the above embodiments.

[0085] In another aspect, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, it implements the remote operation safety behavior evaluation method for quay cranes according to any of the above embodiments.

[0086] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0087] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0089] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A method for evaluating the safety behavior of remote operation of a quayside container crane, characterized in that, The method includes: Obtain the sequence of driver actions and the corresponding work environment state trajectory to construct a multimodal time series dataset; Based on the maximum entropy inverse reinforcement learning framework, the multimodal time series dataset is used to call the safety behavior benchmark model that has been pre-trained through the historical operation trajectory of skilled drivers, so as to calculate the behavioral normative index of the operation trajectory of the driver to be evaluated under the expert policy. The spatiotemporal attention mechanism is used to encode the sequence of operational actions in order to extract and quantify the spatial risk sensitivity and temporal criticality corresponding to different operational stages; A dynamic safety potential field model is constructed, and the repulsive potential energy between the lifting equipment and the load and environmental obstacles is calculated in real time based on the working environment state trajectory to generate a trajectory safety margin index. By integrating the behavioral normative indicators, the spatial risk sensitivity and time criticality, and the trajectory safety margin indicators, a comprehensive evaluation result of the driver's safety behavior is generated. The training process of the safety behavior benchmark model includes: The historical operating trajectories of skilled drivers are abstracted into Markov decision processes, and the underlying reward function of the safety behavior benchmark model is defined as a linear combination of state features, expressed as: in, For the feature weight vector, It is the feature mapping function for state-action pairs; By using maximum likelihood estimation, the expected behavioral features generated by the model are made to align with the expected behavioral features of the historical operating trajectories of the skilled driver, thus obtaining the feature weight vector representing the expert's operating logic. ; The calculation yields the behavioral normative indicators of the driver's operation trajectory under the expert strategy, specifically: According to the feature weight vector Calculate the driving trajectory of the driver to be evaluated. Log-likelihood probability under expert policy distribution , as the behavioral normative indicator.

2. The method for evaluating the safe operation of a quayside container crane remotely according to claim 1, characterized in that, The process of obtaining the sequence of driver actions and the corresponding work environment state trajectory includes: The system synchronously collects operation sequence data including handle opening, button status, and duration of action, as well as operational environment status trajectory data including spreader spatial coordinates, lifting speed, trolley displacement, load swing angle, and real-time distance to the nearest obstacle.

3. The method for evaluating the safe operation of a quayside container crane remotely according to claim 1, characterized in that, The method of encoding the sequence of operational actions using a spatiotemporal attention mechanism includes: The sequence of operational actions is encoded using a bidirectional long short-term memory network to extract temporal features; A spatiotemporal attention mechanism is coupled into the output layer of the bidirectional long short-term memory network to dynamically calculate the attention weights at different sampling times.

4. The method for evaluating the safe operation of a quayside container crane remotely according to claim 3, characterized in that, The spatiotemporal attention mechanism, in the time dimension, is used to identify critical operational transients that contribute significantly to security risks; and calculates the attention weights. The methods include: in, Let be the hidden state at time t. , and For learnable parameters, For the corresponding to the first The unnormalized attention score at each moment.

5. The method for evaluating the safe operation of a quayside container crane remotely according to claim 3 or 4, characterized in that, The extraction and quantification of spatial risk sensitivity and temporal criticality corresponding to different operational stages includes: Based on the attention weight Generate a comprehensive vector that reflects the characteristics of key risk stages. , Let be the hidden state at time t; Based on the distribution of the comprehensive vector and the attention weights in the high-sensitivity phase, the time criticality is calculated, and the risk entropy distribution in different coordinate areas is calculated by dynamically dividing the work area into grids to quantify the spatial risk sensitivity.

6. The method for evaluating the safe operation of a quayside container crane remotely according to claim 1, characterized in that, The construction of the dynamic safety potential field model, based on the operational environment state trajectory, involves real-time calculation of the repulsive potential energy between the lifting equipment and the suspended load and environmental obstacles, including: An exponentially decreasing repulsive field is established with the center of the obstacle as the potential source, the repulsive potential energy being... Represented as: in, This refers to the real-time distance between the spreader and the obstacle. The normal velocity component of the spreading device as it approaches the obstacle. , , This is the scaling factor. The maximum influence distance of the potential field; The trajectory safety margin index is obtained by integrating the repulsive potential energy along the driving trajectory of the driver to be evaluated. Or its function transformation.

7. The method for evaluating the safe operation of a quayside container crane remotely according to claim 1, characterized in that, The method integrates the behavioral normative indicators, the spatial risk sensitivity and time criticality, and the trajectory safety margin indicators to generate a comprehensive safety behavior evaluation result, including: Regarding the aforementioned behavioral normative indicators The key stage stability index calculated based on the aforementioned spatial risk sensitivity and time criticality. and the physical safety redundancy index calculated based on the trajectory safety margin index. Perform weighted fusion; The comprehensive evaluation results of the safety behaviors It is generated by the following formula: in, , , The preset fusion weight coefficients are used, and Norm() is the normalization function.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method for evaluating the safe operation of a quayside container crane as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Embedded system performance evaluation technical proposal based on interactive Markov chain model detection

    CN101593149A

  • Track learning method based on continuous inverse optimal control

    CN117472056A