Container operation behavior recognition and analysis method and system based on AI
By employing an AI-based method for recognizing and analyzing container operations, and utilizing video data and multimodal feature fusion technology, the problem of insufficient accuracy in container operation recognition in existing technologies has been solved. This enables efficient behavior monitoring and safety early warning, thereby improving the intelligence and safety level of container operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 郑州综合交通运输研究院有限公司
- Filing Date
- 2025-10-10
- Publication Date
- 2026-05-08
AI Technical Summary
Existing container operation management systems rely on manual monitoring, making it difficult to achieve real-time analysis and intelligent identification of complex operational behaviors. In particular, they are prone to misjudgment or omission in highly dynamic scenarios, failing to meet the needs of intelligent and safe operations in inland hubs.
An AI-based method for container operation behavior recognition and analysis is adopted. Action data is extracted from video frame sequences to generate an initial behavior sequence and further subdivide it into action segments. By combining temporal feature analysis and multimodal data fusion, feature weighting technology and dynamic optimization algorithm are used to generate progressive recognition results and iteratively update them to improve recognition accuracy and stability.
Real-time monitoring and risk warning of operator behavior were achieved in the complex and ever-changing container operation environment, which improved the management safety and efficiency of container operations and enhanced the robustness and adaptability of the system.
Smart Images

Figure CN121148018B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and artificial intelligence technology, and in particular to an AI-based method and system for recognizing and analyzing container operation behavior, which can be widely used in the fields of container operation behavior analysis and intelligent monitoring in ports and logistics hubs. Background Technology
[0002] With the rapid development of global logistics and international land transportation, container operations have become one of the most critical links in international trade. Container operations not only include hoisting, handling, and stacking, but also involve complex personnel collaboration and equipment interaction. These operational processes place increasingly higher demands on the actions, behaviors, operating procedures, and working environment of personnel. However, current inland hub container operation management systems mainly rely on manual monitoring, making it difficult to achieve real-time analysis and intelligent identification of complex operational behaviors.
[0003] Currently, the monitoring and management of container operations at inland hubs mainly rely on manual observation and experience-based judgment, which suffers from problems such as poor real-time performance, insufficient recognition accuracy, and limited ability to issue early warnings for abnormal behavior. While existing behavior recognition technologies can achieve certain recognition results in static or single-scene conditions, they often struggle to accurately capture subtle movements and behavioral changes of operators in complex and ever-changing container operation environments. This is especially true in highly dynamic scenarios, where misjudgments or omissions are prone to occur. Furthermore, they lack the ability to dynamically optimize recognition models, failing to meet the needs of intelligent and safe operations at inland hubs.
[0004] Therefore, this application proposes an AI-based method and system for identifying and analyzing container operation behavior. By utilizing computer vision and deep learning technologies, combined with temporal feature analysis and multimodal data fusion, it can accurately and efficiently identify various operational behaviors in container operations under different lighting conditions, environments, and complex operating scenarios, thereby improving the safety level and operational efficiency of hub operations and promoting the development of hub operations towards intelligence and automation. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides an AI-based method and system for identifying and analyzing container operation behaviors. This system analyzes and optimizes the identification results of various operational behaviors during container operations in real time, providing a more intelligent safety monitoring and operation optimization solution for container operations.
[0006] Firstly, this application provides an AI-based method for recognizing and analyzing container operation behavior, the method comprising:
[0007] Step S1: Obtain the initial behavior sequence of the target object in the container operation scenario, and obtain the action fragment of the target object based on the initial behavior sequence;
[0008] Step S2: Based on the action segment, generate the behavior change rate of the target object, determine whether the behavior change rate exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro process stage based on the significant feature vector.
[0009] Step S3: Based on the salient feature vector, generate a feature map, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the salient feature vector through feature weighting technology to obtain a weighted feature combination;
[0010] Step S4: The macro-process stage of the target object is refined by the weighted feature combination to obtain the micro-action category. The micro-action category is matched with the pre-established reference behavior library to generate the matching degree. It is determined whether the matching degree reaches the preset similarity threshold. If it does, the progressive recognition result is output and the behavior optimization feedback sequence is generated.
[0011] Step S5: Based on the behavior optimization feedback sequence, generate an optimized behavior recognition framework through an optimization algorithm, and determine whether the behavior recognition framework is suitable for the specific scenario in the container operation scenario;
[0012] Step S6: If it is not suitable, generate an adjusted rate of behavior change and re-determine whether the adjusted rate of behavior change is greater than the preset rate threshold until the final progressive recognition result is output.
[0013] Secondly, this application provides an AI-based container operation behavior recognition and analysis system, the system comprising:
[0014] The data preprocessing module is used to obtain the initial behavior sequence of the target object in the container operation scenario, and obtain the action fragments of the target object based on the initial behavior sequence;
[0015] The analysis and identification module is used to generate the rate of change of the target object's behavior based on the action segment, determine whether the rate of change of behavior exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro process stage based on the significant feature vector.
[0016] The feature weighting module is used to generate a feature map based on the salient feature vector, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the salient feature vector through feature weighting technology to obtain a weighted feature combination.
[0017] The action matching module is used to refine the macro-process stage of the target object through the weighted feature combination to obtain micro-action categories, match the micro-action categories with a pre-established reference behavior library to generate a matching degree, determine whether the matching degree reaches a preset similarity threshold, and if so, output progressive recognition results and generate a behavior optimization feedback sequence.
[0018] The framework optimization module is used to generate an optimized behavior recognition framework based on the behavior optimization feedback sequence using an optimization algorithm, and to determine whether the behavior recognition framework is suitable for a specific scenario in the container operation scenario.
[0019] The iterative output module is used to generate an adjusted rate of behavior change when the behavior recognition framework is not adapted to the specific scenario, and to re-determine whether the adjusted rate of behavior change is greater than the preset rate threshold, until the final progressive recognition result is output.
[0020] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the aforementioned AI-based container operation behavior recognition and analysis method.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0022] This application provides an AI-based method and system for container operation behavior recognition and analysis. It establishes a complete framework for recognizing personnel actions and behaviors in container operation scenarios, encompassing data acquisition, feature extraction, and dynamic optimization. First, it extracts personnel action data from video frame sequences, generates an initial behavior sequence, and subdivides it into action segments. The calculated rate of behavior change is compared with a preset rate threshold to ensure the capture of significant action changes in complex scenarios. Then, it utilizes feature weighting techniques to analyze the time decay pattern and adjust the weights of salient feature vectors, further improving the distinguishability of features across different scenarios. Subsequently, it performs action matching using a reference behavior database, obtains progressive recognition results through similarity measurement, and generates a feedback sequence based on real-time deviation correction to achieve dynamic self-adjustment. Finally, feedback is incorporated into the model parameters through a sequential iterative update method, and risk probability assessment technology is used to determine whether the recognition framework is suitable for specific scenarios, thereby improving its robustness to changing operating environments. If the model is still not suitable, it enters the iterative output stage, generating new initial behavior sequences and adjusted behavior change rates, and combining historical data comparison and trend prediction to finally output reliable macro-process stage labels and confidence levels, achieving convergence and closed-loop operation recognition. This improves the recognition accuracy and stability in complex container operation environments, enabling real-time monitoring and risk warning of operator behavior, and enhancing the management safety and operational efficiency of container operations. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of the AI-based container operation behavior recognition and analysis method in the embodiments of this application;
[0025] Figure 2 This is a schematic diagram of the structure of the AI-based container operation behavior recognition and analysis system in the embodiments of this application. Detailed Implementation
[0026] This application provides an AI-based method and system for recognizing and analyzing container operation behavior. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0027] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the AI-based container operation behavior recognition and analysis method in this application includes:
[0028] Step S1: Obtain the initial behavior sequence of the target object in the container operation scenario, and obtain the action fragments of the target object based on the initial behavior sequence.
[0029] In step S1, obtaining the action fragment of the target object includes:
[0030] Based on video data acquisition, the motion data of the target object is collected. The target object includes the operator, crane boom or other working equipment.
[0031] Limb posture decomposition technology is used to analyze motion data to obtain the initial sequence of limb movements of the target object. The limb posture decomposition technology determines the joint position through key point detection technology.
[0032] Based on gesture trajectory tracking technology, hand movement trajectories are identified from an initial sequence. The gesture trajectory tracking technology generates trajectory data by tracking the spatiotemporal changes of key points on the hand.
[0033] Based on the initial sequence of limb movements and hand movement trajectories, combined with preset thresholds for movement amplitude and time intervals, movement segments are obtained. The movement segments consist of a combination of continuous limb movements and hand gestures, and reflect the behavior of the target object within a specific time period.
[0034] Specifically, this step, as a crucial link in behavior recognition, firstly involves acquiring real-time video streams of operators using cameras installed at the container operation site. These streams are then converted into a series of continuous frames, each containing limb movement information of the target object, such as the operator, crane boom, or other operating equipment. Secondly, limb posture decomposition technology is used with key point detection methods to identify the positions of key human joints such as the shoulder, elbow, and knee, constructing a spatiotemporal sequence to provide basic limb movement features for subsequent behavior recognition. Finally, gesture trajectory tracking technology is used to generate hand trajectory data by tracking changes in key hand points in the spatiotemporal dimension. This trajectory data is represented by a series of three-dimensional coordinate point sequences, accurately reflecting the dynamic path of gesture movements, such as grasping and pointing.
[0035] The segmentation of motion segments is based on thresholds for motion amplitude and time interval. For example, when the change in limb joint position exceeds a preset amplitude threshold and the duration exceeds a preset time interval threshold, this continuous motion data will be marked as an independent motion segment for subsequent temporal analysis and behavior pattern recognition. In container operation scenarios, motion segment segmentation is crucial for high-risk operations such as crane boom fine-tuning, as it can accurately identify subtle operator movements, thereby reducing the occurrence of misidentification.
[0036] To ensure the effectiveness of this technical solution in different environments, this application also specifically considers the impact of low-light environments on video acquisition and behavior recognition. For example, in nighttime or low-light scenarios, video data acquisition can be supplemented by infrared cameras to enhance image visibility. During the limb posture decomposition process, the training dataset for low-light images has been enhanced. A custom dataset containing low-light images is used for training to improve the accuracy of joint position detection and effectively improve the accuracy of action recognition under low-light conditions. This is helpful for the safety monitoring of high-risk operations, especially for the accurate recognition of operations such as boom fine-tuning.
[0037] Furthermore, this application allows for the adjustment of relevant parameters based on different operational scenarios. For example, in daylight scenarios with high illumination, the time interval threshold can be appropriately set to a shorter value to capture rapid action sequences; while in low-light environments such as at night, the time interval threshold can be appropriately adjusted to a longer value to accommodate image acquisition delays. Through these reasonable parameter adjustments, the subtle movements of operators can be identified more accurately, ensuring the accuracy of behavior change rate calculation and behavior unit segmentation in complex scenarios.
[0038] The above technical solution combines video data acquisition, body posture decomposition, and gesture trajectory tracking technologies to reasonably divide action segments and dynamically adjust relevant parameters according to scene requirements, ensuring accurate recognition in changing environments, improving recognition accuracy, and effectively supporting safety monitoring and operation optimization in complex container operation scenarios.
[0039] Step S2: Based on the action fragment, generate the rate of change of the target object's behavior, determine whether the rate of change of behavior exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro-process stage based on the significant feature vector.
[0040] In step S2, the rate of change of the target object's behavior is generated based on the action fragment, including:
[0041] The action segment is divided into multiple time windows based on the time interval and the amplitude of the action. The division of the time window is based on the continuity of the action segment. Eye focus analysis technology is used to obtain the gaze direction data of the target object in each time window. The eye focus analysis technology determines the gaze direction by locating key eye points.
[0042] Based on gaze direction data, the frequency of gaze changes within a time window is generated, and combined with the spatiotemporal features of the action segment, a temporal feature vector is generated.
[0043] Based on temporal feature vectors, the rate of behavior change is generated. The temporal feature vectors include joint features of action speed and gaze change. The rate of behavior change is obtained through gradient analysis of the temporal feature vectors, representing the degree of change in the target object's behavior within each time window. The salient feature vectors reflect the dynamic behavior pattern of the target object.
[0044] Specifically, the division of time windows is based on the continuity of action segments to ensure the coherence of actions within each time window and avoid inaccurate identification caused by action interruptions. For example, a 5-second container crane operation segment is divided into three 1.67-second time windows by the system to more accurately capture behavioral changes within each time period. This process is achieved by quantitatively analyzing the duration and amplitude of action segments to ensure that there are no obvious breaks between actions within each window.
[0045] Eye focus analysis technology utilizes key point detection techniques for the eyes, such as eye landmark detection based on convolutional neural networks, to determine the position of the pupil and eyelids, infer the direction of gaze, calculate the gaze direction vector, and generate a gaze direction data sequence. For example, under low light conditions, the detection of key points for the eyes improves detection accuracy by enhancing image contrast. This can effectively capture the operator's gaze deviation during container operations, especially during nighttime operations, and helps improve the detection of anomalies in operations such as boom fine-tuning.
[0046] The frequency of gaze changes is calculated by performing differential analysis on gaze direction data to determine the number of gaze direction changes per unit time. For example, when the gaze deviation exceeds a certain preset threshold, the gaze change frequency is marked as a valid change. This process helps identify whether the operator is reacting to changes in the work environment and captures abnormal gaze deviations, indicating possible abnormal operational behaviors. The temporal feature vector integrates the joint features of action speed and gaze change. Action speed is extracted from video frames using optical flow, while the gaze change frequency provides quantitative information on the operator's visual concentration. The combination of temporal feature vectors allows the system to comprehensively consider both action and visual data, improving the overall accuracy of recognition. For example, if the operator's action speed is fast and the gaze frequently shifts during boom fine-tuning, the system can identify this dynamic behavior and indicate potential errors or lack of focus.
[0047] Finally, by differentiating the temporal feature vector along the time axis, the degree of behavioral change within each time window is analyzed. The rate of change of each dimension is calculated using the finite difference method, thereby obtaining the gradient value of the behavioral change rate. The rate of behavioral change is calculated using the gradient analysis method. The behavioral change rate can accurately reflect the intensity of the target object's behavior within each time window.
[0048] In step S2, it is determined whether the rate of change of behavior exceeds a preset rate threshold. If it does, a significant feature vector is generated, including:
[0049] The rate of change of behavior in each time window is compared with a preset rate threshold. If the rate of change of behavior exceeds the preset rate threshold, the time window is defined as a high dynamic window. The preset rate threshold is determined based on the statistical distribution of the rate of change of behavior. A high dynamic window represents a period of time in which the behavior changes drastically.
[0050] The system obtains salient feature vectors based on high dynamic windows. The salient feature vectors include gait rhythm features and tool interaction features. Gait rhythm segmentation technology is used to obtain gait features based on motion periodicity analysis. Tool interaction detection technology is used to identify the interaction between the target object and the tool to obtain tool interaction features.
[0051] Specifically, the rate of behavioral change within each time window is obtained by performing gradient analysis on the temporal feature vector. The temporal feature vector integrates the joint features of action speed and line of sight changes, which can comprehensively reflect the dynamic degree of the operator's actions in the container operation scenario. The rate of behavioral change in each time window is compared with a preset rate threshold. If the rate of behavioral change exceeds the preset threshold, the window is marked as a high-dynamic window, indicating a period of drastic action change. This is helpful for focusing on key actions in high-risk operations such as crane fine-tuning and cargo handling in container operations. The preset rate threshold is determined based on the statistical distribution of historical operation video data. For example, it is set by calculating the mean and standard deviation of the rate of behavioral change to cover the range of normal behavioral changes and highlight abnormal dynamics, ensuring adaptability under different lighting conditions and operating environments.
[0052] Gait rhythm segmentation technology, based on motion periodicity analysis, extracts periodic features of operator limb postures from video frame sequences. It converts time-domain signals into frequency-domain information using Fourier transform or autocorrelation functions to identify repetitive motion patterns, such as the periodic frequency of operator arm raising or movement, forming a gait rhythm feature vector to distinguish between routine operations and sudden movements. In tool interaction detection, optical flow is used to track hand trajectories, and the interaction between the operator and the tool is determined by combining the tool boundary position, such as the duration and direction of hand contact with the hook, generating a tool interaction feature vector. Gait rhythm features and tool interaction features are combined through vector concatenation or multi-dimensional fusion to form a salient feature vector. For example, gait period and amplitude are combined with interaction intensity and angle into a single vector. This vector reflects both macro-level process stages, such as boom preparation and execution, and provides micro-level motion change information for subsequent macro-level process judgment and behavior optimization feedback sequence generation.
[0053] In container operation scenarios, such as when fine-tuning the crane boom at night in low light conditions, the video frames are first enhanced to improve the reliability of detecting limb and eye key points, thereby generating accurate temporal feature vectors. Then, gradient analysis is used to calculate the rate of behavior change. When the rate exceeds a threshold, high dynamic windows are marked and significant feature vectors are extracted to detect the operator's rapid movements and incoordination between gaze and hand movements. This technical solution ensures that abnormal operator behavior can be captured in a timely manner in complex container operation environments, such as crane boom adjustment and cargo handling, providing data support for subsequent behavior optimization. This enables the system to achieve accurate behavior recognition and real-time feedback, improving the safety and efficiency of container operations.
[0054] Step S3: Based on the salient feature vectors, generate a feature map, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the salient feature vectors through feature weighting techniques to obtain a weighted feature combination.
[0055] In step S3, it is determined whether a time decay pattern exists in the feature map. If it does, the feature weights of the significant feature vectors are adjusted using feature weighting techniques to obtain a weighted feature combination, including:
[0056] The head rotation recognition technology processes the salient feature vectors to generate a feature map. The head rotation recognition technology identifies head movements by analyzing the changes in the angles of key points on the head. The feature map represents the spatiotemporal distribution of the salient feature vectors, reflecting the behavioral changes of the target object in different time windows.
[0057] Based on the feature map, a time decay pattern detection algorithm is used to determine whether a time decay pattern exists. If it exists, the feature weights of the significant feature vectors are dynamically adjusted based on the change range of the limb joint angle using joint angle quantization technology to obtain a weighted feature combination. The time decay pattern is determined by the temporal decay law of the feature intensity in the feature map, reflecting the feature that the behavior of the target object gradually weakens over time. The weighted feature combination is used for subsequent micro-action recognition.
[0058] Specifically, a salient feature vector (SMR) is a vector representation extracted within a high dynamic range window that reflects the core behavioral characteristics of a target object. It is used to transform complex behavioral data in a video frame sequence into a structured numerical representation, enabling the system to capture key dynamics in the macroscopic process phase during subsequent analysis. Specifically, it is extracted from gait rhythm and tool interaction information within the high dynamic range window. First, head rotation recognition technology is introduced to process components related to attention changes, extracting angular changes at key points such as the forehead, nose tip, and chin, and mapping them chronologically into a feature map. This provides a visual representation of the distribution of salient features along the time and spatial axes. After the feature map is formed, attenuation detection is performed on the feature intensity sequence at each spatiotemporal location. A sliding window is used to observe the continuous change in intensity from high to low, distinguishing between transient noise and continuous changes. When an attenuation pattern is determined, joint angle quantization technology is introduced to obtain the amplitude of angular changes in joints such as the shoulder, elbow, and knee. This amplitude is converted into weighting coefficients to dynamically weight the posture-related channels in the feature map, thereby obtaining a weighted feature combination.
[0059] The weighted logic is built upon the relationship between behavioral attention and limb coordination: head rotation reflects gaze shift, and joint angles reflect operational force and stability. These two aspects are aligned temporally, with timestamps used to achieve one-to-one feature correspondence and synchronous updates, thereby suppressing short-term jitter and highlighting substantive actions relevant to the task. In container hoisting scenarios, after the operator looks around the equipment, they make minor adjustments to the boom. The feature map shows a decay trajectory in the head channel, first rising and then falling, while the upper limb joint channel shows a synchronous increase in amplitude. The weighted strategy accordingly increases the weight of the upper limb channel and decreases the weight of the head channel, making the weighted feature combination more closely resemble the actual state of the work process. The method integrates salient features, feature maps, decay detection, and weight updates. By simultaneously aligning and fusing multi-source features, it solves the false detection problems caused by low light, occlusion, and background interference, improving the stability of macro-level process judgments and providing lower-noise, more discriminative input data for subsequent micro-level action refinement and template matching.
[0060] Step S4: Refine the macro-process stage of the target object by weighted feature combination to obtain micro-action categories. Match the micro-action categories with the pre-established reference behavior library to generate matching degree. Determine whether the matching degree reaches the preset similarity threshold. If it does, output the progressive recognition result and generate the behavior optimization feedback sequence.
[0061] In step S4, it is determined whether the matching degree reaches a preset similarity threshold. If it does, a progressive recognition result is output, and a behavior optimization feedback sequence is generated, including:
[0062] Feature vectors are obtained based on micro-action categories. By comparing the feature vectors with action template data in a pre-established reference behavior library, a matching degree is generated. Micro-action categories represent the type and intensity of specific action behaviors such as boom swing and operator hand operation. The matching degree represents the degree of similarity between the target object's behavior and the corresponding action in the reference behavior library.
[0063] If the matching degree is greater than or equal to the preset similarity threshold, a progressive recognition result is generated, which includes the action category label and confidence level.
[0064] The progressive recognition results are adjusted by real-time deviation correction technology to generate a behavior optimization feedback sequence. The real-time deviation correction technology updates the matching results based on error analysis of historical matching data, and the progressive recognition results are used for subsequent behavior optimization.
[0065] Specifically, the key features extracted from the weighted feature combination include multi-dimensional dynamic behavioral data such as head rotation, joint angles, gait rhythm, and tool interaction, and maintain continuity with the action segments on the timeline, ensuring that each feature vector strictly corresponds to a specific time window. A multi-layer classification model is used to process the weighted feature combination. The macro-process classification layer identifies the overall stage, determining whether the current process belongs to macro-processes such as crane positioning, container loading and unloading, or gait movement. The micro-action classification layer not only relies on joint angle sequences and gesture trajectories but also combines facial key point displacement features captured by subtle facial expression changes to refine the action category into specific operational actions and their intensity.
[0066] Then, the feature vectors corresponding to the micro-action categories are compared with a pre-established reference behavior library. The reference behavior library stores template data for different action categories in container operations. These templates are standardized vector sets obtained by decomposing and normalizing a large amount of historical video data, and a hierarchical index is constructed according to operation type, lighting conditions, and equipment parameters to ensure retrieval efficiency and contextual accuracy of matching. During the comparison process, cosine similarity and other measurement methods are used to calculate the matching degree. The equal weight comparison of data from different channels is ensured by a one-to-one correspondence of vector dimensions. At the same time, for low-light conditions at night, the weight of trajectory components is reduced in the similarity calculation, and the utilization of stable features such as gait rhythm and gaze direction is enhanced to improve robustness. If the matching degree reaches a preset similarity threshold, the system outputs a progressive recognition result, which includes the action category label and confidence level.
[0067] To reduce cumulative biases over long-term use, the progressive recognition results are dynamically adjusted before output using error analysis based on historical matching data. If historical data shows that a certain type of action is frequently misidentified as a similar category, the confidence level of that category is reduced in the current results, or the weights of the corresponding channels are corrected. For example, in nighttime scenes, historical data indicates that actions involving tool interactions are easily confused with actions involving moving hands. The correction module quantifies the frequency and magnitude of this bias and correspondingly reduces the confidence level of tool interactions, ultimately generating a corrected behavior optimization feedback sequence. This sequence not only includes the optimized label sequence but also includes risk warnings and operational suggestions, such as prompts to reduce the crane boom rotation speed or maintain eye contact. This feedback is then fed back to subsequent recognition processes to update model parameters and template weights. This ensures that the entire system maintains high recognition accuracy and stability in remote scenes, low-light conditions at night, and rapid dynamic operations.
[0068] Step S5: Based on the behavior optimization feedback sequence, generate an optimized behavior recognition framework through an optimization algorithm, and determine whether the behavior recognition framework is suitable for the specific scenario in the container operation scenario.
[0069] In step S5, an optimized behavior recognition framework is generated, and it is determined whether the behavior recognition framework is suitable for the specific scenario in container operation, including:
[0070] Based on the behavior optimization feedback sequence, the sequence iterative update technique is used to optimize the work process model, generate an optimized behavior recognition framework, obtain a new macro process stage based on the behavior recognition framework, and update the model parameters through the temporal characteristics of the feedback sequence.
[0071] The optimized behavior recognition framework uses risk probability assessment technology to determine whether it is suitable for a specific scenario. The risk probability assessment technology obtains the risk value based on the probability distribution of scenario features and action categories.
[0072] If the risk value is lower than the preset risk threshold, the optimized behavior recognition framework is determined to be suitable for specific scenarios, including action detection in low light, complex physical environments, etc.
[0073] Specifically, the behavior optimization feedback sequence includes action category labels, confidence change trajectories, joint angle corrections, and temporal continuity indicators. This data, as an input sequence, is processed by a Long Short-Term Memory (LSTM) network to capture the temporal dependencies in action evolution. During iterative updates, the LSM network receives feedback vectors frame by frame, accumulates historical features using hidden states, dynamically adjusts model parameters such as classification thresholds and weight coefficients, and ultimately outputs a new parameter vector. Through multiple rounds of iterative training, the workflow model gradually converges, forming an optimized behavior recognition framework. This framework integrates updated classification strategies and feature weights, establishing a more accurate mapping between macro-level workflow segmentation and micro-level action category recognition.
[0074] After the optimized behavior recognition framework is generated, the probability distribution of scene features and action categories is extracted using risk probability assessment technology. Scene features include the average brightness, noise distribution, and occlusion rate under low light conditions. The probability distribution comes from the softmax function of the framework's output layer, mapping the feature vectors of action categories to probability values. Risk probability assessment combines scene features and category probabilities to calculate a risk value, which reflects the degree of uncertainty of the recognition result under the current environmental conditions. For example, in a container crane boom fine-tuning scenario at night, if the light level is below a certain level and the probability distribution shows insufficient differences between multiple action categories, the risk value increases, indicating low adaptability. If the risk value is below a preset risk threshold, the optimized behavior recognition framework is determined to be able to operate stably in this specific scenario and can be directly used to update the macro-process stage, ensuring the accuracy of the stage division results in low light or complex physical environments. For example, during nighttime operations at a hub, the average brightness level of video frames is insufficient to clearly distinguish action trajectories, but the iteratively updated framework can enhance its sensitivity to gait rhythm and joint angle features. If the risk value given by the risk probability assessment is below the threshold, it confirms that the framework is adapted to the scenario, and the updated stage labels are applied to the real-time recognition results. For example, in complex physical environments with equipment obstruction, the risk probability assessment increases the weight of noise distribution. Even if the calculated risk value is still below the threshold, the framework maintains the stability of action category output, thereby reducing misidentification caused by environmental changes. In this way, a dynamic balance between iterative optimization of the recognition framework and scene adaptation is ensured, which also improves the robustness and operational safety of the recognition system.
[0075] Step S6: If it is not suitable, generate an adjusted rate of behavior change and re-determine whether the adjusted rate of behavior change is greater than the preset rate threshold until the final progressive recognition result is output.
[0076] In step S6, the adjusted rate of behavioral change is generated, and it is re-evaluated whether the adjusted rate of behavioral change is greater than a preset rate threshold until the final progressive recognition result is output, including:
[0077] Based on the new macro-process stage, a new initial behavior sequence is obtained, and based on the new initial behavior sequence, an adjusted rate of behavior change is generated;
[0078] If the adjusted rate of change of behavior is greater than the preset rate threshold, then a significant feature vector is obtained based on the high dynamic window data and through historical data comparison technology.
[0079] The predictive trend generation technology is used to map salient feature vectors to macro-process stages and output the final progressive identification results. These results are used to optimize work processes and improve work safety and efficiency.
[0080] Specifically, the system acquires new initial behavior sequences based on the updated macro-process stages. These sequences are extracted from action data within the stage, including the operator's body posture, gait rhythm, and tool interaction trajectory. New initial behavior sequences are generated by calculating statistical features, providing the basic input for subsequent temporal feature processing. The weights of the temporal features are dynamically allocated based on scene features. For example, in low-light scenes, the weight of eye focus features is increased to compensate for feature weakening caused by insufficient ambient brightness; in windy scenes, the weight of tool interaction features is increased to reflect the impact of external disturbances on the operation. The weighted temporal features are then analyzed using gradients to generate adjusted behavior change rates, which are compared with preset rate thresholds to determine whether a high-dynamic window has been entered.
[0081] If the adjusted rate of change in behavior exceeds a threshold, the system invokes historical data comparison technology within a high-dynamic window to match the current window features with action data in a historical sample library. The sample library contains historical behavior sequences accumulated under different operational conditions. The matching process generates significant feature vectors through feature similarity calculation. Then, a trend prediction generation technique is used to analyze the trend of the significant feature vectors. On one hand, methods such as linear regression are used to fit the curves of feature values over time to obtain the directional trend of behavior change. On the other hand, combined with predefined stage boundary conditions, the system maps the feature change trend along the time axis to a macro-process stage and outputs progressive recognition results, including stage labels and confidence levels. If the adjusted rate of change in behavior does not exceed the threshold, the new initial behavior sequence is maintained, and the parameter adjustment and gradient analysis process is repeated under the same scene feature constraints. This iterative process checks whether the threshold conditions are met until progressive recognition results that reflect the actual operational state are generated.
[0082] In daytime scenarios, if the salient feature vector is similar to historical high-efficiency loading and unloading operation data, the prediction trend generation technology predicts that the current sequence is developing towards the loading and unloading preparation stage. By comparing with historical samples of high-efficiency loading and unloading operations, it outputs the label and confidence level of the loading and unloading preparation stage, providing early warning of potential risks. In nighttime low-light scenarios, the system dynamically increases the weight of gesture trajectory and gait rhythm features, and generates fine-tuned error risk labels by combining them with a nighttime sample library, reducing the misjudgment rate under low-light conditions. Through the above process, it achieves the generation of adjusted behavior change rates based on new macro-process stages, and outputs progressive recognition results by combining historical data comparison and trend mapping in a high-dynamic window. This not only ensures the consistency between macro-process stages and micro-action categories, but also enhances the framework's adaptability in complex environments, thereby optimizing container operation processes and improving the safety and efficiency of container operations.
[0083] This application also provides an AI-based container operation behavior recognition and analysis system, implemented through the methods described above, such as... Figure 2 As shown, the system includes:
[0084] The data preprocessing module is used to obtain the initial behavior sequence of the target object in the container operation scenario, and based on the initial behavior sequence, obtain the action fragments of the target object.
[0085] The analysis and recognition module is used to generate the rate of change of the target object's behavior based on action segments, determine whether the rate of change of behavior exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro-process stage based on the significant feature vector.
[0086] The feature weighting module is used to generate a feature map based on salient feature vectors, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the salient feature vectors through feature weighting technology to obtain a weighted feature combination.
[0087] The action matching module is used to refine the macro-process stage of the target object through weighted feature combination to obtain micro-action categories. Based on the micro-action categories, it matches them with a pre-established reference behavior library to generate a matching degree. It determines whether the matching degree reaches a preset similarity threshold. If it does, it outputs progressive recognition results and generates a behavior optimization feedback sequence.
[0088] The framework optimization module is used to generate an optimized behavior recognition framework based on the behavior optimization feedback sequence and through optimization algorithms, and to determine whether the behavior recognition framework is suitable for the specific scenario in the container operation scenario.
[0089] The iterative output module is used to generate an adjusted rate of behavior change when the behavior recognition framework is not adapted to a specific scenario, and to re-determine whether the adjusted rate of behavior change is greater than a preset rate threshold, until the final progressive recognition result is output.
[0090] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of an AI-based container operation behavior recognition and analysis method.
[0091] In summary, this application establishes a complete framework for personnel action and behavior recognition in container operation scenarios, encompassing data acquisition, feature extraction, and dynamic optimization. First, it extracts personnel action data from video frame sequences, generates an initial behavior sequence, and subdivides it into action segments. The calculated rate of behavior change is compared with a preset rate threshold to ensure the capture of significant action changes in complex scenarios. Then, feature weighting techniques are used to analyze the time decay pattern and adjust the weights of salient feature vectors, further improving the distinguishability of features across different scenarios. Subsequently, action matching is performed using a reference behavior database, and progressive recognition results are obtained through similarity measurement. Based on real-time deviation correction, a feedback sequence is generated to achieve dynamic self-adjustment. Finally, feedback is incorporated into the model parameters through a sequential iterative update method, and risk probability assessment technology is used to determine whether the recognition framework is suitable for specific scenarios, thereby improving its robustness to changing operating environments. If the model is still not suitable, it enters the iterative output stage, generating new initial behavior sequences and adjusted behavior change rates, and combining historical data comparison and trend prediction to finally output reliable macro-process stage labels and confidence levels, achieving convergence and closed-loop operation recognition. This improves the recognition accuracy and stability in complex container operation environments, enabling real-time monitoring and risk warning of operator behavior, and enhancing the management safety and operational efficiency of container operations.
[0092] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0093] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0094] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for identifying and analyzing container operation behavior based on AI, characterized in that, include: S1: Obtain the initial behavior sequence of the target object in the container operation scenario, and obtain the action fragments of the target object based on the initial behavior sequence; S2: Based on the action fragment, generate the rate of change of the target object's behavior, determine whether the rate of change of behavior exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro process stage based on the significant feature vector. S3: Based on the significant feature vectors, generate a feature map, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the significant feature vectors through feature weighting techniques to obtain a weighted feature combination; S4: The macro-process stage of the target object is refined by weighted feature combination to obtain micro-action categories. Based on the micro-action categories, they are matched with a pre-established reference behavior library to generate matching degree. It is determined whether the matching degree reaches the preset similarity threshold. If it does, the progressive recognition result is output and the behavior optimization feedback sequence is generated. S5: Based on the behavior optimization feedback sequence, an optimized behavior recognition framework is generated through an optimization algorithm, and it is determined whether the behavior recognition framework is suitable for the specific scenario in the container operation scenario. S6: If it is not suitable, generate an adjusted rate of change of behavior and re-determine whether the adjusted rate of change of behavior is greater than the preset rate threshold until the final progressive recognition result is output. Determine whether the rate of change of behavior exceeds a preset rate threshold. If it does, generate a significant feature vector, including: comparing the rate of change of behavior in each time window with the preset rate threshold; if the rate of change of behavior exceeds the preset rate threshold, defining the time window as a high-dynamic window; obtaining a significant feature vector based on the high-dynamic window, the significant feature vector includes gait rhythm features and tool interaction features; using gait rhythm segmentation technology to obtain gait features based on motion periodicity analysis; and using tool interaction detection technology to identify the interaction between the target object and the tool to obtain tool interaction features. The process of generating an optimized behavior recognition framework and determining its suitability for specific scenarios in container operations includes: optimizing the operation process model based on the behavior optimization feedback sequence using sequence iterative update technology to generate an optimized behavior recognition framework; obtaining new macro-process stages based on the behavior recognition framework; determining whether the optimized behavior recognition framework is suitable for specific scenarios using risk probability assessment technology, which obtains risk values based on the probability distribution of scenario features and action categories; if the risk value is lower than a preset risk threshold, the optimized behavior recognition framework is deemed suitable for specific scenarios, including action detection in low-light and complex physical environments; the behavior optimization feedback sequence includes action category labels, confidence change trajectories, joint angle correction amounts, and temporal continuity indicators. The process of generating an adjusted rate of behavioral change and re-evaluating whether the adjusted rate of behavioral change exceeds a preset rate threshold, until the final progressive recognition result is output, includes: obtaining a new initial behavioral sequence based on the new macro-process stage; generating an adjusted rate of behavioral change based on the new initial behavioral sequence; if the adjusted rate of behavioral change exceeds the preset rate threshold, obtaining a significant feature vector based on high dynamic window data and historical data comparison technology; and mapping the significant feature vector to the macro-process stage using a predictive trend generation technology to output the final progressive recognition result.
2. The method according to claim 1, characterized in that, Obtaining motion segments of the target object includes: acquiring motion data of the target object based on video data, including the operator and the crane boom; analyzing the motion data using limb posture decomposition technology to obtain the initial sequence of limb movements of the target object, with the limb posture decomposition technology determining joint positions through key point detection technology; identifying hand movement trajectories from the initial sequence based on gesture trajectory tracking technology, which generates trajectory data by tracking the spatiotemporal changes of hand key points; and obtaining motion segments based on the initial sequence of limb movements and hand movement trajectories, combined with preset thresholds for movement amplitude and time intervals. The motion segments consist of a combination of continuous limb movements and gesture movements.
3. The method according to claim 1, characterized in that, Based on action segments, the rate of change of the target object's behavior is generated, including: dividing the action segment into multiple time windows based on time intervals and action amplitude; using eye focus analysis technology to obtain the gaze direction data of the target object in each time window; eye focus analysis technology determines the gaze direction by locating key eye points; based on the gaze direction data, the frequency of gaze change within the time window is generated; combined with the spatiotemporal characteristics of the action segment, a temporal feature vector is generated; based on the temporal feature vector, the rate of change of behavior is generated, wherein the temporal feature vector includes the joint features of action speed and gaze change, and the rate of change of behavior is obtained through gradient analysis of the temporal feature vector.
4. The method according to claim 1, characterized in that, The algorithm determines whether a time decay pattern exists in the feature map. If it does, it adjusts the feature weights of the salient feature vectors using feature weighting techniques to obtain a weighted feature combination. This includes: processing the salient feature vectors using head rotation recognition technology to generate a feature map; the head rotation recognition technology identifies head movements by analyzing the angle changes of key head points; and based on the feature map, using a time decay pattern detection algorithm to determine whether a time decay pattern exists. If it does, it dynamically adjusts the feature weights of the salient feature vectors based on the angle change amplitude of limb joints using joint angle quantization technology to obtain a weighted feature combination. The time decay pattern is determined by the temporal decay law of feature intensity in the feature map.
5. The method according to claim 1, characterized in that, The process involves determining whether the matching degree reaches a preset similarity threshold. If it does, a progressive recognition result is output, and a behavior optimization feedback sequence is generated. This includes: obtaining feature vectors based on micro-action categories; comparing the feature vectors with action template data in a pre-established reference behavior library to generate a matching degree; where micro-action categories represent the type and intensity of specific actions such as boom swing and operator hand operations; if the matching degree is greater than or equal to the preset similarity threshold, a progressive recognition result is generated, which includes the action category label and confidence level; and adjusting the progressive recognition result using real-time deviation correction technology to generate a behavior optimization feedback sequence.
6. An AI-based container operation behavior recognition and analysis system, used to implement the AI-based container operation behavior recognition and analysis method as described in any one of claims 1-5, characterized in that, include: The data preprocessing module is used to obtain the initial behavior sequence of the target object in the container operation scenario, and obtain the action fragments of the target object based on the initial behavior sequence; The analysis and recognition module is used to generate the rate of change of the target object's behavior based on action segments, determine whether the rate of change of behavior exceeds a preset rate threshold, and if it does, generate a significant feature vector and determine the corresponding macro process stage based on the significant feature vector. The feature weighting module is used to generate a feature map based on the salient feature vectors, determine whether there is a time decay pattern in the feature map, and if so, adjust the feature weights of the salient feature vectors through feature weighting technology to obtain a weighted feature combination. The action matching module is used to refine the macro-process stage of the target object through weighted feature combination to obtain micro-action categories. Based on the micro-action categories, it matches them with a pre-established reference behavior library to generate a matching degree. It judges whether the matching degree reaches a preset similarity threshold. If it does, it outputs progressive recognition results and generates a behavior optimization feedback sequence. The framework optimization module is used to generate an optimized behavior recognition framework based on the behavior optimization feedback sequence and through optimization algorithms, and to determine whether the behavior recognition framework is suitable for the specific scenario in the container operation scenario. The iterative output module is used to generate an adjusted rate of behavior change when the behavior recognition framework is not adapted to a specific scenario, and to re-determine whether the adjusted rate of behavior change is greater than a preset rate threshold, until the final progressive recognition result is output.
7. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Video monitoring personnel behavior identification method and system
CN119580352A
AI-based unsafe behavior identification method and system
CN119830068A