Container frame spreader large-amplitude shaking detection and early warning system
Through visual signal acquisition and image processing technology combined with deep learning models, the large shaking of container frame spreaders is detected and warned in real time, which solves the problems of high coupling, large resource consumption and insufficient detection accuracy in traditional technologies, and achieves efficient and accurate shaking detection and early warning.
Patent Information
- Application Number
- CN202510357287.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional container frame spreader rocking detection and anti-shake technology has problems such as high coupling and maintenance complexity, large resource consumption and insufficient detection accuracy, especially when instantaneous large shaking is difficult to warn in time.
The visual signal acquisition module is equipped with a high-resolution industrial camera, which detects shaking in real time through image processing technology and feature point analysis, combines long and short-term memory network deep learning models to perform shaking warnings, and designs a model training and update mechanism to ensure the accuracy and adaptability of the prediction model.
Real-time detection and early warning of large shaking of container frame spreaders is realized, which reduces the complexity of system structure and maintenance, improves detection accuracy and robustness, and significantly reduces the overall energy consumption of the system by allocating resources on demand.
Smart Images

Figure CN120208091A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial automation and intelligent monitoring, and particularly to a system for detecting and warning of large-amplitude shaking of a container frame spreader. Background Art
[0002] During the container loading and unloading process, the stability of the container frame spreader directly affects the loading and unloading efficiency and the service life of the equipment. Traditional shaking detection and anti-shaking technologies mainly rely on sensors such as inertial measurement units (IMUs), and identify the motion state of the equipment through continuous data monitoring. However, this method has the following main problems: 1. High coupling degree and maintenance complexity: Relying on additional hardware sensors such as IMUs increases the system coupling degree, and later maintenance and upgrading become more complex.
[0003] 2. High resource consumption: The continuously running anti-shaking algorithm requires a large amount of computing resources. However, the equipment is in a stable state most of the time, and only requires high-intensity processing during short periods of large-amplitude shaking, resulting in low resource utilization.
[0004] 3. Insufficient detection accuracy: Traditional methods have low detection accuracy when dealing with instantaneous (less than 1 second) large-amplitude shaking, which may lead to early warning delays and thus fail to take protective measures in a timely manner.
[0005] Therefore, a system for detecting and warning of large-amplitude shaking of a container frame spreader is needed. Summary of the Invention
[0006] Based on the existing technical problems, the present invention proposes a system for detecting and warning of large-amplitude shaking of a container frame spreader.
[0007] A system for detecting and warning of large-amplitude shaking of a container frame spreader proposed by the present invention includes a visual signal acquisition module, a shaking detection module, a feature point coordinate estimation module of the container frame spreader, a shaking warning module, and a model training and updating mechanism module. The visual signal acquisition module is equipped with a high-resolution industrial camera to collect image data of the container frame spreader in real time.
[0008] The shaking detection module detects in real time whether large-amplitude shaking occurs through image processing technology and feature point analysis, and labels the data of the shaking event; it includes data collection and preprocessing, evaluation of the degree of image blurring, evaluation of the quantity and quality of feature points, shaking determination logic, multi-frame continuous detection, and recording and feedback of shaking events.
[0009] The feature point coordinate estimation module of the container frame spreader estimates the 3D coordinates of the feature points of the container frame spreader for input to the shaking warning module of the model.
[0010] The shaking warning module predicts the possibility of future shaking based on historical shaking data using a long short-term memory network deep learning model.
[0011] The model training and updating mechanism module regularly uses the data labeled by the shaking detection module for training and updating the shaking prediction model to ensure the accuracy and adaptability of the prediction model.
[0012] Preferably, the specific implementation steps of the shaking detection module are as follows: S1. Data collection and preprocessing: a1. Data collection: Obtain continuous video frames from a visual sensor at 30 frames per second; a2. Preprocessing: First, convert the color image to a grayscale image to simplify subsequent processing. For a discrete image, the grayscale conversion can be expressed as: ; where represents the grayscale value at position ; represents the th channel value of the RGB image at position ; , , where is the image height, is the image width, correspond to the R, G, and B channels respectively; then apply Gaussian filtering to reduce image noise and improve the accuracy of blur detection. For a discrete image, the filtering operation using a 5×5 Gaussian kernel is expressed as: ; where is a 5×5 Gaussian kernel that satisfies ; , is set to be between 0.8 and 1.6; where represents the result after Gaussian filtering and denoising at position , represents summing from position -2 to 2 in the row direction and then from position -2 to 2 in the column direction; The Gaussian kernel is pre-discretized into a 5×5 matrix: ; For the edge pixels of the image, an edge filling method is adopted, such as mirror filling: , ; , ; S2. Evaluation of the degree of blurriness of the screen: Use the variance method of the Laplacian operator to quantify the degree of blurriness of the image: b1. Laplace transform; The formula used is: ; where represents the second-order differential, represents the partial derivative; specifically, on the discrete denoised image, the Laplacian operator is implemented through a convolution kernel: ; where, represents the position of the Laplace transform result; is expressed as the Laplace convolution kernel, represents the value of the Laplace convolution kernel at ; specifically, the following is adopted: ; this convolution kernel corresponds to the discrete form of the Laplacian operator: ; b2. Variance calculation; Calculate the variance of the Laplace transform, and the formula used is: ; where, represents the result of the variance calculation, and the mean of the Laplace transform is ; where, and are the height and width of the image respectively; is the total number of pixels; b3. Blur metric; The lower the variance value, the blurrier the image; Set the threshold , and the shake determination rule: ; where, is the blur judgment function, True indicates blur, False indicates clarity, I is the image, and L is the Laplace transform result of the image; for images with a 1080p resolution, is set to a value between 80 and 120.
[0013] Preferably, S3. Feature point quantity and quality evaluation: c1. Feature point set definition; Define the set of feature points in the image as: ; where, is the pixel coordinate of the th feature point; is the feature point response value calculated by the ORB algorithm; is the total number of feature points; c2. Feature point detection process; The feature point detection process includes two steps: FAST key point detection and BRIEF descriptor generation; FAST key point detection: Use the FAST algorithm to detect key points in the image; The specific formula is as follows: ; where are the coordinates of the detected key points; is the key point response value, defined as: ; max means taking the maximum value, represents the number of consecutive bright or dark points detected in the clockwise direction, represents the number of consecutive bright or dark points detected in the counterclockwise direction; respectively represent the offsets on the x and y axes of the k-th point on a circle with a radius of 3; the determination condition requires that there are at least 12 consecutive points on the circumference that are at least units brighter or at least units darker than the center point; BRIEF descriptor generation: For each detected key point , the BRIEF descriptor generation process is as follows: Step 1, generation of sampling point pairs; First, define 256 pairs of sampling points within the image area around the key point ; the sampling points are randomly sampled within a window of size : ; Among them, is the offset of the first sampling point in the -th pair relative to the key point ; is the offset of the second sampling point in the -th pair relative to the key point ; the offset follows a Gaussian distribution to ensure that the sampling points are distributed around the key point; Step 2, smoothing processing: To reduce the influence of noise, apply small-size Gaussian smoothing to the sampling area; for each sampling point , use the pixel values of the denoised grayscale image : ; among them, represents the result of Gaussian denoising within the neighborhood of the point at position , and denoising is equivalent to smoothing, is the smoothed image; Step 3, pixel comparison; perform intensity comparison on 256 pairs of sampling points respectively to generate a 256-bit binary string; for the -th pair of sampling points, the comparison result is: ; Step 4, binary string construction; generate a BRIEF descriptor for each key point , and concatenate the 256 comparison results in order into a binary string, which is the BRIEF descriptor: ; During actual storage, these 256 bits will be packed into 32 bytes; among them represents the bit string concatenation operator, ; represents the bit string concatenation operation, which connects 256 binary values in sequence; c3. Feature point quality evaluation index; To evaluate the quantity and quality of feature points, the following two indexes are defined: Number of features: Number of feature points: ; Average response intensity: Calculate the average response value of all feature points, and the formula is: ; Set the threshold of the number of feature points and the threshold of the average response intensity ; When the number of feature points or the average response intensity is less than, it is determined that the quality of the feature points has decreased.
[0014] Preferably, S4. Shake determination logic: Comprehensive consideration of the degree of picture blurring and the quantity and quality of feature points for shake determination: Single-frame determination: If and or , then there is a shake in the current frame; among them is the blurring threshold.
[0015] Preferably, S5. Multi-frame continuous detection; Multi-frame continuous shake determination: To reduce false positives, a time window mechanism is introduced: Time window setting: Dynamic time window; Automatically adjust the window size according to the current frame rate: ; among them is the window size, and FPS is the frame rate; Determination logic within the window: Within the time window, if shake is detected in the most recent consecutive frames, it is considered a real shake event, otherwise, it is determined not to be a real shake.
[0016] Preferably, S6. Shake event recording and feedback: After detecting a shake, record the relevant event information and trigger subsequent processing through system feedback; Event recording: Record the timestamp, duration, and impact degree of the shake; such as the degree of blurring and the decline amplitude of feature points; Warning and triggering: Send the shake event information to the monitoring center or the fleet management system through the wireless transmission module; Trigger the anti-shake or de-blurring algorithm to ensure that the camera picture returns to clarity.
[0017] Preferably, the feature point coordinate estimation module of the container frame spreader uses the OpenVINS system to track the two-dimensional feature points of the container frame spreader, and further estimates its three-dimensional coordinates in the real-world coordinate system to construct the input tensor for the shaking prediction of the model; where the dynamic state quantity represents the state of the feature point at time t, i is the feature point index number, and the feature point state contains m events, represents the m-th event; Initialization of the distinction between the trolley frame and the background: Initial frame annotation: In the first frame, manually annotate the polygonal area of the trolley frame, denoted as , and the remaining area is the background, denoted as ; Dynamic classification: Based on the initial polygonal area, the feature points tracked by OpenVINS will be classified as trolley frame points or background points for subsequent analysis; for each feature point , judge whether it is located within : ; where is the feature point classification function, is the trolley frame point, is the background point; these feature points are further divided into two categories: the feature points within the polygonal area of the container frame spreader are recorded as ; the feature points in the remaining non-container frame spreader area , assuming that several feature points have been marked on the ground as the real-world coordinate origin and the positioning point , are recorded as ; d1. 2D tracking of feature points: First, use OpenVINS for visual-inertial simultaneous localization and mapping of the container frame spreader to achieve real-time tracking and matching of feature points in the 2D image; Feature point extraction: It is divided into two parts: using the feature point within the polygonal area of the container frame spreader as the feature points for initial tracking; within the area outside the container frame spreader, use the feature extraction module in OpenVINS to extract the key feature point set from the continuous video frames: ; each feature point contains the two-dimensional image coordinates and its descriptor; Feature matching and tracking: Establish the correspondence of feature points between consecutive frames through the optical flow method or the feature descriptor matching algorithm to generate the matching set: ; where represents the set of feature point matches, represents the feature points of the current frame, represents the feature points of the previous frame; d2. Use OpenVINS to estimate 3D coordinates; After obtaining consistent 2D feature point matches in consecutive frames, utilize the visual-inertial fusion ability of OpenVINS to convert the 2D feature points into coordinate points in 3D space ; (1) Camera internal and external parameter calibration: Camera internal parameter calibration: The camera internal parameters describe the internal geometry and optical characteristics of the camera, including: the camera internal parameter matrix; the camera internal parameter matrix describes the intrinsic geometric properties of the camera, in the form of: ; where, and are the focal lengths in the x and y directions in pixel units respectively; and are the base point coordinates; is the parameter describing pixel skew, assumed to be 0; Distortion parameters; The camera lens introduces radial distortion and tangential distortion, which need to be corrected through a distortion model: Radial distortion correction: ; Tangential distortion correction: ; where, , are the radial distortion coefficients; are the tangential distortion coefficients; Camera internal parameter calibration method: The specific steps are as follows: Step 1. Prepare a calibration board with a known geometry; Step 2. Take multiple images of the calibration board from different angles; Step 3. Detect the corner points of the calibration board in the images; Step 4. Establish the correspondence between the image coordinates of the corner points and the actual coordinates on the calibration board; Step 5. Solve for the camera internal parameters and distortion parameters by minimizing the reprojection error; The reprojection error is defined as: ; where is the error, is the point detected in the image, is the point obtained by projecting the corresponding 3D point onto the image using the estimated camera parameters; Camera external parameter calibration; The camera external parameters define the spatial relationship of the camera coordinate system relative to the world coordinate system, including: the rotation matrix , A 3×3 rotation matrix that describes the rotation of the camera coordinate system relative to the world coordinate system. The rotation matrix is represented by Euler angles, quaternions, or axis angles. For example, the basic rotation matrices about the x, y, and z axes are , , , which are expressed as follows: ; ; ; where is the rotation angle; The complete rotation matrix is obtained by combining the basic rotation matrices; Translation vector ; A 3×1 vector that describes the position of the origin of the camera coordinate system in the world coordinate system: ; External parameter calibration method: The specific steps are as follows: Step 1, for a static scene: Use marked points or feature objects with known 3D coordinates; Step 2, for a dynamic system: Estimate the camera pose through camera motion. First, use feature matching between consecutive frames; then use the fundamental matrix or essential matrix for pose estimation; finally, combine IMU data for more robust estimation; In OpenVINS, the external parameters are dynamically estimated through the fusion of visual and inertial information: First, use visual feature tracking to calculate the relative motion of the camera; then obtain acceleration and angular velocity measurements from the IMU; and fuse visual and inertial data through an extended Kalman filter; finally, continuously optimize and update the external parameters of the camera; (2) Depth estimation and triangulation; Depth estimation is a key step in converting 2D image feature points into 3D space points. In the OpenVINS system, the depth of feature points is accurately estimated using the triangulation method through camera motion and feature point matching between consecutive frames; Mathematical modeling; Assume there is a 3D point , that is observed at two camera positions, obtaining image coordinates and ; Each camera position has its own external parameters and ; Projection equation; The camera projection model is expressed as: ; where are homogeneous image coordinates, are homogeneous world coordinates, is the scale factor depth; Linear triangulation; For two perspectives, there is: ; The equation is rewritten as a linear system: ; where is a matrix composed of camera parameters and image coordinates, is the corresponding constant term; Expanding in detail, the following linear equations are constructed: For the first camera: ; For the second camera: ; where, represents the -th row of the rotation matrix of the -th camera, is the -th element of the translation vector; Nonlinear optimization; due to measurement errors and numerical calculation limitations; define the reprojection error: ; where is the function that projects the 3D point onto the image; by minimizing the reprojection error, more accurate 3D point coordinates are obtained; The complete transformation process from the 2D feature point to the 3D point is as follows: Step 1. For the feature point , first apply the inverse transformation of the camera intrinsic matrix to obtain the normalized coordinates: ; where is the normalized coordinate, is the inverse matrix of the camera intrinsics; Step 2. Use multi-frame observation data and camera pose information for triangulation to obtain the 3D point coordinates in the camera coordinate system; where is the three-dimensional coordinate; Step 3. Convert the 3D point in the camera coordinate system to the world coordinate system: ; where, is the 3D point position in the world coordinate system, is the rotation transformation matrix, is the translation transformation matrix; in OpenVINS, this process is further enhanced through visual-inertial fusion, using IMU data to provide additional motion constraints and scale information; (3) Coordinate system transformation; Coordinate system definition: In the OpenVINS system, the following coordinate systems are involved: Image coordinate system: A 2D plane coordinate system with the origin at the upper left corner of the image and the unit being pixels; The coordinate representation is: ; The x-axis points to the right and the y-axis points downwards; Normalized camera coordinate system: A coordinate system that removes the influence of camera intrinsics; Coordinate representation: ; The optical center of the camera is at the origin, and the z-axis is along the positive direction of the camera optical axis; Camera coordinate system: A 3D coordinate system with the origin at the optical center of the camera; Coordinate representation: ; The x-axis points to the right, the y-axis points downwards, and the z-axis points forward; World coordinate system: The global reference coordinate system, with respect to which all camera poses and 3D points are represented; Coordinate representation: ; Select the initial camera position or reference point as the origin; d3. Construct the input tensor of the feature information; By obtaining the 3D position of each feature point in the world coordinate system i.e., , combined with the dynamic state quantity and historical shaking events provided by the shaking detection module, construct the input tensor for shaking prediction ; Dynamic state quantity: Utilize the current moment and the 3D coordinates at the moment of the most recent large shake to calculate the dynamic state quantity: ; Where is the moment when the most recent large shake occurred, is the feature point index; At the current t moment, is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; At the moment of the most recent large shake , is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; Historical shaking events: Collect the most recent shaking events , where each event records the dynamic state quantity of the shake that occurred at the moment ; · Input tensor construction: Integrate the current dynamic state quantity with historical shaking events to form a complete input tensor: ; This input tensor will be used as the input feature of the shaking prediction module for training and predicting future shaking events.
[0018] Preferably, the shaking warning module predicts whether a large-scale shake will occur in the future based on historical shaking data; Specifically, it includes the following steps: Step 1. Construction of input features; According to the input tensor , construct effective features, including the following parts: (1) Dynamic state quantity ; (2) Jiggle event sequence ; (3) Jiggle displacement data: average displacement of two jiggles; average displacement in horizontal direction: ; average displacement in vertical direction: ; average displacement of the current moment from the nearest previous jiggle: average displacement in horizontal direction: ; average displacement in vertical direction: ; where respectively represent the horizontal coordinates of the jiggle event feature points at times Tm, Tm-1, and t; respectively represent the vertical coordinates of the jiggle event feature points at times Tm, Tm-1, and t; Step 2. Time series analysis and feature extraction; Jiggle events have obvious time series characteristics; By analyzing the displacement patterns of historical jiggle data, the regularity of jiggles is captured, and thus the occurrence of future jiggles is predicted; Long short-term memory network; LSTM is a recurrent neural network suitable for processing and predicting time series data, capable of capturing long-term dependencies; By constructing an LSTM network and combining historical jiggle data, it is predicted whether a large-amplitude jiggle will occur within the next 1 second; Step 3. Model architecture; (1) Input layer: Accept the constructed feature vectors, including dynamic state quantities, historical jiggle events, and displacement data; (2) LSTM layer: First LSTM layer: Number of units: 128; Return sequence: True; Activation function: tanh; Dropout: 0.2; Second LSTM layer: Number of units: 64; Return sequence: False; Activation function: tanh; Dropout: 0.2; (3) Fully connected layer: First fully connected layer: Number of neurons: 64; Activation function: ReLU; Dropout: 0.2; (4) Output layer: Output the predicted displacement differences dx, dy, and dz in the horizontal and vertical directions; Explanation of the output dx and dy: Describe the changes in the horizontal directions x and y of the next large horizontal jiggle relative to the previous large jiggle; dx represents the displacement difference in the x direction, and dy represents the displacement difference in the y direction; When taking the straight track on the ground as the x-axis, y is defaulted to 0; dz: Describe the change in the vertical direction z of the next large vertical jiggle relative to the previous large jiggle.
[0019] Preferably, model training and prediction; (1) Data preprocessing: Denoising processing: Apply moving average filtering or denoising methods to smooth the 3D coordinate data and reduce the influence of measurement noise; Feature standardization: Normalize the displacement data to ensure that each feature is input into the model at the same scale; (2) Training process: Label definition: Based on historical data, label whether shaking occurs within the next 1 second; Loss function: Mean Squared Error; Optimization algorithm: Use the Adam optimizer with a learning rate set to 0.001; Regularization measures: Dropout: Add Dropout layers in the LSTM layer and fully connected layer to prevent overfitting; L2 regularization: Introduce L2 regularization in the Dense layer to limit the weight size; Data splitting: Training set: 70%; Validation set: 15%; Test set: 15%; (3) Model prediction; Input the feature data of the current moment and historical shaking events, and the model outputs the predicted ; Output explanation: dx and dy describe the changes in the horizontal direction (x, y) of the next large horizontal shake relative to the previous large shake; dz describes the change in the next large vertical shake relative to the previous large shake; Output decision: No significant change: When and and at the same time, it means that there is no significant change in the next shake compared to the previous one; where represents the feature thresholds in three directions; Significant change: When any one of the output values exceeds the corresponding threshold, it means that there will be a significant change in the next shake and anti-shake measures need to be taken.
[0020] Preferably, the model training and updating mechanism module: includes the following steps: Step 1, Data annotation and collection: Shaking detection module; Real-time detect and record shaking events, including the timestamp, duration, blur degree, and feature point drop amplitude information of the shaking occurrence; Data storage; Store the detected shaking event data in the database as the labels of the training data; Step 2, Model training: Periodic training; Regularly retrain the prediction model using the latest collected shaking event data; Incremental learning: Adopt incremental learning technology to gradually add new data to the training set and continuously optimize the model parameters; Training process: Data preparation: Extract the latest shaking event data from the database to construct the training set and validation set; Data preprocessing: Perform data normalization, missing value handling, and data augmentation to ensure the quality of the training data; Model training: Use the designed LSTM model architecture for model training; Model validation: Evaluate the model performance on the validation set, adjust the model parameters and structure, and optimize the performance; Model saving and deployment: Save the best-trained model and deploy it to the real-time prediction system; Step 3, Online update and adjustment: Real-time feedback: Dynamically adjust the model parameters and thresholds according to the actual prediction results and subsequent verification to improve the prediction accuracy; Adaptive adjustment: Use reinforcement learning technology to achieve the adaptive adjustment of the model to environmental changes and improve the robustness and stability of the system; Step 4: Model Evaluation and Optimization: Evaluation Metrics: Mean Squared Error: Measures the average squared error between the predicted value and the true value; Mean Absolute Error: Measures the average absolute error between the predicted value and the true value; R²-score: Measures the ability of the model to explain the variance of the data; Model Tuning: Optimize the model performance by adjusting the following parameters: Number of LSTM units: Increase or decrease the number of units in the LSTM layer to find the optimal model complexity; Number of time steps: Adjust the number of time steps of the input sequence and observe the impact on the model performance; Learning rate: Optimize the learning rate of the Adam optimizer to ensure the model converges quickly; Regularization parameter: Adjust the Dropout rate and the L2 regularization strength to prevent overfitting; Early Stopping Mechanism: During the training process, use the early stopping mechanism. When the performance of the validation set no longer improves, stop the training early to prevent the model from overfitting; Step 5: Continuous Learning and Optimization: Data Augmentation: Generate more jittery data through data augmentation techniques to improve the generalization ability of the model; Handling of Abnormal Data: Reasonably handle noisy data and missing data to improve the adaptability of the model to actual working conditions; Closed-loop Feedback Mechanism: Implement a closed loop of detection - warning - optimization to ensure the adaptive adjustment and continuous stable operation of the system.
[0021] The beneficial effects of the present invention are as follows: 1. Through the collaborative work of the visual signal acquisition module, jitter detection module, and jitter warning module, this system realizes the real-time detection of the jitter of the spreader and the warning of future jitter; through image processing technology and feature point analysis, it can detect in real time whether a large-scale jitter occurs currently and annotate the data of the jitter event; based on historical jitter data, using deep learning models such as Long Short-Term Memory (LSTM), it predicts the possibility of future jitter; the data annotated by the jitter detection module is regularly used for training and updating the jitter prediction model to ensure the accuracy and adaptability of the prediction model.
[0022] 2. Abandon the traditional method that requires additional sensors such as IMU, and realize jitter detection and prediction only based on visual signals, which greatly simplifies the system structure and maintenance complexity; innovatively combines image blur metric and feature point trajectory analysis to form a mutually verified detection process, improving the detection accuracy and robustness; by setting thresholds to intelligently trigger high-energy-consuming algorithms, it realizes the on-demand allocation of resources and significantly reduces the overall energy consumption of the system; using deep learning models such as LSTM for jitter prediction enables the system to change from passive response to active prevention and proactively intervene in possible jitter risks; a regular training and updating mechanism is designed to enable the model to continuously learn and adapt to different working environments and jitter patterns. Description of the Drawings
[0023] Figure 1It is a flowchart of the shaking detection module of a large-amplitude shaking detection and warning system for a container frame spreader; Figure 2 It is an overall flowchart from 2D feature points to 3D points of a large-amplitude shaking detection and warning system for a container frame spreader; Figure 3 It is a flowchart of the shaking warning module of a large-amplitude shaking detection and warning system for a container frame spreader; Figure 4 It is a discrete pixel grid table diagram of a large-amplitude shaking detection and warning system for a container frame spreader. Specific implementation mode
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0025] Refer to Figures 1-4 , a large-amplitude shaking detection and warning system for a container frame spreader, including a visual signal acquisition module, a shaking detection module, a feature point coordinate estimation module of the container frame spreader, a shaking warning module, and a model training and updating mechanism module. The visual signal acquisition module is equipped with a high-resolution industrial camera to collect image data of the container frame spreader in real time.
[0026] The shaking detection module detects whether a large-amplitude shaking occurs in real time through image processing technology and feature point analysis, and labels the data of the shaking event; it includes data acquisition and preprocessing, evaluation of the degree of image blurring, evaluation of the quantity and quality of feature points, shaking determination logic, multi-frame continuous detection, and recording and feedback of shaking events.
[0027] The specific implementation steps of the shaking detection module are: S1. Data acquisition and preprocessing: a1. Data acquisition: Obtain continuous video frames from the visual sensor, 30 frames per second; a2. Preprocessing: First, convert the color image to a grayscale image to simplify subsequent processing. For a discrete image, the grayscale conversion is expressed as: ; Among them, represents the grayscale value at position ; represents the th channel value of the RGB image at position ; , , where is the image height, is the image width, Correspond to the R, G, and B channels respectively; then apply Gaussian filtering to reduce image noise and improve the accuracy of blur detection. For a discrete image, the filtering operation with a 5×5 Gaussian kernel is expressed as: ; where is a 5×5 Gaussian kernel that satisfies ; , is set to be between 0.8 and 1.6; where represents the result after Gaussian filtering and denoising at position ; means summing from position -2 to 2 in the row direction and then from position -2 to 2 in the column direction; The Gaussian kernel is pre - discretized into a 5×5 matrix: ; For the edge pixels of the image, an edge padding method is adopted, mirror padding: , ; , ; S2. Evaluation of the degree of image blur: The variance method of the Laplacian operator is used to quantify the degree of image blur: b1. Laplace transform; The formula used is: ; where represents the second - order differential, represents the partial derivative; Specifically, on the discrete denoised image, the Laplacian operator is implemented through a convolution kernel: ; where represents the result of the Laplace transform at position ; represents the Laplace convolution kernel, represents the value of the Laplace convolution kernel at ; Specifically, it is adopted: ; This convolution kernel corresponds to the discrete - form Laplacian operator: ; b2. Variance calculation; Calculate the variance of the Laplace transform, and the formula used is: ; where represents the result of the variance calculation, and the mean of the Laplace transform is ; where and are the height and width of the image respectively; is the total number of pixels; b3. Blur metric; The lower the variance value, the blurrier the image; Set a threshold , Shaking determination rule: ; among which, is a fuzzy judgment function, True indicates fuzzy, False indicates clear, I is the image, and L is the result of the Laplace transform of the image; for an image with a resolution of 1080p, is set to a value between 80 and 120.
[0028] S3. Evaluation of the number and quality of feature points: c1. Definition of the feature point set; the feature point set in the image is defined as: ; among which, is the pixel coordinate of the -th feature point; is the response value of the feature point calculated by the ORB algorithm; is the total number of feature points; c2. Feature point detection process; the feature point detection process includes two steps: FAST key point detection and BRIEF descriptor generation; FAST key point detection: Use the FAST algorithm to detect key points in the image; the specific formula is as follows: ; among which are the coordinates of the detected key points; is the key point response value, defined as: ; max means taking the maximum value, represents the number of consecutive bright or dark points detected in the clockwise direction, represents the number of consecutive bright or dark points detected in the counterclockwise direction; respectively represent the offsets on the x and y axes of the k-th point on a circle with a radius of 3; the determination condition requires that there are at least 12 consecutive points on the circumference that are at least units brighter or at least units darker than the center point.
[0029] Implementation based on discrete images is as follows: Definition of circumferential pixel points: For a candidate point in a discrete image, define 16 sampling points on a circle with a radius of 3 around it: ; among which, for the 16 sampling points on a circle with a radius of 3 around, the offset is defined as Figure 4 shown, and these offsets approximately form a circle with a radius of 3, which is applicable to the discrete pixel grid.
[0030] Continuity determination: Define the function to represent the brightness difference state of the -th circumferential point relative to the center point: ; among which: is the pixel value of the image at the position ; is the brightness difference threshold, usually set to 20.
[0031] Continuity judgment: For a point in the image, consider the extended circumferential point sequence: ; Define the continuous counting function: ; where, is the coordinate of the current detection point, represents the brighter or darker state, is the traversal variable. For the circumferential point sequence , traverse from all starting points where m = 0 to m = 15; for each starting point m, check the length n of the continuous subsequence starting from this starting point, such that the states of all points are equal to v; find the length of the longest continuous subsequence among all possible starting points; return this maximum length, indicating the number of the most continuous bright or dark points around the point.
[0032] BRIEF descriptor generation: For each detected key point , the BRIEF descriptor generation process is as follows: Step 1, generation of sampling point pairs; First, define 256 pairs of sampling points in the image area around the key point ; These sampling points are usually randomly sampled within a -sized window: ; where, is the offset of the first sampling point in the -th pair relative to the key point ; is the offset of the second sampling point in the -th pair relative to the key point ; These offsets usually follow a Gaussian distribution , ensuring that the sampling points are mainly distributed around the key point; Step 2, smoothing processing: To reduce the influence of noise, usually apply small-size Gaussian smoothing to the sampling area; for each sampling point , use the pixel value of the denoised grayscale image : ; where, represents the result of Gaussian denoising within the neighborhood range of the position point , and denoising is equivalent to smoothing, is the smoothed image; Step 3, pixel comparison; perform intensity comparison on 256 pairs of sampling points respectively to generate a 256-bit binary string; for the -th pair of sampling points, the comparison result is: ; Step 4, Binary string construction; For each key point Generate a BRIEF descriptor , and concatenate 256 comparison results in order to form a binary string, namely the BRIEF descriptor: ; In actual storage, these 256 bits will be packed into 32 bytes; where represents the bit string concatenation operator, ; represents the bit string concatenation operation, connecting 256 binary values in order.
[0033] Specific example: If , , , ; Then ; In the BRIEF descriptor, compare 256 pairs of points: If the comparison result of the point pair is , , ..., ; Then the final descriptor is a 256-bit binary string: .
[0034] c3. Feature point quality evaluation index; In order to evaluate the quantity and quality of feature points, define the following two indexes: Feature quantity: The number of feature points: ; Average response intensity: Calculate the average response value of all feature points, and the formula is: ; Set the feature point quantity threshold and the average response intensity threshold ; When the number of feature points or the average response intensity , it is determined that the quality of the feature points has decreased.
[0035] Adaptive threshold setting mechanism: To ensure the stability of the system in different scenarios and lighting conditions, design an adaptive threshold setting mechanism; Initialization stage; When the system starts, collect the video data of the initial 5 minutes as the reference calibration set ; Calculate the set of the number of feature points for each frame of image ; Calculate the set of average response intensities ; Calculate the set of blur evaluation values .
[0036] Robust Threshold Determination: Using trimmed mean for the number of feature points: ; Using trimmed mean for the average response intensity: ; Using trimmed mean for the blur evaluation value: .
[0037] Adaptive Threshold Calculation: Threshold for the number of feature points: ; where ; Threshold for the average response intensity: ; where ; Threshold for the blur: ; where .
[0038] Threshold Dynamic Adjustment: Recalculate the statistical values for the recent period regularly or when a significant change in the lighting environment is detected; Apply exponential moving average for smooth update: ; ; ; where represents the weight of the historical threshold; represents the newly calculated threshold; represents the historical threshold obtained in the previous round of calculation; is the statistical average of the detected number of key points, is the average or maximum value of the response, is the average or maximum value of the image blur degree; is the influence factor for adjusting the recent statistical values.
[0039] S4. Jitter Judgment Logic: Comprehensive consideration of the degree of picture blur and the quantity and quality of feature points for jitter judgment: Single-frame Judgment: If and or , then there is jitter in the current frame; where is the blur threshold.
[0040] S5. Multi-frame Continuous Detection; Multi-frame Continuous Jitter Judgment: To reduce false positives, introduce a time window mechanism: Time window setting: Dynamic time window; Automatically adjust the window size according to the current frame rate: ; where is the window size and FPS is the frame rate; Judgment logic within the window: Within the time window, if jitter is detected in the most recent consecutive frames, it is considered a real jitter event, otherwise, it is determined not to be a real jitter.
[0041] S6. Jitter Event Recording and Feedback: When jitter is detected, record the relevant event information and trigger subsequent processing through system feedback; Event record: Record the timestamp, duration, and impact degree of the shaking; such as the blur degree and the decrease amplitude of feature points; Warning and triggering: Through the wireless transmission module, send the shaking event information to the monitoring center or the fleet management system; trigger the anti-shake or de-blurring algorithm to ensure that the camera image is restored clearly.
[0042] Comprehensive shaking score Calculation formula: ; where the weighted judgment score: blur judgment score ; key point judgment score ; response value judgment score ; where the weights are default set to the blur weight , key point weight , response value weight ; current blur threshold , current key point threshold , current response value threshold .
[0043] After detecting the shaking, record the relevant event information in detail: Event record: Record the timestamp, duration, and impact degree of the shaking; calculate the shaking severity score : ; the value range is 1-10 points; record the percentage of feature changes before and after the shaking: ; ; where, is the percentage change in the number of key points, is the percentage change in the response value of key points; is the number of key points detected before the shaking, is the number of key points detected during the shaking; is the average response value of all detected key points before the shaking, is the average response value of all detected key points during the shaking; Warning and triggering: Issue graded warnings according to the shaking severity; Mild shaking 1-3 points: Record but do not trigger an alarm; Moderate shaking 4-7 points: Trigger a low-priority alarm; Severe shaking 8-10 points: Trigger a high-priority alarm and start the anti-shake algorithm.
[0044] The feature point coordinate estimation module of the container frame spreader estimates the 3D coordinates of the feature points of the container frame spreader, which is used as the input for the shaking warning module of the model.
[0045] The feature point coordinate estimation module of the container frame spreader uses the OpenVINS system to track the 2D feature points of the container frame spreader and further estimates its 3D coordinates in the real-world coordinate system to construct the input tensor for the shaking prediction of the model; where the dynamic state quantity represents the state of the feature point at time t, i is the feature point index number, and the feature point state contains m events, representing the m-th event; Initialization of the distinction between the trolley frame and the background: Initial frame annotation: In the first frame, manually annotate the polygonal area of the trolley frame, denoted as , and the rest of the area is the background, denoted as ; Dynamic classification: Based on the initial polygonal area, the feature points tracked by OpenVINS will be classified as trolley frame points or background points for subsequent analysis; for each feature point , judge whether it is located in : ; where is the feature point classification function, is a trolley frame point, is a background point; these feature points are further divided into two categories: the feature points within the polygonal area of the container frame spreader are recorded as ; the feature points in the remaining non-container frame spreader area , assuming that several feature points have been marked on the ground as the real-world coordinate origin and the positioning points , are recorded as ; d1. 2D tracking of feature points: First, use OpenVINS for visual-inertial simultaneous localization and mapping of the container frame spreader to achieve real-time tracking and matching of feature points in 2D images; Feature point extraction: Divided into two parts: Use the feature points within the polygonal area of the container frame spreader as the feature points for initial tracking; within the area outside the container frame spreader, use the feature extraction module in OpenVINS to extract a set of key feature points from consecutive video frames: ; each feature point contains the two-dimensional image coordinates and its descriptor; Feature matching and tracking: Establish the correspondence of feature points between consecutive frames through the optical flow method or the feature descriptor matching algorithm to generate a matching set: ; where Represents the set of feature point matches, Represents the feature points of the current frame, Represents the feature points of the previous frame; d2. Use OpenVINS to estimate 3D coordinates; After obtaining consistent 2D feature point matches in consecutive frames, utilize the visual-inertial fusion ability of OpenVINS to convert the 2D feature points into coordinate points in 3D space ; (1) Camera intrinsic and extrinsic calibration: Camera intrinsic calibration: The camera intrinsics describe the internal geometry and optical characteristics of the camera, including: the camera intrinsic matrix; the camera intrinsic matrix Describes the intrinsic geometric properties of the camera, in the form of: ; Among them, and are the focal lengths in the x and y directions in pixel units respectively; and are the base point coordinates; Is the parameter describing pixel skew, assumed to be 0; Distortion parameters; The camera lens introduces radial and tangential distortions, which need to be corrected through a distortion model: Radial distortion correction: ; Tangential distortion correction: ; Among them, , Are the radial distortion coefficients; Is the tangential distortion coefficient; Camera intrinsic calibration method: The specific steps are as follows: Step 1. Prepare a calibration board with a known geometry; Step 2. Take multiple images of the calibration board from different angles; Step 3. Detect the corner points of the calibration board in the images; Step 4. Establish the correspondence between the image coordinates of the corner points and the actual coordinates on the calibration board; Step 5. Solve for the camera intrinsics and distortion parameters by minimizing the reprojection error; The reprojection error is defined as: ; Among them Is the error, Is the point detected in the image, Is the point obtained by projecting the corresponding 3D point onto the image using the estimated camera parameters; Camera extrinsic calibration; The camera extrinsics define the spatial relationship of the camera coordinate system relative to the world coordinate system, including: the rotation matrix , A 3×3 rotation matrix that describes the rotation of the camera coordinate system relative to the world coordinate system. The rotation matrix is represented by Euler angles, quaternions, or axis angles. For example, the basic rotation matrices around the x, y, and z axes are , , , which are expressed as follows: ; ; ; where is the rotation angle; The complete rotation matrix is obtained by combining the basic rotation matrices; Translation vector ; A 3×1 vector that describes the position of the origin of the camera coordinate system in the world coordinate system: ; Extrinsic parameter calibration method: The specific steps are as follows: Step 1. For a static scene: Use marked points or feature objects with known 3D coordinates; Step 2. For a dynamic system: Estimate the camera pose through camera motion. First, use feature matching between consecutive frames; then use the fundamental matrix or essential matrix for pose estimation; finally, combine IMU data for more robust estimation; In OpenVINS, the extrinsic parameters are dynamically estimated through the fusion of visual and inertial information: First, use visual feature tracking to calculate the relative motion of the camera; then obtain acceleration and angular velocity measurements from the IMU; and fuse visual and inertial data through an extended Kalman filter; finally, continuously optimize and update the extrinsic parameters of the camera; (2) Depth estimation and triangulation; Depth estimation is a key step in converting 2D image feature points into 3D space points. In the OpenVINS system, the depth of feature points is accurately estimated using the triangulation method through the camera motion and feature point matching between consecutive frames; Mathematical modeling; Assume there is a 3D point , that is observed at two camera positions, obtaining image coordinates and ; Each camera position has its own extrinsic parameters and ; Projection equation; The camera projection model is expressed as: ; where are the homogeneous image coordinates, are the homogeneous world coordinates, is the scale factor depth; Linear triangulation; For two viewpoints, there is: ; The equation is rewritten as a linear system: ; where is a matrix composed of camera parameters and image coordinates, is the corresponding constant term; Expanding in detail, the following linear equations are constructed: For the first camera: ; For the second camera: ; where represents the -th row of the rotation matrix of the -th camera, is the -th element of the translation vector; Nonlinear optimization; due to measurement errors and numerical calculation limitations; define the reprojection error: ; where is the function that projects the 3D point onto the image; by minimizing the reprojection error, more accurate 3D point coordinates are obtained; The complete transformation process from the 2D feature point to the 3D point is as follows: Step 1. For the feature point , first apply the inverse transformation of the camera intrinsic matrix to obtain the normalized coordinates: ; where is the normalized coordinate, is the inverse matrix of the camera intrinsics; Step 2. Use multi-frame observation data and camera pose information for triangulation to obtain the 3D point in the camera coordinate system; obtain the 3D point coordinates in the camera coordinate system; where are the three-dimensional coordinates; Step 3. Convert the 3D point in the camera coordinate system to the world coordinate system: ; where, is the position of the 3D point in the world coordinate system, is the rotation transformation matrix, is the translation transformation matrix; in OpenVINS, this process is further enhanced through visual-inertial fusion, using IMU data to provide additional motion constraints and scale information; (3) Coordinate system transformation; Coordinate system definition: In the OpenVINS system, the following coordinate systems are involved: Image coordinate system: A 2D plane coordinate system with the origin at the upper left corner of the image and the unit being pixels; The coordinate representation is: ; The x-axis points to the right and the y-axis points downwards; Normalized camera coordinate system: A coordinate system that removes the influence of camera internal parameters; Coordinate representation: ; The camera optical center is at the origin, and the z-axis is along the positive direction of the camera optical axis; Camera coordinate system: A 3D coordinate system with the origin at the camera optical center; Coordinate representation: ; The x-axis is to the right, the y-axis is downward, and the z-axis is forward; World coordinate system: A global reference coordinate system, with all camera poses and 3D points represented relative to this coordinate system; Coordinate representation: ; Select the initial camera position or reference point as the origin.
[0046] Coordinate transformation chain: From 2D image feature points to 3D points in the world coordinate system, the complete transformation chain is.
[0047] Image coordinates to normalized camera coordinates: ; Where is the camera internal parameter matrix; Normalized camera coordinates to camera coordinates; Adding depth information ; Where is the depth value obtained by triangulation; Camera coordinates to world coordinates: ; Where is the rotation matrix from the camera coordinate system to the world coordinate system, is the translation vector; It should be noted that and are related to the camera pose as follows: ; .
[0048] Coordinate transformation example: To more clearly understand the conversion process from 2D feature points to 3D world coordinates, the following is a numerical example: Assume: Camera internal parameter matrix: ; At time t, the camera external parameters; Rotation matrix ; Translation vector: ; Image feature point: ; Depth obtained by triangulation: .
[0049] Calculation steps: Convert image coordinates to normalized coordinates: ; Calculate the 3D point in the camera coordinate system: ; Calculate the camera-to-world transformation: ; ; Convert the 3D point in the camera coordinate system to the world coordinate system: ; In this way, the conversion from image feature points Transformation of 3D points into the world coordinate system .
[0050] d3. Construct the input tensor with characteristic information; by obtaining the 3D positions of each feature point in the world coordinate system i.e., , combine the dynamic state quantity provided by the shaking detection module and historical shaking events to construct the input tensor for shaking prediction ; Dynamic state quantity: Utilize the 3D coordinates at the current moment and the moment of the last major shake to calculate the dynamic state quantity: ; where is the moment when the last major shake occurred, is the feature point index; at the current t moment, is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; at the moment of the last major shake , is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; Historical shaking events: Collect the recent shaking events , where each event records the shaking dynamic state quantity that occurred at the moment ; Input tensor construction: Integrate the current dynamic state quantity and historical shaking events to form a complete input tensor: ; This input tensor will be used as the input feature of the shaking prediction module for training and predicting future shaking events.
[0051] The said shaking warning module predicts the possibility of future shaking based on historical shaking data using a long short-term memory network deep learning model.
[0052] The shaking warning module predicts whether a large-scale shake will occur in the future based on historical shaking data; specifically, it includes the following steps: Step 1. Construction of input features; according to the input tensor , construct effective features, including the following parts; (1) Dynamic state quantity ; (2) Shaking event sequence ; (3) Shaking displacement data: the average displacement of two shakes; average displacement in the horizontal direction: ; average displacement in the vertical direction: ; average displacement from the current moment to the nearest previous shake: average displacement in the horizontal direction: ; average displacement in the vertical direction: ; where respectively represent the horizontal coordinates of the shaking event feature points at times Tm, Tm-1, and t; respectively represent the vertical coordinates of the shaking event feature points at times Tm, Tm-1, and t; Step 2: Time series analysis and feature extraction; Shaking events have obvious time series characteristics; By analyzing the displacement patterns of historical shaking data, the regularity of shaking is captured, and then the occurrence of future shaking is predicted; Long short-term memory network; LSTM is a recurrent neural network suitable for processing and predicting time series data, which can capture long-term dependencies; By constructing an LSTM network and combining historical shaking data, it is predicted whether a large-amplitude shake will occur within the next 1 second. Step 3: Model architecture; (1) Input layer: Accept the constructed feature vectors, including dynamic state quantities, historical shaking events, and displacement data; (2) LSTM layer: The first LSTM layer: number of units: 128; return sequence: True; activation function: tanh; Dropout: 0.2; The second LSTM layer: number of units: 64; return sequence: False; activation function: tanh; Dropout: 0.2; (3) Fully connected layer: The first fully connected layer: number of neurons: 64; activation function: ReLU; Dropout: 0.2; (4) Output layer: Output the predicted displacement differences dx, dy, and dz in the horizontal and vertical directions; Explanation of the output dx and dy: Describe the changes in the horizontal directions x and y of the next large horizontal shake relative to the previous large shake; dx represents the displacement difference in the x direction, and dy represents the displacement difference in the y direction; When taking the straight track on the ground as the x-axis, y is defaulted to 0; dz: Describe the change in the vertical direction z of the next large vertical shake relative to the previous large shake.
[0053] Model training and prediction; (1) Data preprocessing: Denoising processing: Apply moving average filtering or other denoising methods to smooth the 3D coordinate data and reduce the influence of measurement noise; Feature standardization: Normalize the displacement data to ensure that each feature is input into the model at the same scale. (2)Training process: Label definition: Based on historical data, label whether shaking occurs within the next 1 second; Loss function: Mean Squared Error; Optimization algorithm: Use the Adam optimizer with a learning rate set to 0.001; Regularization measures: Dropout: Add Dropout layers in the LSTM layer and the fully connected layer to prevent overfitting; L2 regularization: Introduce L2 regularization in the Dense layer to limit the weight size; Data splitting: Training set: 70%; Validation set: 15%; Test set: 15%; (3)Model prediction; Input the feature data of the current moment and historical shaking events, and the model outputs the predicted ; Output interpretation: dx and dy describe the changes in the horizontal direction (x, y) of the next large horizontal shaking relative to the previous large shaking; dz describes the change in the next large vertical shaking relative to the previous large shaking; Output decision: No significant change: When and and it means that there is no significant change in the next shaking compared to the previous one; where represents the feature thresholds in three directions; Significant change: When any one of the output values exceeds the corresponding threshold, it means that there will be a significant change in the next shaking and anti-shake measures need to be taken.
[0054] The model training and updating mechanism module regularly uses the data labeled by the shaking detection module r for training and updating the shaking prediction model to ensure the accuracy and adaptability of the prediction model.
[0055] Model training and updating mechanism module: Includes the following steps: Step 1. Data annotation and collection: Shaking detection module; Real-time detect and record shaking events, including the timestamp, duration, blur degree, and feature point drop amplitude information of the shaking occurrence; Data storage; Store the detected shaking event data in the database as the labels of the training data; Step 2. Model training: Periodic training; Regularly retrain the prediction model using the latest collected shaking event data; Incremental learning: Adopt incremental learning technology to gradually add new data to the training set and continuously optimize the model parameters; Training process: Data preparation: Extract the latest shaking event data from the database to construct the training set and the validation set; Data preprocessing: Perform data normalization, missing value handling, and data augmentation to ensure the quality of the training data; Model training: Use the designed LSTM model architecture for model training; Model validation: Evaluate the model performance on the validation set, adjust the model parameters and structure to optimize the performance; Model saving and deployment: Save the best model after training and deploy it to the real-time prediction system; Step 3. Online Update and Adjustment: Real-time Feedback: Dynamically adjust the model parameters and thresholds according to the actual prediction results and subsequent verification to improve the prediction accuracy; Adaptive Adjustment: Utilize reinforcement learning techniques to achieve the adaptive adjustment of the model to environmental changes and improve the robustness and stability of the system; Step 4. Model Evaluation and Optimization: Evaluation Metrics: Mean Squared Error: Measure the average squared error between the predicted value and the true value; Mean Absolute Error: Measure the average absolute error between the predicted value and the true value; R²-score: Measure the ability of the model to explain the variance of the data; Model Tuning: Optimize the model performance by adjusting the following parameters: Number of LSTM Units: Increase or decrease the number of units in the LSTM layer to find the optimal model complexity; Number of Time Steps: Adjust the number of time steps of the input sequence and observe the impact on the model performance; Learning Rate: Optimize the learning rate of the Adam optimizer to ensure the rapid convergence of the model; Regularization Parameter: Adjust the Dropout rate and the L2 regularization strength to prevent overfitting; Early Stopping Mechanism: During the training process, use the early stopping mechanism. When the performance of the validation set no longer improves, stop the training in advance to prevent the model from overfitting; Step 5. Continuous Learning and Optimization: Data Augmentation: Generate more jittery data through data augmentation techniques to improve the generalization ability of the model; Handling of Abnormal Data: Reasonably handle noisy data and missing data to improve the adaptability of the model to actual working conditions; Closed-loop Feedback Mechanism: Implement a closed loop of detection - warning - optimization to ensure the adaptive adjustment and continuous stable operation of the system.
[0056] This system realizes the real-time detection of the spreader jitter and the warning of future jitter through the collaborative work of the visual signal acquisition module, the jitter detection module, and the jitter warning module; Through image processing technology and feature point analysis, it can detect in real time whether a large-scale jitter occurs currently and label the data of the jitter event; Based on historical jitter data, use deep learning models such as Long Short-Term Memory (LSTM) to predict the possibility of future jitter; Regularly use the data labeled by the jitter detection module for training and updating the jitter prediction model to ensure the accuracy and adaptability of the prediction model.
[0057] Abandoning the traditional method that requires additional sensors such as IMUs, it realizes jitter detection and prediction only based on visual signals, greatly simplifying the system structure and maintenance complexity; innovatively combining image blur metrics with feature point trajectory analysis to form a mutually verified detection process, improving the detection accuracy and robustness; intelligently triggering high-performance algorithms by setting thresholds to achieve on-demand allocation of resources and significantly reducing the overall system energy consumption; using deep learning models such as LSTM for jitter prediction to transform the system from passive response to active prevention and proactively intervening in possible jitter risks; designing a regular training and updating mechanism to enable the model to continuously learn and adapt to different working environments and jitter patterns As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A container frame spreader large-scale shaking detection and early warning system, characterized by: It includes a visual signal acquisition module, a shaking detection module, a characteristic point coordinate estimation module of the container frame spreader, a shaking warning module and a model training and updating mechanism module. The visual signal acquisition module is equipped with a high-resolution industrial camera to collect image data of the container frame spreader in real time; The shaking detection module detects whether a large shaking is currently occurring in real time through image processing technology and feature point analysis, and performs data annotation on the shaking event; Including data collection and preprocessing, image blur assessment, feature point quantity and quality assessment, shake determination logic, multi-frame continuous detection, and shake event recording and feedback; The feature point coordinate estimation module of the container frame spreader estimates the 3D coordinates of the feature points of the container frame spreader for input into the shaking warning module of the model; The shaking warning module predicts future shaking based on historical shaking data using a long short-term memory network deep learning model; The model training and updating mechanism module uses the data marked by the shaking detection module to train and update the shaking prediction model regularly to ensure the accuracy and adaptability of the prediction model.
2. A container frame spreader large-amplitude shaking detection and early warning system according to claim 1, characterized in that: The specific implementation steps of the shaking detection module are: S1. Data collection and preprocessing: a1. Data acquisition: Acquire continuous video frames from the visual sensor at 30 frames per second; a2. Preprocessing: First convert the color image into a grayscale image to simplify subsequent processing. For discrete images, the grayscale conversion is expressed as: ; in, Indicates location The gray value at ; Indicates location The RGB image at Channel values; , ,in is the image height, is the image width, Corresponding to R, G and B channels respectively; Gaussian filtering is then applied to reduce image noise and improve the accuracy of blur detection. For discrete images, the filtering operation using a 5×5 Gaussian kernel is expressed as: ;in, is a 5×5 Gaussian kernel, satisfying ; , Set to between 0.8 and 1.6; Indicates at location The result after Gaussian filtering and denoising is: It means that the value is accumulated from position -2 to position 2 in the row direction, and then accumulated from position -2 to position 2 in the vertical direction; The Gaussian kernel is pre-discretized into a 5×5 matrix: ; For the edge pixels of the image, use the edge filling method and mirror filling: , ; , ; S2. Image blur evaluation: The Laplace operator variance method is used to quantify the blur degree of the image: b1, Laplace transform; the formula used is: ;in represents the second-order differential, Represents partial derivative; specifically, on a discrete denoised image, the Laplacian operator is implemented through a convolution kernel: ;in, Indicates location The Laplace transform result at ; Represented as the Laplace convolution kernel, Indicates that the Laplace convolution kernel is The value at; specifically: ; This convolution kernel corresponds to the discrete form of the Laplacian operator: ; b2. Variance calculation: The variance of Laplace transform is calculated using the formula: ;in, It is expressed as the result of variance calculation, and the mean of Laplace transform is ;in, and are the height and width of the image respectively; is the total number of pixels; b3, blur measure; the lower the variance value, the blurrier the image; set the threshold , shaking judgment rules: ;in, is the blur judgment function, True means blur, False means clear, I is the image, and L is the Laplace transform result of the image; for an image with a resolution of 1080p, Set to a value between 80-120.
3. A container frame spreader large-amplitude shaking detection and early warning system according to claim 2, characterized in that: S3. Feature point quantity and quality assessment: c1. Definition of feature point set: The feature point set in the image is defined as: ;in, For the The pixel coordinates of feature points; It is the feature point response value calculated by ORB algorithm; is the total number of feature points; c2. Feature point detection process: The feature point detection process includes two steps: FAST key point detection and BRIEF descriptor generation; FAST key point detection: Use the FAST algorithm to detect key points in the image; the specific formula is as follows: ;in are the detected key point coordinates; is the key point response value, defined as: ; max means taking the maximum value, Indicates the number of consecutive bright or dark spots detected in a clockwise direction. Indicates the number of consecutive bright or dark spots detected in a counterclockwise direction; They represent the offset of the kth point on the x and y axes on a circle with a radius of 3. The judgment condition requires that there are at least 12 consecutive points on the circumference that are at least 1 / 2 of the brighter than the center point. units or darker units; BRIEF descriptor generation: For each detected key point , the BRIEF descriptor generation process is as follows: Step 1: Generate sampling point pairs; First, at the key point 256 pairs of sampling points are defined in the surrounding image area; the sampling points are in a Random sampling in a window of size: ; in, It is Center the first sample point relative to the key point The offset; It is Center the second sampling point relative to the key point The offset follows a Gaussian distribution , ensure that the sampling points are distributed around the key points; Step 2: Smoothing: In order to reduce the impact of noise, a small-size Gaussian smoothing is applied to the sampling area; for each sampling point , using a denoised grayscale image Pixel value: ;in, express Location The result after Gaussian denoising within the neighborhood. Denoising and smoothing are equivalent. To smooth the image; Step 3: Pixel comparison: compare the intensity of each of the 256 pairs of sampling points to generate a 256-bit binary string; For the sampling points, compare the results for: ; Step 4: Binary string construction; for each key point Generate BRIEF descriptor , concatenate the 256 comparison results in sequence into a binary string, namely the BRIEF descriptor: ; When actually stored, these 256 bits will be packed into 32 bytes; represents the bit string concatenation operator, ;Number It represents the bit string concatenation operation, which connects 256 binary values in sequence; c3. Feature point quality evaluation index: In order to evaluate the quantity and quality of feature points, the following two indicators are defined: Number of features: Number of feature points: ; Average response intensity: Calculate the average response value of all feature points. The formula is: ; Set the feature point number threshold and the average response intensity threshold ; When the number of feature points or average response intensity When , it is determined that the quality of the feature points has deteriorated.
4. A container frame spreader large-amplitude shaking detection and early warning system according to claim 3, characterized in that: S4, shake determination logic: shake determination is performed based on the degree of image blur and the number and quality of feature points: Single frame judgment: If and or , then the current frame has shaking, where is the fuzzy threshold.
5. A container frame spreader large-amplitude shaking detection and early warning system according to claim 4, characterized in that: S5, multi-frame continuous detection; multi-frame continuous shaking judgment: in order to reduce misjudgment, the time window mechanism is introduced: time window setting: dynamic time window; automatically adjust the window size according to the current frame rate: ;in is the window size, FPS is the frame rate; Window judgment logic: within the time window, if the most recent continuous If the frame detects shaking, it is considered as a real shaking event, otherwise it is determined not to be a real shaking.
6. A container frame spreader large-amplitude shaking detection and early warning system according to claim 5, characterized in that: S6, shaking event recording and feedback: when shaking is detected, relevant event information is recorded and subsequent processing is triggered through system feedback; Event record: record the timestamp, duration, impact degree of the shaking; blur degree and the drop range of feature points; Early warning and triggering: Send shaking event information to the monitoring center or fleet management system through the wireless transmission module; trigger the anti-shake or deblurring algorithm to ensure that the camera image returns to clarity.
7. The container frame spreader large-amplitude shaking detection and early warning system according to claim 1 is characterized by: The feature point coordinate estimation module of the container frame spreader uses the OpenVINS system to track the two-dimensional feature points of the container frame spreader and further estimates its three-dimensional coordinates in the real world coordinate system to construct the input tensor The characteristic information of is used for the sway prediction of the model; the dynamic state quantity Indicates the state of the feature point at time t, i is the index number of the feature point, and the feature point state contains m events. represents the mth event; Initialize the distinction between the car frame and the background: Initial frame annotation: In the first frame, manually annotate the polygonal area of the trolley frame, denoted as , the rest of the area is the background, denoted as ; Dynamic classification: Based on the initial polygonal area, the feature points tracked by OpenVINS will be classified as small frame points or background points for subsequent analysis; for each feature point , to determine whether it is located in Inside: ;in is the feature point classification function, For the trolley frame, are background points; these feature points are further divided into two categories: polygonal areas of container frame spreaders The feature points within are recorded as ; Other non-container frame spreader areas Assume that several feature points have been marked on the ground as the origin of the real-world coordinates. and anchor points , recorded as ; d1. 2D tracking of feature points: First, OpenVINS is used to perform visual inertial synchronous positioning and map construction on the container frame spreader to achieve real-time tracking and matching of feature points in the 2D image; Feature point extraction: divided into two parts: polygonal area of container frame spreader Use feature points within As the feature point for initial tracking; in the area of non-container frame spreader In the example above, the feature extraction module in OpenVINS is used to extract a set of key feature points from continuous video frames: Each feature point Contains 2D image coordinates and its descriptors; Feature matching and tracking: Use the optical flow method or feature descriptor matching algorithm to establish the correspondence between feature points between consecutive frames and generate a matching set: ;in represents the feature point matching set, Represents the feature points of the current frame, Indicates the feature points of the previous frame; d2. Use OpenVINS to estimate 3D coordinates; after obtaining consistent 2D feature point matches in consecutive frames, use the visual inertial fusion capability of OpenVINS to convert 2D feature points into coordinate points in 3D space ; (1) Camera intrinsic and extrinsic calibration: Camera intrinsic calibration: Camera intrinsic parameters describe the internal geometry and optical properties of the camera, including: camera intrinsic matrix; camera intrinsic matrix Describes the intrinsic geometric properties of the camera in the form of: ; in, and are the focal lengths in the x and y directions in pixel units respectively; and are the base point coordinates; is a parameter describing the pixel tilt, assumed to be 0; Distortion parameters; The camera lens introduces radial distortion and tangential distortion, which need to be corrected by the distortion model: Radial distortion correction: ; Tangential distortion correction: ;in, , is the radial distortion coefficient; is the tangential distortion coefficient; Camera intrinsic calibration method: The specific steps are as follows: Step 1: Prepare a calibration plate with a known geometric shape; Step 2: Take multiple images of the calibration plate from different angles; Step 3: Detect the corner points of the calibration plate in the image; Step 4: Establish the corresponding relationship between the image coordinates of the corner points and the actual coordinates on the calibration plate; Step 5: Solve the camera intrinsic parameters and distortion parameters by minimizing the reprojection error; The reprojection error is defined as: ;in is the error, is the point detected on the image, The camera extrinsic parameters define the spatial relationship between the camera coordinate system and the world coordinate system, including: rotation matrix ,3×3 rotation matrix, describes the rotation of the camera coordinate system relative to the world coordinate system. The rotation matrix is expressed in Euler angles, quaternions or axis angles; for example, the basic rotation matrices around the x, y, and z axes are , , , which is expressed as follows: , , ;in is the rotation angle; The complete rotation matrix Obtained through the combination of basic rotation matrices; Translation Vector ; A 3×1 vector describing the position of the origin of the camera coordinate system in the world coordinate system: ; External parameter calibration method: The specific steps are as follows: Step 1: For static scenes: use markers or feature objects with known 3D coordinates; Step 2: For dynamic systems: estimate the camera pose through camera motion; first use feature matching between consecutive frames; then use the basic matrix or essential matrix for pose estimation; finally combine IMU data for a more robust estimate; In OpenVINS, extrinsic parameters are dynamically estimated by fusing visual and inertial information: first, the relative motion of the camera is calculated using visual feature tracking; then the acceleration and angular velocity measurements are obtained from the IMU; and the visual and inertial data are fused through the extended Kalman filter; finally, the camera's extrinsic parameters are continuously optimized and updated; (2) Depth estimation and triangulation: Depth estimation is a key step in converting 2D image feature points into 3D spatial points. In the OpenVINS system, the depth of feature points is accurately estimated using triangulation methods through camera motion and feature point matching between consecutive frames. Mathematical modeling; suppose there is a 3D point , Observed at two camera positions, the image coordinates are obtained and ; Each camera position has its own extrinsic parameters and ; Projection equation; the camera projection model is expressed as: ; in are homogeneous image coordinates, is the homogeneous world coordinate, is the scale factor depth; Linear triangulation; for two views, we have: ; The equations are rewritten as a linear system: ;here is a matrix consisting of camera parameters and image coordinates, is the corresponding constant term; Expanding in detail, the following linear equation is constructed: For the first camera: ; For the second camera: ;in, Indicates The rotation matrix of the camera OK, is the translation vector elements; Nonlinear optimization; Due to measurement errors and numerical limitations; Define the reprojection error: ;in Is the 3D point Function projected onto the image; by minimizing the reprojection error, more accurate 3D point coordinates are obtained; From 2D feature points To 3D point The complete transformation process is as follows: Step 1: For feature points , first apply the inverse transformation of the camera intrinsic parameter matrix to obtain the normalized coordinates: ;in is the normalized coordinate, is the inverse matrix of the camera intrinsic parameters; Step 2: Use multi-frame observation data and camera pose information to perform triangulation to obtain the 3D point coordinates in the camera coordinate system ;in is the three-dimensional coordinate; Step 3: Convert the 3D point in the camera coordinate system to the world coordinate system: ;in, is the 3D point position in the world coordinate system, is the rotation transformation matrix, is the translation transformation matrix; in OpenVINS, this process is further enhanced by visual-inertial fusion, using IMU data to provide additional motion constraints and scale information; (3) Coordinate system conversion; Coordinate system definition: In the OpenVINS system, the following coordinate systems are involved: Image coordinate system: 2D plane coordinate system, the origin is at the upper left corner of the image, the unit is pixel; the coordinate is expressed as: ; The x-axis points to the right, and the y-axis points downward; Normalized camera coordinate system: a coordinate system that removes the influence of camera intrinsic parameters; coordinate representation: ; The optical center of the camera is at the origin, and the z-axis is along the positive direction of the camera optical axis; Camera coordinate system: 3D coordinate system, the origin is at the camera optical center; coordinate representation: ; The x-axis points to the right, the y-axis points downward, and the z-axis points forward; World coordinate system: global reference coordinate system, all camera poses and 3D points are expressed relative to this coordinate system; coordinate representation: ;Select the initial camera position or reference point as the origin; d3. Build input tensor feature information; by obtaining the 3D position of each feature point in the world coordinate system Right now , combining the dynamic state and historical shaking events provided by the shaking detection module to construct the input tensor for shaking prediction ; Dynamic state quantity: using the current moment And the most recent big shake moment 3D coordinates, calculate the dynamic state: ;in This is when the most recent big shake occurred. is the feature point index; at the current time t, is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; the most recent large shake At this moment, is the position of the i-th feature point on the x-axis, is the position of the i-th feature point on the y-axis, is the position of the i-th feature point on the z-axis; Historical shaking events: Collect recent Shaking event , where each event Records what happened at the time The sway dynamic state quantity; Input tensor construction: Integrate the current dynamic state with the historical shaking events to form a complete input tensor: ; This input tensor It will be used as input features of the slosh prediction module for training and predicting future slosh events.
8. A container frame spreader large-amplitude shaking detection and early warning system according to claim 7, characterized in that: The shaking warning module predicts whether a large shaking will occur in the future based on historical shaking data; specifically, it includes the following steps: Step 1: Construction of input features; based on the input tensor , constructing effective features, including the following parts: (1) Dynamic state quantity ; (2) Shaking event sequence ; (3) Shaking displacement data: average displacement of two shaking events; average horizontal displacement: ; Average vertical displacement: ; The average displacement from the current moment to the most recent shake: Average horizontal displacement: ; Average vertical displacement: ;in, Respectively represent the horizontal coordinates of the characteristic points of the shaking event at time Tm, Tm-1, and t; Respectively represent the vertical coordinates of the characteristic points of the shaking event at time Tm, Tm-1, and t; Step 2: Time series analysis and feature extraction; Shaking events have obvious time series characteristics; By analyzing the displacement pattern of historical shaking data, the regularity of shaking can be captured, and the occurrence of future shaking can be predicted; Long short-term memory network; LSTM is a recurrent neural network suitable for processing and predicting time series data, which can capture long-term dependencies; By constructing an LSTM network and combining historical shaking data, it is possible to predict whether a large-scale shaking will occur within the next 1 second; Step 3, model architecture; (1) Input layer: accepts the constructed feature vector, including dynamic state quantity, historical shaking events and displacement data; (2) LSTM layer: first layer LSTM: number of units: 128; return sequence: True; activation function: tanh; Dropout: 0.2; second layer LSTM: number of units: 64; return sequence: False; activation function: tanh; Dropout: 0.2; (3) Fully connected layer: first layer fully connected: number of neurons: 64; activation function: ReLU; Dropout: 0.2; (4) Output layer: output predicted horizontal and vertical displacement differences dx, dy and dz; output explanation dx and dy: describe the changes in the horizontal direction x and y of the next large horizontal shaking relative to the previous large shaking; dx represents the displacement difference in the x direction, and dy represents the displacement difference in the y direction; when the straight track on the ground is taken as the x-axis, y defaults to 0; dz: describes the changes in the vertical direction z of the next large vertical shaking relative to the previous large shaking.
9. A container frame spreader large-amplitude shaking detection and early warning system according to claim 8, characterized in that: Model training and prediction; (1) Data preprocessing: De-noising: Apply moving average filtering or denoising methods to smooth 3D coordinate data and reduce the impact of measurement noise; Features Standardization: Normalize the displacement data to ensure that all features are input into the model at the same scale; (2) Training process: Label definition: Based on historical data, mark whether shaking will occur in the next 1 second; Loss function: mean square error; Optimization algorithm: Use Adam optimizer, and set the learning rate to 0.001; Regularization measures: Dropout: Add Dropout layer in LSTM layer and fully connected layer to prevent overfitting; L2 regularization: Introduce L2 regularization in Dense layer to limit the weight size; Data segmentation: Training set: 70%; Validation set: 15% Test set: 15%; (3) Model prediction: Input the characteristic data of the current time and historical shaking events, and the model outputs the predicted ; Output explanation: dx and dy describe the change in the horizontal direction (x, y) of the next large horizontal sway relative to the previous large sway; dz describes the change in the vertical direction relative to the previous large sway; Output decision: No significant change: When and and When , it means that the next shaking has no significant change from the previous one; Indicates the feature thresholds in three directions; Significant change: When any output value exceeds the corresponding threshold, it means that the next shake will change significantly and anti-shake measures need to be taken.
10. The container frame spreader large-amplitude shaking detection and early warning system according to claim 1, characterized in that: The model training and updating mechanism module includes the following steps: Step 1: Data labeling and collection: Shake detection module; real-time detection and recording of shake events, including the timestamp, duration, blur level, and feature point drop amplitude of the shake; data storage; storing the detected shake event data in the database as a label for training data; Step 2: Model training: Periodic training: Regularly retrain the prediction model using the latest collected shaking event data; Incremental learning: Using incremental learning technology, gradually add new data to the training set and continuously optimize model parameters; Training process: Data preparation: extract the latest shaking event data from the database to build training sets and validation sets; Data preprocessing: perform data normalization, missing value processing and data enhancement to ensure the quality of training data; Model training: use the designed LSTM model architecture to train the model; Model validation: evaluate model performance on the validation set, adjust model parameters and structure, and optimize performance; Model preservation and deployment: save the best trained model and deploy it to the real-time prediction system; Step 3: Online update and adjustment: Real-time feedback: Dynamically adjust model parameters and thresholds based on actual prediction results and subsequent verification to improve prediction accuracy; Adaptive adjustment: Use reinforcement learning technology to achieve adaptive adjustment of the model to environmental changes and improve the robustness and stability of the system; Step 4. Model evaluation and optimization: Evaluation indicators: Mean square error: measures the average square error between the predicted value and the true value; Mean absolute error: measures the average absolute error between the predicted value and the true value; R²-score: measures the ability of the model to explain the variance of the data; Model tuning: Optimize model performance by adjusting the following parameters: Number of LSTM units: increase or decrease the number of units in the LSTM layer to find the optimal model complexity; Number of time steps: adjust the number of time steps of the input sequence to observe the impact on model performance; Learning rate: optimize the learning rate of the Adam optimizer to ensure rapid convergence of the model; Regularization parameters: adjust the Dropout rate and L2 regularization strength to prevent overfitting; Early stopping mechanism: During the training process, the early stopping mechanism is used. When the performance of the validation set no longer improves, the training is stopped in advance to prevent the model from overfitting. Step 5. Continuous learning and optimization: Data enhancement: Generate more data with shaking through data enhancement technology to improve the generalization ability of the model; Abnormal data processing: Reasonably handle noise data and missing data to improve the adaptability of the model to actual working conditions; Closed-loop feedback mechanism: Realize the closed loop of detection-early warning-optimization to ensure the adaptive adjustment and continuous and stable operation of the system.