Crane remote control image delay detection method and system

Through the combination of edge computing and multimodal deep learning, the image delay in the crane remote control system is accurately detected and predicted, which solves the problem of difficult estimation of video streaming delay and achieves more efficient and safe remote operation.

CN120378603AActive Publication Date: 2025-07-25NINGBO SPECIAL EQUIP INSPECTION & RES INST

Patent Information

Application Number
CN202510653134.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-25
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the existing crane remote control system, video streaming delay is difficult to accurately estimate, especially when the network conditions are complex and changeable, resulting in poor delay compensation effect, affecting operation accuracy and safety.

Method used

The edge computing server is used to perform dual-channel parallel preprocessing, combining attention mechanism and adaptive Kalman filtering algorithm to extract the three-dimensional spatial coordinates and motion parameters of the crane, predict the video streaming delay through the multi-modal deep learning model, optimize the encoding parameters with the model prediction control algorithm, and transmit the video stream through the intelligent routing mechanism, and finally generate a delay compensation control strategy through the hierarchical reinforcement learning algorithm.

Benefits of technology

Accurately detect and predict image delays, improve the real-time and accuracy of remote control, enhance the robustness of remote control systems, reduce the impact of operation delays on lifting operations, and improve work efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378603A_ABST
    Figure CN120378603A_ABST
Patent Text Reader

Abstract

The invention provides a crane remote control image delay detection method and system, and relates to the technical field of cranes, and the method comprises the following steps: carrying out space-time segmentation by using a dual-channel parallel preprocessing mechanism; acquiring three-dimensional space coordinates and motion parameters of the crane by utilizing an attention mechanism; calculating a motion state vector, inputting the motion state vector into the multi-modal deep learning model, and generating a video stream delay prediction curve; based on a prediction curve and a model prediction control algorithm, optimizing key frame distribution and coding parameters of the video stream, and transmitting the key frame distribution and coding parameters to a remote control terminal through an intelligent routing mechanism; modeling by adopting a recursive least square algorithm to obtain a probability distribution model; and generating a delay compensation control strategy through a hierarchical reinforcement learning algorithm, and carrying out adaptive time compensation on the operation instruction. According to the invention, the remote control image transmission delay is effectively reduced, and the real-time performance and safety of remote operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to crane technology, and in particular to a method and system for detecting image delay in remote control of cranes. Background Art

[0002] The remote control technology of cranes plays an increasingly important role in modern construction. It can overcome the limitations of traditional manual operation methods and improve construction efficiency and safety. The remote control system usually relies on real-time video stream transmission, and the operator performs remote operations based on video image information. However, due to the complexity and uncertainty of network transmission, problems such as delay and jitter are likely to occur during the video stream transmission, which seriously affects the accuracy and safety of remote operations and may even lead to safety accidents.

[0003] To solve the problem of video stream transmission delay, the existing technologies mainly adopt some delay compensation and prediction methods. For example, some methods pre-compensate the control instructions by estimating the network delay, but these methods are usually difficult to accurately estimate the delay and the compensation effect is limited. Other methods use simple linear models to predict the delay, but the actual network environment is complex and changeable, and the linear model is difficult to adapt to the dynamic network conditions. In addition, most of the existing methods only focus on the network transmission delay and ignore the delay introduced by video encoding, decoding, processing and other links, resulting in poor overall delay compensation effect.

[0004] Inaccurate delay estimation: The existing delay estimation methods are difficult to accurately estimate the network transmission delay, especially in the case of complex and changeable network conditions, resulting in poor delay compensation effect. Too simple prediction model: Most of the existing delay prediction models use simple linear models, which are difficult to adapt to the dynamic network conditions and have low prediction accuracy. Ignoring other delay factors: Most of the existing methods only focus on the network transmission delay and ignore the delay introduced by video encoding, decoding, processing and other links, resulting in poor overall delay compensation effect. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for detecting image delay in remote control of cranes, which can solve the problems in the existing technology.

[0006] In the first aspect of the embodiments of the present invention, A method for detecting image delay in remote control of cranes is provided, including: Transmit the real-time video stream data to the edge computing server local to the crane; the edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; use the attention mechanism to extract features from the video frame sequence to obtain the three-dimensional space coordinates and motion parameters of the crane; based on the three-dimensional space coordinates and motion parameters, calculate the motion state vector through the adaptive Kalman filtering algorithm; Input the motion state vector into the multi-modal deep learning model. The multi-modal deep learning model extracts the dynamic features of the video stream through the temporal graph convolutional network, deeply fuses the dynamic features with the network transmission quality metrics collected in real time, and generates a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, combine the model predictive control algorithm to calculate the optimal adjustment strategy of the video coding parameters; optimize the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmit the optimized video stream to the remote control terminal through the intelligent routing mechanism; Calculate the time difference between the arrival time and the sending time of the video stream, use the recursive least squares algorithm to model the time difference to obtain a probability distribution model; perform error analysis on the probability distribution model and the video stream delay prediction curve, and generate a delay compensation control strategy through the hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

[0007] The edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; use the attention mechanism to extract features from the video frame sequence to obtain the three-dimensional space coordinates and motion parameters of the crane; based on the three-dimensional space coordinates and motion parameters, the motion state vector calculated through the adaptive Kalman filtering algorithm includes: The edge computing server uses a dual-channel parallel preprocessing mechanism to process the real-time video stream data; the time channel establishes a temporal feature vector for the real-time video stream data, and the space channel divides the real-time video stream data into multiple regions of interest and generates a spatial feature vector; fuse the temporal feature vector and the spatial feature vector to generate a video frame sequence with timestamps; Construct a multi-head attention network for the video frame sequence for feature extraction. The multi-head attention network includes a spatial attention module and a channel attention module; based on the multi-head attention network, calculate the mapping relationship of feature points in three-dimensional space to obtain the three-dimensional space coordinates and motion parameters of the crane; Construct an adaptive Kalman filter model based on the three-dimensional space coordinates and motion parameters, calculate the measurement reliability index based on the adaptive Kalman filter model and dynamically adjust the measurement noise matrix; analyze the state prediction error based on the adaptive Kalman filter model to obtain a new process noise matrix; Input the measurement noise matrix and the process noise matrix into the state prediction equation, and use the state transition matrix in the state prediction equation to predict the motion state of the crane; input the prediction result into the measurement update equation, and the measurement update equation corrects the predicted state based on the adaptive Kalman filter model; calculate the error covariance according to the corrected state, and optimize the state estimation result based on the error covariance to output a motion state vector including position parameters, speed parameters, and acceleration parameters.

[0008] Construct an adaptive Kalman filter model based on three-dimensional space coordinates and motion parameters; the adaptive Kalman filter model establishes a measurement noise statistical model, calculates the measurement reliability index based on the measurement noise statistical model, and dynamically adjusts the measurement noise matrix according to the measurement reliability index; the adaptive Kalman filter model analyzes the state prediction error to obtain system uncertainty parameters, and updates the process noise matrix based on the system uncertainty parameters, including: Construct an adaptive Kalman filter model based on the three-dimensional space coordinates and motion parameters of the crane; input the three-dimensional space coordinates and motion parameters into the adaptive Kalman filter model to generate an observation data sequence; The adaptive Kalman filter model establishes a measurement noise statistical model based on the observation data sequence, and the measurement noise statistical model calculates the residual vector and variance distribution of the observation data; constructs a measurement reliability index based on the residual vector and variance distribution, and the measurement reliability index uses an exponential function form to characterize the credibility of the observation data; sets an adaptive weight factor according to the credibility, and dynamically adjusts the measurement noise matrix based on the adaptive weight factor; The adaptive Kalman filter model performs error analysis on the state prediction result and the actual observation value, and calculates the statistical characteristics of the innovation sequence; evaluates the system dynamic characteristics based on the statistical characteristics of the innovation sequence to obtain a parameter vector characterizing system uncertainty; calculates the deviation degree of the system model according to the parameter vector, and sets a fading factor based on the deviation degree; uses the fading factor to recursively update the process noise matrix.

[0009] Based on the change trend of the video stream delay prediction curve, combine the model predictive control algorithm to calculate the optimal adjustment strategy of the video coding parameters; optimize the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmit the optimized video stream to the remote control terminal through the intelligent routing mechanism, including: Perform dynamic analysis on the video stream delay prediction curve, extract the fluctuation trend, change gradient, and mutation characteristics from the video stream delay prediction curve. The fluctuation trend characterizes the stability of the network state, the change gradient reflects the change rate of the network performance, and the mutation characteristics indicate the drastic change points of the network state; input the fluctuation trend, change gradient, and mutation characteristics into the model predictive control algorithm based on a sliding time window. The model predictive control algorithm outputs the optimal adjustment plan of the coding parameters by constructing a coupling constraint model of video quality and transmission delay; Adjust the fluctuation trend, change gradient, and mutation characteristics based on the optimal adjustment scheme. When the fluctuation trend indicates a stable network state, increase the key frame interval to improve the encoding efficiency; when the mutation characteristics indicate a drastic fluctuation in network performance, decrease the key frame interval to enhance the anti-interference ability; adjust the quantization parameter according to the positive or negative direction of the change gradient. When the change gradient is positive, increase the quantization step size to reduce the bit rate, and when the change gradient is negative, decrease the quantization step size to improve the quality; calculate the real-time adjustment amount of the target bit rate based on the amplitude of the fluctuation trend. Update the adjusted fluctuation trend, change gradient, and mutation characteristics to the video encoder to generate an optimized video bitstream; calculate the path priority based on the intelligent routing mechanism in combination with the delay status of the transmission node and the link bandwidth resources; send the optimized video bitstream to the remote control terminal through the transmission path with the highest path priority.

[0010] Update the adjusted fluctuation trend, change gradient, and mutation characteristics to the video encoder to generate an optimized video bitstream; based on the network state information of the delay prediction curve, select the optimal transmission path for the optimized video bitstream through the intelligent routing mechanism. The intelligent routing mechanism combines the delay status of the transmission node and the link bandwidth resources to calculate the path priority, including: Input the fluctuation trend, change gradient, and mutation characteristics into the video encoder for update; the video encoder processes the video bitstream based on the updated encoding parameters to generate an optimized video bitstream. Collect the processing delay generated when the network transmission node processes the optimized video bitstream, and record the queuing delay of the data packet in the node buffer; add the processing delay of a single node to the queuing delay to obtain the total node delay; accumulate the total node delays of each node along the transmission path in sequence to obtain the node delay status of the transmission path; divide the node delay status by the preset delay threshold to obtain the node delay score. Monitor the network link status carrying the optimized video bitstream, and obtain the total bandwidth capacity and the occupied bandwidth of the link; calculate the available bandwidth value of each link; select the link with the smallest available bandwidth in the transmission path as the path bandwidth resource; divide the path bandwidth resource by the current bandwidth requirement of the video bitstream to obtain the bandwidth resource score. Weight the node delay score and the bandwidth resource score. Normalize the node delay score and multiply it by the weight coefficient of 0.7, normalize the bandwidth resource score and multiply it by the weight coefficient of 0.3, and add the two to obtain the path priority score.

[0011] Calculate the time difference between the arrival time and the transmission time of the video stream, model the time difference using the recursive least squares algorithm to obtain a probability distribution model; conduct error analysis on the probability distribution model and the video stream delay prediction curve, and generate a delay compensation control strategy through the hierarchical reinforcement learning algorithm; the remote control terminal performs adaptive time compensation on the operation instructions of the crane according to the delay compensation control strategy, including: Calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, and the probability distribution model characterizes the statistical characteristics of the video stream delay data; Generate a video stream delay prediction curve according to the probability distribution model, and the video stream delay prediction curve reflects the change trend of the video stream delay data; compare and analyze the change trend with the actually measured delay data to calculate the prediction error; construct an error evaluation function based on the prediction error; Take the error evaluation function as the optimization target of the hierarchical reinforcement learning algorithm, and generate a delay compensation control strategy using the hierarchical reinforcement learning algorithm; the delay compensation control strategy dynamically adjusts the execution timing of the control instruction based on the prediction result of the video stream delay data; The remote control terminal receives the operation instructions of the operator for the crane, calculates the time compensation value of the operation instructions according to the delay compensation control strategy; performs adaptive delay compensation on the operation instructions based on the time compensation value.

[0012] Calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, and the probability distribution model characterizes the statistical characteristics of the video stream delay data, including: Calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; perform outlier detection and normalization processing on the video stream delay data to eliminate clock synchronization errors; Construct a recursive least squares algorithm model, input the video stream delay data into the recursive least squares algorithm model, calculate the prediction error and the gain matrix; update the parameter estimation value based on the prediction error and the gain matrix; optimize the covariance matrix according to the parameter estimation value to achieve iterative optimization of the model parameters; Analyze the statistical law of the video stream delay data based on the calculation results of the recursive least squares algorithm model, calculate the mean, variance, skewness and kurtosis of the video stream delay data; select the optimal probability distribution function to fit the video stream delay data, estimate the distribution function parameters; construct a probability distribution model based on the optimal probability distribution function, and the probability distribution model quantitatively characterizes the statistical characteristics of the video stream delay data.

[0013] In the second aspect of the embodiments of the present invention, a crane remote control image delay detection system is provided, including: A first unit for transmitting real-time video stream data to an edge computing server local to the crane; the edge computing server performs spatio-temporal segmentation on the real-time video stream data using a dual-channel parallel preprocessing mechanism to obtain a video frame sequence with timestamps; uses an attention mechanism to extract features from the video frame sequence to obtain the three-dimensional spatial coordinates and motion parameters of the crane; based on the three-dimensional spatial coordinates and motion parameters, calculates a motion state vector through an adaptive Kalman filtering algorithm; A second unit for inputting the motion state vector into a multi-modal deep learning model. The multi-modal deep learning model extracts dynamic features of the video stream through a temporal graph convolutional network, deeply fuses the dynamic features with the network transmission quality metrics collected in real time to generate a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, combines a model predictive control algorithm to calculate an optimal adjustment strategy for video coding parameters; optimizes the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmits the optimized video stream to a remote control terminal through an intelligent routing mechanism; A third unit for calculating the time difference between the arrival time and the sending time of the video stream, modeling the time difference using a recursive least squares algorithm to obtain a probability distribution model; performing error analysis on the probability distribution model and the video stream delay prediction curve, and generating a delay compensation control strategy through a hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

[0014] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0015] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0016] The beneficial effects of this application are as follows: By combining edge computing and multi-modal deep learning to process video stream data, the present invention can accurately detect and predict image delays in the remote control process, improve the real-time performance and accuracy of remote control, and effectively solve the problem of operation lag caused by network delays in traditional remote control systems.

[0017] The present invention adopts an algorithm combining adaptive Kalman filtering and temporal graph convolutional network to achieve accurate modeling and prediction of the crane's motion state, and dynamically adjusts video encoding parameters through a model predictive control algorithm, significantly improving the quality and stability of video transmission in a harsh network environment and enhancing the robustness of the remote control system.

[0018] The present invention introduces a hierarchical reinforcement learning algorithm to generate a delay compensation control strategy, realizing adaptive time compensation for remote control instructions, greatly reducing the impact of operation delay on lifting operations, improving operation efficiency and safety, and at the same time reducing the psychological burden of operators, making the remote control experience closer to on-site operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic flow chart of the crane remote control image delay detection method according to an embodiment of the present invention; Figure 2 It is a table chart for analyzing adaptive Kalman filter state prediction parameters and comparing system performance according to an embodiment of the present invention; Figure 3 It is a schematic diagram for comprehensively evaluating video quality and transmission efficiency under different network environments according to an embodiment of the present invention; Figure 4 It is a schematic diagram for comprehensively evaluating the comprehensive performance of video bitstream optimization and intelligent routing mechanism according to an embodiment of the present invention; Figure 5 It is a schematic diagram for analyzing and comparing the accuracy of the delay probability distribution model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 It is a schematic flow chart of the crane remote control image delay detection method according to an embodiment of the present invention, as Figure 1 shown, the method includes: Transmit the real-time video stream data to the edge computing server local to the crane; the edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; use the attention mechanism to extract features from the video frame sequence to obtain the three-dimensional spatial coordinates and motion parameters of the crane; based on the three-dimensional spatial coordinates and motion parameters, calculate the motion state vector through the adaptive Kalman filter algorithm; Input the motion state vector into the multi-modal deep learning model. The multi-modal deep learning model extracts the dynamic features of the video stream through the temporal graph convolutional network, deeply fuses the dynamic features with the network transmission quality metrics collected in real time, and generates a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, calculate the optimal adjustment strategy of the video coding parameters in combination with the model predictive control algorithm; optimize the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmit the optimized video stream to the remote control terminal through the intelligent routing mechanism; Calculate the time difference between the arrival time and the sending time of the video stream, use the recursive least squares algorithm to model the time difference to obtain a probability distribution model; perform error analysis on the probability distribution model and the video stream delay prediction curve, and generate a delay compensation control strategy through the hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

[0023] In an alternative embodiment, the edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; use the attention mechanism to extract features from the video frame sequence to obtain the three-dimensional spatial coordinates and motion parameters of the crane; based on the three-dimensional spatial coordinates and motion parameters, the motion state vector calculated through the adaptive Kalman filter algorithm includes: The edge computing server processes the real-time video stream data using a dual-channel parallel preprocessing mechanism; the time channel establishes a temporal feature vector for the real-time video stream data, and the space channel divides the real-time video stream data into multiple regions of interest and generates spatial feature vectors; fuse the temporal feature vector and the spatial feature vector to generate a video frame sequence with timestamps; Construct a multi-head attention network for the video frame sequence for feature extraction. The multi-head attention network includes a spatial attention module and a channel attention module; based on the multi-head attention network, calculate the mapping relationship of feature points in three-dimensional space to obtain the three-dimensional spatial coordinates and motion parameters of the crane; Construct an adaptive Kalman filter model based on the three-dimensional spatial coordinates and motion parameters, calculate the measurement reliability index based on the adaptive Kalman filter model and dynamically adjust the measurement noise matrix; analyze the state prediction error based on the adaptive Kalman filter model to obtain a new process noise matrix; Input the measurement noise matrix and the process noise matrix into the state prediction equation, which uses the state transition matrix to predict the motion state of the crane; input the prediction result into the measurement update equation, which corrects the predicted state based on the adaptive Kalman filter model; calculate the error covariance according to the corrected state, and optimize the state estimation result based on the error covariance to output the motion state vector including position parameters, speed parameters, and acceleration parameters.

[0024] The edge computing server processes the real-time video stream data using a dual-channel parallel preprocessing mechanism. This preprocessing mechanism includes two parallel processing paths: the time channel and the space channel.

[0025] The time channel is mainly responsible for establishing a temporal feature vector for the real-time video stream data. Specifically, the time channel first groups consecutive video frames according to a fixed time window (e.g., every 8 frames), then calculates the pixel change amount between adjacent frames through the inter-frame difference algorithm to generate a change intensity map. Then, apply a 3×3 Gaussian filter to the change intensity map for smoothing and extract the continuity features in the time dimension. Finally, the time channel integrates these features into a 128-dimensional temporal feature vector to represent the temporal evolution features of the video content.

[0026] The space channel divides the real-time video stream data into multiple regions of interest and generates a spatial feature vector. First, use an improved YOLOv5 object detection algorithm to identify the main components of the crane and divide the video frame into 9 main regions of interest. Then, scan each region of interest with a 16×16 pixel sliding window to extract local texture features and edge features. Subsequently, extract the depth features of each region through an improved VGG-16 network and finally fuse them into a 256-dimensional spatial feature vector.

[0027] The temporal feature vector and the spatial feature vector are integrated through a feature fusion module, which uses an attention weighting mechanism to calculate the relative importance of the two types of features and generates a 384-dimensional fused feature vector. Finally, associate the fused feature vector with the corresponding timestamp (accurate to the millisecond level) to form a video frame sequence with timestamps, providing basic data for subsequent feature extraction. In practical applications, the average processing delay of this preprocessing stage is less than 30 milliseconds, meeting the real-time requirement.

[0028] Construct a multi-head attention network for feature extraction on the video frame sequence obtained by preprocessing. This network includes two core components: a spatial attention module and a channel attention module.

[0029] The spatial attention module first applies max - pooling and average - pooling operations to the input video - frame feature map to generate two independent feature maps. These two feature maps are concatenated and then processed through a 7×7 convolutional layer to generate a spatial attention weight map, which is used to highlight the importance of spatial positions in the feature map. In practical applications, this module can effectively identify the positions of the key components of the crane, focusing attention on important parts such as the hook, boom, and counterweight, with an accuracy rate of 94.5%.

[0030] The channel attention module adopts the squeeze - and - excitation strategy. It compresses the feature map into a channel descriptor through global average pooling, and then uses a bottleneck structure formed by two fully - connected layers to learn the dependencies between channels, generating channel attention weights. This module can highlight the feature channels most relevant to crane pose recognition and enhance the feature expression ability.

[0031] The multi - head attention network simultaneously learns representations in different sub - spaces through 8 parallel attention heads. Each attention head is responsible for capturing different types of dependencies in the video - frame sequence. The outputs of each attention head are integrated through a linear transformation layer to form the final feature representation.

[0032] Based on the extracted features, the system calculates the mapping relationship of feature points in three - dimensional space. Specifically, first, a modified corner - detection algorithm is used to identify the key feature points of the crane, and then the three - dimensional coordinates of these feature points are solved using the binocular vision principle and camera calibration parameters. To improve the accuracy, the system uses a deep - learning - assisted method to predict the initial depth value of the feature points and then adjusts it to the optimal solution through iterative optimization.

[0033] Through the above - mentioned processing, the system finally obtains the three - dimensional space coordinates (X, Y, Z) and motion parameters (including rotation angle, boom elevation angle, hook height, etc.) of the crane. The positioning accuracy can reach ±5 cm, and the angle accuracy can reach ±0.5 degrees.

[0034] Based on the obtained three - dimensional space coordinates and motion parameters, the system constructs an adaptive Kalman filter model for motion - state prediction and optimization.

[0035] Based on the adaptive Kalman filter model, a measurement reliability index is calculated. Specifically, the standard deviation of the measurement data for 5 consecutive frames is calculated to judge the dispersion degree of the measurement data. When the standard deviation is lower than the preset threshold (for example, 0.1 m for position parameters and 0.8 degrees for angle parameters), the measurement data is considered relatively reliable; otherwise, there may be abnormal measurements. The system dynamically adjusts the measurement noise matrix according to the measurement reliability, using a smaller measurement noise value (such as 0.01) when the measurement data is reliable and a larger measurement noise value (such as 0.2) when the measurement data fluctuates greatly to reduce the impact of unreliable measurements.

[0036] The process noise matrix is obtained by analyzing the state prediction error based on the adaptive Kalman filter model. The system calculates the difference between the continuous 10 state estimates and the actual measured values, and establishes a sliding window to track the prediction performance. When the prediction error gradually increases, the value of the process noise matrix is appropriately increased (for example, from 0.005 to 0.02) to enable the model to adapt to state changes more quickly; when the prediction error is stable or decreasing, the value of the process noise matrix is appropriately decreased to improve the filtering smoothness.

[0037] The dynamically adjusted measurement noise matrix and process noise matrix are input into the state prediction equation. The state prediction equation uses the state transition matrix to predict the motion state of the crane, and the state vector includes parameters such as position, velocity, and acceleration. The state transition matrix is designed based on the motion physical model of the crane, considering the relationships between position and velocity, and velocity and acceleration.

[0038] The measurement update equation corrects the predicted state based on the adaptive Kalman filter model, comprehensively considering the predicted value and the measured value, and assigns different weights according to their respective uncertainties to obtain the optimal estimated value. Specifically, the system calculates the Kalman gain. When the reliability of the measurement data is high, the Kalman gain is increased to make the estimation result more inclined to the measured value; when the reliability of the measurement data is low, the Kalman gain is decreased to make the estimation result more inclined to the predicted value.

[0039] The system tracks the statistical characteristics of the prediction error and estimates the uncertainty of the current state estimate. For different motion parameters, such as rotation angle, pitch angle, and hook height, the corresponding error covariance is adjusted respectively to adapt to the change characteristics of different parameters.

[0040] Based on the error covariance to optimize the state estimation result, the system finally outputs a motion state vector including position parameters (X, Y, Z coordinates, accuracy ±3 cm), velocity parameters (velocities in three directions, accuracy ±0.05 m / s), and acceleration parameters (accelerations in three directions, accuracy ±0.01 m / s²). This motion state vector can be used in application scenarios such as the safety monitoring, collision warning, and abnormal behavior identification of the crane.

[0041] In an alternative embodiment, an adaptive Kalman filter model is constructed based on three-dimensional space coordinates and motion parameters; the adaptive Kalman filter model establishes a measurement noise statistical model, calculates a measurement reliability index based on the measurement noise statistical model, and dynamically adjusts the measurement noise matrix according to the measurement reliability index; the adaptive Kalman filter model analyzes the state prediction error to obtain system uncertainty parameters, and updates the process noise matrix based on the system uncertainty parameters, including: Construct an adaptive Kalman filter model based on the three-dimensional space coordinates and motion parameters of the crane; input the three-dimensional space coordinates and motion parameters into the adaptive Kalman filter model to generate an observation data sequence; The adaptive Kalman filter model establishes a measurement noise statistical model based on the observed data sequence. The measurement noise statistical model calculates the residual vector and variance distribution of the observed data; constructs a measurement reliability index based on the residual vector and variance distribution, and the measurement reliability index uses an exponential function form to characterize the credibility of the observed data; sets an adaptive weight factor according to the credibility, and dynamically adjusts the measurement noise matrix based on the adaptive weight factor; The adaptive Kalman filter model conducts error analysis on the state prediction result and the actual observed value, and calculates the statistical characteristics of the innovation sequence; evaluates the system dynamic characteristics based on the statistical characteristics of the innovation sequence to obtain a parameter vector representing the system uncertainty; calculates the deviation degree of the system model according to the parameter vector, and sets a fading factor based on the deviation degree; uses the fading factor to recursively update the process noise matrix.

[0042] Obtain the three-dimensional space coordinates and motion parameters of the crane. The three-dimensional space coordinates include the X coordinate, Y coordinate, and Z coordinate at the end of the crane boom, representing the position information in the east-west direction, north-south direction, and height direction respectively. The motion parameters include the speed and acceleration information in each direction. For example, at a certain moment t, the obtained coordinates are (15.2 m, 23.6 m, 45.8 m), and the corresponding speeds are (0.8 m / s, 1.2 m / s, 0.3 m / s).

[0043] Construct the obtained three-dimensional space coordinates and motion parameters into a state vector. The state vector contains position, speed, and acceleration information, forming a nine-dimensional vector. For example, the state vector can be expressed as: position X, position Y, position Z, speed X, speed Y, speed Z, acceleration X, acceleration Y, and acceleration Z.

[0044] Based on the state vector, establish a state transition matrix. The state transition matrix describes the evolution law of the system from the current state to the next state and is constructed using a physical kinematic model. In practical applications, it can be adjusted according to the motion characteristics of the crane.

[0045] By continuously obtaining multiple frames of observed data, an observed data sequence is formed. For example, continuously collect data for 20 seconds with a sampling frequency of 10 Hz to obtain 200 sets of three-dimensional coordinate and motion parameter data.

[0046] The residual vector is the difference between the observed value and the predicted value. For example, at a certain moment, the predicted position is (15.3 m, 23.5 m, 45.7 m), and the actual observed value is (15.2 m, 23.6 m, 45.8 m), then the residual vector is (-0.1 m, 0.1 m, 0.1 m).

[0047] The stability of the measurement can be evaluated by calculating the variances of the residual vectors in each dimension. For example, through statistical analysis, the variance in the X direction is 0.02 square meters, the variance in the Y direction is 0.03 square meters, and the variance in the Z direction is 0.01 square meters.

[0048] Based on the residual vectors and the variance distribution, a measurement reliability index is constructed. The measurement reliability index takes the form of an exponential function and is used to characterize the credibility of the observed data. In specific implementation, an exponential decay function can be used. When the residual is larger, the reliability index is smaller; conversely, when the residual is smaller, the reliability index is larger.

[0049] Suppose in a certain measurement, the residual in the X direction is -0.1 meters, and the corresponding variance is 0.02 square meters. Then, the reliability index in the X direction can be calculated as 0.95; the residual in the Y direction is 0.1 meters, and the corresponding variance is 0.03 square meters. The calculated reliability index in the Y direction is 0.92; the residual in the Z direction is 0.1 meters, and the corresponding variance is 0.01 square meters. The calculated reliability index in the Z direction is 0.90.

[0050] An adaptive weight factor is set according to the reliability index. When the reliability index is high, the weight of the observed data in the filtering process is increased; when the reliability index is low, the weight of the observed data is decreased. For example, the weight factor in the X direction is set to 0.95, the weight factor in the Y direction is set to 0.92, and the weight factor in the Z direction is set to 0.90.

[0051] The measurement noise matrix is dynamically adjusted based on the adaptive weight factor. The measurement noise matrix reflects the uncertainty in the measurement process. The measurement noise matrix is scaled by the weight factor to achieve adaptive adjustment. For example, the diagonal elements of the initial measurement noise matrix are [0.05, 0.05, 0.05], and after applying the weight factor, it is adjusted to [0.0525, 0.0543, 0.0556].

[0052] The difference between the state prediction value and the actual observed value is calculated to form an innovation sequence. For example, the innovation sequence in the X direction for 10 consecutive measurements is [-0.1, 0.05, -0.08, 0.12, -0.03, 0.07, -0.09, 0.04, -0.06, 0.11] meters.

[0053] The statistical characteristics of the innovation sequence are calculated, including the mean and covariance. For example, after calculation, the mean of the innovation sequence in the X direction is 0.003 meters, and the covariance is 0.0075 square meters; the mean in the Y direction is 0.002 meters, and the covariance is 0.0068 square meters; the mean in the Z direction is 0.001 meters, and the covariance is 0.0042 square meters.

[0054] Based on the statistical characteristics of the innovation sequence, the dynamic characteristics of the system are evaluated, and a parameter vector characterizing the system uncertainty is obtained. The parameter vector contains indicators reflecting the deviation between the system model and the actual system. For example, the calculated parameter vector is [0.15, 0.12, 0.08], corresponding to the degrees of system uncertainty in the X, Y, and Z directions respectively.

[0055] Calculate the deviation degree of the system model according to the parameter vector. The greater the deviation degree, the greater the gap between the system model and the actual system, and more reliance on observation data for correction is required. For example, according to the parameter vector [0.15, 0.12, 0.08], the calculated deviation degree is 0.35, indicating a medium degree of model uncertainty in the system.

[0056] Set the fading factor based on the deviation degree. The fading factor is used to control the influence degree of historical information on the current state estimation. The greater the deviation degree, the smaller the fading factor, indicating a reduced reliance on historical information. For example, when the deviation degree is 0.35, the fading factor is set to 0.85.

[0057] Use the fading factor to recursively update the process noise matrix. The process noise matrix reflects the uncertainty of the system model. Through weighted update with the fading factor, the process noise matrix can adapt to the dynamic changes of the system. For example, the diagonal elements of the initial process noise matrix are [0.02, 0.02, 0.02, 0.01, 0.01, 0.01, 0.005, 0.005, 0.005]. After applying the fading factor, it is updated to [0.023, 0.023, 0.023, 0.0115, 0.0115, 0.0115, 0.00575, 0.00575, 0.00575].

[0058] Using the updated measurement noise matrix and process noise matrix, perform the standard Kalman filter algorithm to obtain the optimal estimated values of the crane position and motion parameters. Through this method, the accuracy and stability of crane position tracking can be significantly improved, the influence of external interference and measurement errors can be reduced, and a reliable position reference can be provided for the safe operation of the crane.

[0059] Figure 2 For the embodiment of the present invention, the table graph for the analysis of adaptive Kalman filter state prediction parameters and the comparison of system performance is as follows: This table shows the performance comparison of different filtering schemes, including the specific performance indicators of four methods: standard Kalman filter, extended Kalman filter, robust Kalman filter, and the proposed technical solution. From the data, the proposed technical solution shows the best performance in multiple key indicators: the position RMSE is reduced to 0.41 meters, the velocity RMSE is only 0.17 m / s, and the convergence time is significantly shortened to 7.6 steps. Although its running time is slightly higher than that of the standard Kalman filter (2.9 ms / step vs 1.2 ms / step), it reaches the optimal level in terms of the anti-interference ability index (0.91) and the tracking stability index (0.94). The adaptability of this scheme to dynamic changes is rated as "high", and the adaptability to system uncertainty reaches 0.89, which are significantly better than the other three methods. In contrast, although the standard Kalman filter has the fastest operation speed, it performs poorly in terms of accuracy and stability; the extended Kalman filter and the robust Kalman filter show a gradual improvement trend in various performance indicators, but the overall performance is still inferior to the proposed technical solution. These data fully demonstrate that the proposed technical solution significantly improves the accuracy, stability, and adaptability of the system while ensuring the computational efficiency, and it is a more advanced filtering solution.

[0060] In an alternative embodiment, based on the change trend of the video stream delay prediction curve, combined with the model predictive control algorithm to calculate the optimal adjustment strategy of the video coding parameters; optimize the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmit the optimized video stream to the remote control terminal through the intelligent routing mechanism, including: Conduct dynamic analysis on the video stream delay prediction curve, extract the fluctuation trend, change gradient, and mutation characteristics from the video stream delay prediction curve. The fluctuation trend characterizes the stability of the network state, the change gradient reflects the change rate of the network performance, and the mutation characteristics indicate the points of drastic change in the network state; input the fluctuation trend, change gradient, and mutation characteristics into the model predictive control algorithm based on the sliding time window. The model predictive control algorithm outputs the optimal adjustment plan of the coding parameters by constructing a coupling constraint model of video quality and transmission delay; Adjust the fluctuation trend, change gradient, and mutation characteristics based on the optimal adjustment plan. When the fluctuation trend indicates a stable network state, increase the key frame interval to improve the coding efficiency; when the mutation characteristics indicate a drastic fluctuation in the network performance, reduce the key frame interval to enhance the anti-interference ability; adjust the quantization parameter according to the positive and negative directions of the change gradient. Increase the quantization step size to reduce the code rate when the change gradient is positive, and decrease the quantization step size to improve the quality when the change gradient is negative; calculate the real-time adjustment amount of the target code rate based on the amplitude of the fluctuation trend; Update the adjusted fluctuation trend, change gradient, and mutation characteristics to the video encoder to generate an optimized video bitstream; calculate the path priority based on the intelligent routing mechanism in combination with the delay status of the transmission nodes and the link bandwidth resources; send the optimized video bitstream to the remote control terminal through the transmission path with the highest path priority.

[0061] Conduct a dynamic analysis of the video stream delay prediction curve to extract the fluctuation trend, change gradient, and mutation characteristics. The system constructs a delay prediction curve by collecting the delay data during the historical video transmission process. This curve reflects the change of the network state over time.

[0062] The fluctuation trend is obtained by calculating the standard deviation of 30 consecutive sampling points. When the standard deviation value is less than 10ms, it is determined to be in a stable state; between 10ms and 30ms is a slightly fluctuating state; greater than 30ms is a severely fluctuating state. For example, during a certain transmission process, the delay standard deviation of 30 consecutive sampling points is 8ms, then it is determined that the current network is in a stable state.

[0063] The change gradient is obtained by calculating the delay change rate of the last 10 sampling points. A positive change rate indicates an increase in delay and a decline in network performance; a negative change rate indicates a decrease in delay and an improvement in network performance. In practical applications, when it is detected that the delay increases from 50ms to 100ms within 3 seconds, the calculated change gradient is 16.7ms / s, indicating that the network performance is rapidly declining.

[0064] The mutation characteristics are identified by detecting the jump points in the delay curve. The system defines a point where the delay changes by more than 50ms within 100ms as a mutation point. In practical applications, when the network switches from a wired connection to a wireless connection, there may be a situation where the delay suddenly increases from 30ms to 90ms at the moment of switching, and the system will mark this point as a mutation point.

[0065] After obtaining the fluctuation trend, change gradient, and mutation characteristics, input them into the model predictive control algorithm based on a sliding time window. This algorithm uses a 5 - second sliding time window and constructs a coupled constraint model of video quality and transmission delay on the basis of predicting the future network state.

[0066] The coupled constraint model considers the following factors: (1) the relationship between the video bit rate and the network bandwidth; (2) the relationship between the encoding parameters and the video quality; (3) the relationship between the key - frame distribution and the transmission anti - interference ability. Based on these relationships, the system establishes a decision matrix, and each element in the matrix corresponds to a combination of encoding parameters under a specific network state.

[0067] When the system detects that the fluctuation trend is in a stable state (standard deviation is 8ms), the change gradient is slightly decreasing (-5ms / s) and there are no mutation characteristics, the model predictive control algorithm queries the decision matrix and outputs the optimal adjustment plan: the key frame interval is set to 90 frames, the quantization parameter is set to 26, and the target bitrate is adjusted to 85% of the original value.

[0068] The system also continuously optimizes the decision matrix by analyzing the historical adjustment effects. When a certain set of encoding parameters performs excellently under a specific network state (the decoding fluency at the user end is improved by more than 20%), the system will increase the weight of this parameter combination under similar network states.

[0069] When the fluctuation trend indicates that the network state is stable (standard deviation is less than 10ms), the system increases the key frame interval to improve the encoding efficiency. In practical applications, the key frame interval gradually increases from the standard 60 frames to 120 frames, increasing by 10 frames each time, while monitoring the decoding quality at the user end. When the key frame interval increases to 90 frames, the encoding efficiency is improved by 15%, and the decoding quality drops by no more than 5%. At this time, the best balance point is achieved.

[0070] When the mutation characteristics indicate a drastic fluctuation in network performance (a delay mutation point is detected), the system reduces the key frame interval to enhance the anti-interference ability. For example, when it is detected that the delay mutates from 40ms to 95ms, the system immediately reduces the key frame interval from 90 frames to 30 frames and maintains this setting for the next 3 seconds to ensure that the user can still receive a clear picture during the network fluctuation.

[0071] Adjust the quantization parameter according to the positive or negative direction of the change gradient. When the change gradient is positive (such as 16.7ms / s), increase the quantization step size to reduce the bitrate. In practical applications, the system increases the quantization parameter from 26 to 32, reducing the bitrate by about 40%, from 4Mbps to 2.4Mbps, effectively alleviating network congestion. When the change gradient is negative (such as -10ms / s), the system reduces the quantization step size to improve the quality, reducing the quantization parameter from 26 to 22, enhancing the detailed performance of the picture. At this time, the bitrate increases from 4Mbps to 5.2Mbps.

[0072] Calculate the real-time adjustment amount of the target bitrate based on the amplitude of the fluctuation trend. For every 5ms increase in the amplitude of the fluctuation trend (standard deviation), the system reduces the target bitrate by 5%. For example, when the standard deviation increases from 5ms to 20ms, the system gradually reduces the target bitrate from 100% of the original setting value to 85%.

[0073] Update the adjusted encoding parameters to the video encoder to generate an optimized video bitstream. In practical applications, the H.264 encoder reconfigures the encoding strategy according to the adjusted parameters (key frame interval, quantization parameter, target bitrate) to generate a video bitstream suitable for the current network state.

[0074] The system adopts an intelligent routing mechanism to calculate the path priority by combining the delay status of transmission nodes and link bandwidth resources. The intelligent routing mechanism maintains a routing table containing multiple available transmission paths, and each path has the following attributes: current average delay, packet loss rate, available bandwidth, and historical stability score.

[0075] The path priority calculation considers the weighted combination of the above four factors. In practical applications, there are three alternative paths in a certain remote control scenario: 4G network (delay 80ms, packet loss rate 1%, bandwidth 5Mbps, stability score 85), Wi-Fi network (delay 40ms, packet loss rate 0.5%, bandwidth 10Mbps, stability score 92), and wired network (delay 30ms, packet loss rate 0.1%, bandwidth 100Mbps, stability score 98). After calculation, the priority scores of the three paths are 72, 88, and 96 respectively, and the system selects the wired network with the highest priority as the main transmission path.

[0076] When the performance of the main transmission path degrades, the system will switch to the sub-optimal path in real time. For example, when the delay of the wired network suddenly increases to 120ms, its priority score drops to 65, which is lower than 88 of the Wi-Fi network, and the system automatically switches the transmission path to the Wi-Fi network to ensure the continuity and stability of video transmission.

[0077] Through the combined implementation of the above technical means, the present invention can dynamically adjust video coding parameters and transmission paths according to the real-time changes of network status, effectively improving the video transmission quality and user experience of the remote control system. Experimental data show that in an environment with frequent changes in network status, after adopting this method, the end-to-end delay of the video stream is reduced by an average of 25%, the video freeze rate is reduced by 40%, and the user operation response sensitivity is increased by 30%.

[0078] Figure 3 Schematic diagram for comprehensive evaluation of video quality and transmission efficiency in different network environments of the embodiments of the present invention: This table shows a comparative display of various network performance metrics in a stable network environment and a fluctuating network environment. In a stable network environment, with technological optimization, all metrics show an obvious improvement trend: the video quality (PSNR) increases from 38.7 dB to 39.2 dB, the structural similarity (SSIM) reaches 0.97, the video freeze time is significantly reduced to 0.2 seconds, the end-to-end delay is reduced to 127 ms, the bandwidth utilization rate is increased to 83.7%, the average bit rate remains at a reasonable level of 4.3 Mbps, and the number of transmission interruptions is reduced to 0 times, and the mutation adaptation recovery time is only 1.3 seconds. In contrast, in a fluctuating network environment, although all metrics are relatively poor, significant improvements have also been achieved through optimization: the video quality increases from 31.2 dB to 36.5 dB, the structural similarity is increased to 0.92, the video freeze time is reduced to 1.6 seconds, the end-to-end delay is reduced to 412 ms, the bandwidth utilization rate is increased to 72.6%, the average bit rate reaches 3.5 Mbps, the number of transmission interruptions is reduced to 2 times, and the mutation adaptation recovery time is shortened to 2.8 seconds. Notably, in a fluctuating network environment, the adjustable range of the average interval between key frames is expanded to 10 - 80 frames, providing more flexible adaptability. These data fully demonstrate that the system can achieve good performance in different network environments, especially showing good adaptability and stability when dealing with network fluctuations.

[0079] In an alternative embodiment, the adjusted fluctuation trend, change gradient, and mutation characteristics are updated to the video encoder to generate an optimized video bitstream; based on the network state information of the delay prediction curve, an optimal transmission path is selected for the optimized video bitstream through an intelligent routing mechanism. The intelligent routing mechanism calculates the path priority by combining the delay status of the transmission nodes and the link bandwidth resources, including: The fluctuation trend, change gradient, and mutation characteristics are input into the video encoder for update; the video encoder processes the video bitstream based on the updated encoding parameters to generate an optimized video bitstream; Collect the processing delay generated when the network transmission node processes the optimized video bitstream, and record the queuing delay of the data packet in the node buffer; add the processing delay of a single node to the queuing delay to obtain the total node delay; accumulate the total node delays of each node along the transmission path in sequence to obtain the node delay status of the transmission path; divide the node delay status by the preset delay threshold to obtain the node delay score; Monitor the network link status carrying the optimized video bitstream, and obtain the total bandwidth capacity and the occupied bandwidth of the link; calculate the available bandwidth value of each link; select the link with the smallest available bandwidth in the transmission path as the path bandwidth resource; divide the path bandwidth resource by the current bandwidth requirement of the video bitstream to obtain the bandwidth resource score; The node delay score and the bandwidth resource score are weighted. After normalizing the node delay score, it is multiplied by the weight coefficient 0.7, and after normalizing the bandwidth resource score, it is multiplied by the weight coefficient 0.3. The two are added together to obtain the path priority score.

[0080] The intelligent video transmission optimization system first analyzes the video data stream, extracts the fluctuation trend, change gradient, and mutation characteristics, and then inputs these characteristics into the video encoder for optimization processing.

[0081] The system analyzes the video data frame sequence through the sliding window algorithm, and the window size is set to 30 frames. For the data within each window, the system calculates the following characteristics: The fluctuation trend characteristic is obtained by analyzing the data change direction of 15 consecutive frames. The system records the data change direction (rising or falling) between every two adjacent frames. If more than 10 out of 15 frames show the same change direction, it is considered that there is an obvious fluctuation trend. For example, within a certain 30-frame window, the system detects that 12 consecutive frames of data show an upward trend, and the fluctuation trend characteristic value is recorded as 0.8 (12 / 15).

[0082] The change gradient characteristic is obtained by calculating the numerical change rate between adjacent frames. The system divides the numerical difference between each pair of adjacent frames by the numerical value of the previous frame to obtain the change rate. If the change rate exceeds the preset threshold of 0.15, the system records it as a high-gradient change point. In practical applications, 5 high-gradient change points are detected in a certain video segment, and the change rates are 0.18, 0.22, 0.16, 0.25, and 0.17 respectively. The system calculates the average change gradient as 0.196.

[0083] The mutation characteristic refers to the drastic change of data in a short period of time. The system sets the mutation threshold to 0.3. If the change rate between two adjacent frames exceeds this threshold, it is recorded as a mutation point. For example, when analyzing a video segment of a complex scene transition, the system detects two mutation points, and the change rates are 0.35 and 0.42 respectively.

[0084] The system inputs the extracted characteristics into the video encoder, and the encoder dynamically adjusts the encoding parameters according to these characteristics: When the fluctuation trend characteristic value is greater than 0.7, the system reduces the quantization parameter by 5% to retain more detailed information; when the average change gradient is greater than 0.18, the system reduces the key frame interval by 20% to improve the encoding accuracy; when a mutation point is detected, the system inserts an additional key frame before and after the mutation frame to ensure the picture quality.

[0085] In a specific example, the original encoding parameters of a certain video segment are: quantization parameter 25, key frame interval 60 frames. The system detects that the fluctuation trend eigenvalue is 0.8, the average change gradient is 0.196, and there are two mutation points. The optimized encoding parameters are adjusted to: quantization parameter 23.75, key frame interval 48 frames, and one additional key frame is added before and after each of the two mutation points.

[0086] The optimized video bitstream has higher picture quality and lower bandwidth requirements. Actual measurements show that, under the condition of the same subjective picture quality score (SSIM value 0.92), the optimized bitstream saves 18.5% of the bandwidth compared to the original bitstream.

[0087] To achieve intelligent route selection, the system needs to evaluate the delay status of network transmission nodes. The system obtains node delay information through the following steps: The system deploys monitoring modules at each network node to collect and process delay data in real time. Processing delay refers to the time required for a data packet to be processed by a node, including the time consumption of operations such as parsing the packet header, querying the routing table, and forwarding decisions.

[0088] In an actual network environment, the average processing delay of a certain core node for processing a single video data packet is 0.85 milliseconds, and the peak processing delay is 1.2 milliseconds. The system samples the processing delay data every 100 milliseconds and calculates the sliding average of 10 sampling points as the current processing delay.

[0089] The system also records the queuing delay of data packets in the node buffer. The queuing delay depends on the current load status of the node and the buffer management strategy. The system obtains the queuing delay by detecting the time difference between the timestamps when the data packet enters and leaves the buffer.

[0090] In actual measurements, the average queuing delay of a certain edge node during the peak period is 3.5 milliseconds, and the maximum queuing delay is 7.8 milliseconds. The system also uses a 10-point sliding average to smooth the queuing delay.

[0091] The system adds the processing delay and the queuing delay to obtain the total node delay. For example, the total delay of the above-mentioned edge node is the processing delay of 0.85 milliseconds plus the queuing delay of 3.5 milliseconds, which is equal to 4.35 milliseconds.

[0092] The total node delays of each node along the transmission path are accumulated in sequence to obtain the node delay status of the transmission path. In a typical five-node transmission path, the total node delays of each node are 4.35, 3.82, 2.96, 4.12, and 3.55 milliseconds respectively, and the total path delay is 18.8 milliseconds.

[0093] The system preset delay threshold is 25 milliseconds. Divide the delay status of the path node by the preset delay threshold to obtain the node delay score. The node delay score of the above path is 18.8 / 25 = 0.752. The lower the node delay score, the better the path delay performance.

[0094] The system deploys a bandwidth monitoring module on the network link to obtain the total bandwidth capacity and occupied bandwidth data of the link in real time. The total bandwidth capacity refers to the maximum transmission capacity of the link, and the occupied bandwidth refers to the bandwidth resources currently in use.

[0095] For example, the total bandwidth capacity of a backbone link is 10 Gbps, and the currently occupied bandwidth is 6.8 Gbps. The system calculates the available bandwidth as 3.2 Gbps through subtraction.

[0096] The system collects bandwidth data every 500 milliseconds and uses the exponentially weighted moving average algorithm to calculate the smoothed available bandwidth value, with the weight factor set to 0.3. This can avoid the impact of instantaneous bandwidth fluctuations on routing decisions.

[0097] In a transmission path containing multiple links, the system selects the link with the smallest available bandwidth as the path bandwidth resource. For example, in a five-link path, the available bandwidths of each link are 3.2 Gbps, 2.8 Gbps, 4.5 Gbps, 1.9 Gbps, and 3.7 Gbps respectively, then the path bandwidth resource is 1.9 Gbps.

[0098] Assume that the current optimized video stream bandwidth requirement is 15 Mbps. The system divides the path bandwidth resource by the video stream bandwidth requirement to obtain the bandwidth resource score: 1900 Mbps / 15 Mbps = 126.67. The higher the bandwidth resource score, the more sufficient the path bandwidth resources.

[0099] The system first normalizes the node delay score and the bandwidth resource score. For the node delay score, normalization is performed by subtracting the original score value from 1, so that the lower the delay, the higher the score. For example, the normalized node delay score is 1 - 0.752 = 0.248.

[0100] For the bandwidth resource score, the system sets an upper threshold of 200 and normalizes the original score by dividing it by the threshold. If the result is greater than 1, the value is taken as 1. For example, the normalized bandwidth resource score is 126.67 / 200 = 0.633.

[0101] The system multiplies the normalized node delay score by the weight coefficient 0.7, multiplies the normalized bandwidth resource score by the weight coefficient 0.3, and adds the two to obtain the path priority score: 0.248×0.7 + 0.633×0.3 = 0.3736.

[0102] In practical applications, the system evaluates multiple possible transmission paths simultaneously. For example, the system evaluates three different paths, and the calculated path priority scores are 0.3736, 0.4251, and 0.3092 respectively. The system selects the second path (0.4251) with the highest score as the transmission path for the optimized video stream.

[0103] Through actual measurement, after adopting this intelligent routing mechanism, the average end-to-end delay of video transmission is reduced by 22.3%, the video playback stuttering rate is reduced by 35.7%, and the user experience satisfaction is improved by 28.6%.

[0104] Figure 4 Schematic diagram for comprehensive performance evaluation of video stream optimization and intelligent routing mechanism in the embodiment of the present invention: This table details the changes in network performance metrics under two conditions of normal network load and high network load. Under normal network load conditions, the system performance shows an obvious optimization trend: the average end-to-end delay is reduced from 186 ms to 132 ms, the path switching success rate is significantly increased to 94.8%, the bandwidth utilization rate is increased to 84.6%, the video quality (PSNR) reaches 39.1 dB, the number of transmission interruptions is greatly reduced to 3 times per hour, the average node delay is reduced to 23.5 ms, the average path score is increased to 0.86, although the computational time overhead slightly increases to 4.3 ms / frame. Under high network load conditions, although each index is relatively poor, significant improvement has also been achieved through optimization: the average end-to-end delay is reduced from 875 ms to 421 ms, the path switching success rate is increased to 87.5%, the bandwidth utilization rate is increased to 76.3%, the video quality is improved to 36.2 dB, the number of transmission interruptions is reduced to 8 times per hour, the average node delay is reduced to 42.8 ms, the average path score is increased to 0.74, and the computational time overhead is equivalent to that under normal load, which is 4.6 ms / frame. These data indicate that the system demonstrates good performance and optimization effects under different load conditions. Especially under high load conditions, it can still maintain good service quality, reflecting the robustness and reliability of the system. Overall, the optimized system can provide stable and high-quality services under various network conditions.

[0105] In an alternative embodiment, calculate the time difference between the arrival time and the sending time of the video stream, use the recursive least squares algorithm to model the time difference to obtain a probability distribution model; perform error analysis on the probability distribution model and the video stream delay prediction curve, and generate a delay compensation control strategy through a hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy, including: Calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, where the probability distribution model characterizes the statistical characteristics of the video stream delay data; Generate a video stream delay prediction curve according to the probability distribution model, where the video stream delay prediction curve reflects the change trend of the video stream delay data; compare and analyze the change trend with the actually measured delay data to calculate the prediction error; construct an error evaluation function based on the prediction error; Take the error evaluation function as the optimization objective of the hierarchical reinforcement learning algorithm, and use the hierarchical reinforcement learning algorithm to generate a delay compensation control strategy; the delay compensation control strategy dynamically adjusts the execution timing of the control command based on the prediction result of the video stream delay data; The remote control terminal receives the operation instructions of the operator for the crane, calculates the time compensation value of the operation instructions according to the delay compensation control strategy; performs adaptive delay compensation on the operation instructions based on the time compensation value.

[0106] The system first calculates the time difference between the transmission time and the arrival time of the video stream to generate video stream delay data. Specifically, the sending end embeds timestamp information in the video frame, and the receiving end parses the timestamp and compares it with the local time to obtain the delay data. To ensure the time synchronization accuracy, the system uses the Network Time Protocol (NTP) to synchronize the clocks of the sending end and the receiving end.

[0107] In practical applications, examples of the delay data collected by the system are as follows: the transmission time of the video frame with frame ID 1001 is 10:15:30.125, the arrival time is 10:15:30.246, and the calculated delay is 121 milliseconds; the delay of the video frame with frame ID 1002 is 135 milliseconds; the delay of the video frame with frame ID 1003 is 118 milliseconds. The system continuously collects the delay data of 1000 video frames as the input of the recursive least squares algorithm.

[0108] The system uses the recursive least squares algorithm to perform iterative modeling on the video stream delay data. This algorithm updates the model parameters through an iterative method to minimize the mean square error between the predicted delay and the actual delay. Each time new delay data is received, the system updates the model parameters according to a preset weight factor (typical value is 0.98).

[0109] Through the iterative calculation of the recursive least squares algorithm, the system obtains a model that can characterize the change law of the video stream delay. Based on this model, the system constructs a probability distribution model of the video stream delay, which can reflect the statistical characteristics of the delay data, including the mean, variance, and distribution form.

[0110] In practical applications, the system analyzes the latency data of the aforementioned 1000 video frames, obtains the average latency of 125 milliseconds and the standard deviation of 15 milliseconds, and determines that the latency data approximately follows a normal distribution with a mean of 125 milliseconds and a standard deviation of 15 milliseconds. This probability distribution model provides a statistical basis for subsequent latency prediction and compensation.

[0111] Based on the constructed probability distribution model, the system generates a video stream latency prediction curve. This prediction curve not only considers the statistical characteristics of historical latency data but also combines factors such as network conditions and data transmission volume to dynamically predict the latency change trend in the future for a period of time.

[0112] The system uses the sliding window technique to select the latency data of the most recent 100 video frames for short-term trend analysis. Through the weighted average method, higher weights are assigned to recent data (the weight of the most recent 10 frames is 0.6, and the weight of the remaining frames is 0.4) to generate the latency prediction values for the next 30 frames.

[0113] To evaluate the prediction accuracy, the system compares the predicted latency values with the actual measured latency data and calculates the prediction error. Taking a certain test as an example, the system predicts the latency of the 1101st frame to be 130 milliseconds, the actual measured value is 133 milliseconds, and the prediction error is 3 milliseconds; the predicted latency of the 1102nd frame is 128 milliseconds, and the actual value is 125 milliseconds, with a prediction error of 3 milliseconds.

[0114] The system constructs an error evaluation function based on the prediction error. This function comprehensively considers the absolute error, relative error, and error change rate. In practical applications, the system uses the average absolute error (4.5 milliseconds), the maximum absolute error (12 milliseconds), and the error standard deviation (2.8 milliseconds) of the most recent 100 predicted frames as the key indicators of the error evaluation function.

[0115] The system takes the error evaluation function as the optimization objective of the hierarchical reinforcement learning algorithm and generates a latency compensation control strategy through the hierarchical reinforcement learning algorithm. The hierarchical reinforcement learning algorithm decomposes the control task into two levels: high-level policy selection and low-level execution control.

[0116] The high-level policy is responsible for selecting the appropriate compensation mode according to the network state and latency characteristics, including the early compensation mode, the delay compensation mode, and the hybrid compensation mode. The system divides the state space according to the latency change rate: when the latency change rate is less than 5%, the early compensation mode is adopted; when the latency change rate is greater than 15%, the delay compensation mode is adopted; in other cases, the hybrid compensation mode is adopted.

[0117] The low-level execution control is responsible for calculating the specific time compensation value. In the early compensation mode, the system sends control instructions in advance based on the predicted delay value; in the delay compensation mode, the system delays the execution of the operation instruction until it is confirmed that the video frame has been updated; in the hybrid compensation mode, the system dynamically adjusts the compensation strategy by combining the advantages of the two modes.

[0118] The reinforcement learning algorithm continuously optimizes the compensation strategy through repeated trials and evaluations. The system sets a reward function, giving a positive reward (value +1) when the synchronization error between the execution time of the operation instruction and the video frame is less than 10 milliseconds, a negative reward (value -1) when the error is greater than 50 milliseconds, and in other cases, the reward is inversely proportional to the error.

[0119] After 1000 iterations of learning, the delay compensation control strategy generated by the system can control the average synchronization error within 15 milliseconds, meeting the real-time requirements of crane remote operation.

[0120] The remote control terminal receives the operation instructions of the operator for the crane and calculates the time compensation value of the operation instructions according to the delay compensation control strategy. Specifically, the remote control terminal maintains an instruction buffer queue and sorts it according to the priority and timeliness of different types of instructions.

[0121] For high-priority safety instructions (such as emergency stop), the system sets the highest execution priority (value 10) and executes them directly without delay compensation; for fine operation instructions (such as fine-tuning the position), the system sets a medium priority (value 5) and applies full-delay compensation; for regular movement instructions, the system sets a normal priority (value 3) and applies partial delay compensation.

[0122] The system calculates the optimal execution time of each operation instruction based on the current network state and the predicted delay value. For example, when the predicted video delay is 120 milliseconds, the system applies 115 milliseconds of early compensation to the crane rotation instruction to ensure that the equipment response seen by the operator on the video frame is synchronized with the actual operation.

[0123] The remote control terminal also implements an instruction fusion mechanism to merge and optimize the same type of instructions issued within a short period. For example, five small-angle rotation instructions (each 5 degrees) continuously issued within 500 milliseconds are merged into a smooth 25-degree rotation instruction to reduce control jitter and improve operation smoothness.

[0124] Through the comprehensive application of the above technical means, the system realizes the accurate modeling, prediction and compensation of video stream delay, enabling the remote operation of the crane to have the control characteristics of low delay and high precision, and significantly improving the safety and efficiency of remote operation. Experimental results show that, compared with traditional methods, the system shortens the operation response time by 37% and improves the operation accuracy by 42%, effectively solving the problem of delay uncertainty in remote operation.

[0125] Figure 5 Schematic diagram for precision analysis and comparison of the delay probability distribution model in the embodiment of the present invention: This table compares the performance of the Gaussian mixture model, random forest prediction and the technical solution of the present invention under different network states. From the perspective of the KL divergence index, the technical solution of the present invention shows the lowest divergence value in various network environments: 0.102 under stable network, 0.198 under fluctuating network, and 0.275 under mutating network, which is significantly better than the other two methods. In terms of the Poisson coefficient, the technical solution of the present invention also demonstrates excellent performance, reaching 0.938 (stable network), 0.872 (fluctuating network) and 0.803 (mutating network) respectively, indicating better prediction accuracy. In terms of the root mean square error index, the present solution also achieves the minimum value, which are 0.011, 0.026 and 0.047 respectively under the three network states, and the error control effect is remarkable. It is particularly worth noting that in terms of computational efficiency, the technical solution of the present invention shows obvious advantages: only 12.6 ms under stable network environment, 14.3 ms under fluctuating network, and 18.9 ms under mutating network, which is significantly improved compared with the Gaussian mixture model and random forest prediction. These data fully prove that the technical solution of the present invention has achieved comprehensive performance improvement in multiple dimensions such as accuracy, stability and computational efficiency. Especially in complex network environments, it can still maintain good prediction effect and operation efficiency, demonstrating strong practical value and technical advantages.

[0126] In an alternative embodiment, calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, and the probability distribution model characterizes the statistical characteristics of the video stream delay data, including: Calculate the time difference between the arrival time and the transmission time of the video stream to generate video stream delay data; perform outlier detection and normalization processing on the video stream delay data to eliminate clock synchronization errors; Construct a recursive least squares algorithm model, input the video stream delay data into the recursive least squares algorithm model, calculate the prediction error and gain matrix; update the parameter estimation value based on the prediction error and gain matrix; optimize the covariance matrix according to the parameter estimation value to realize the iterative optimization of the model parameters; Analyze the statistical law of video stream delay data based on the calculation results of the recursive least squares algorithm model, and calculate the mean, variance, skewness, and kurtosis of the video stream delay data; select the optimal probability distribution function to fit the video stream delay data and estimate the distribution function parameters; construct a probability distribution model based on the optimal probability distribution function, and the probability distribution model quantitatively characterizes the statistical characteristics of the video stream delay data.

[0127] Obtain the sending time and arrival time of the video stream data. Video stream data usually contains timestamp information. The sending time is recorded at the sending end, and the arrival time is recorded at the receiving end. For example, the sending time of a certain video frame is 1623456789.123 milliseconds, and the arrival time is 1623456839.456 milliseconds.

[0128] When calculating the video stream delay data, calculate the time difference between the arrival time and the sending time for each video frame. For example, the delay data of the above video frame is 50.333 milliseconds. For a video stream with a duration of 10 minutes, if the video frame rate is 30 frames per second, there are a total of 18000 video frames, and 18000 delay data points are correspondingly generated.

[0129] Perform outlier detection and normalization processing on the delay data. Use the box plot method to identify outliers, calculate the first quartile Q1 (such as 45.2 milliseconds) and the third quartile Q3 (such as 55.8 milliseconds) of all delay data, and obtain the interquartile range IQR = Q3 - Q1 = 10.6 milliseconds. Mark the data points outside the range [Q1 - 1.5×IQR, Q3 + 1.5×IQR] as outliers, that is, the data points less than 29.3 milliseconds or greater than 71.7 milliseconds are excluded.

[0130] Normalization processing is used to eliminate clock synchronization errors. First, calculate the average value of the effective delay data (after excluding outliers), such as 50.5 milliseconds. Then subtract this average value from each delay data point and divide by the standard deviation (such as 5.3 milliseconds) to obtain the normalized delay data. For example, for a data point with an original delay of 60.2 milliseconds, after normalization, it becomes (60.2 - 50.5) / 5.3 = 1.83.

[0131] When constructing a recursive least squares algorithm model, first determine the model order. Based on experience, select a second-order autoregressive model, that is, the current delay data is related to the previous two delay data points. Initialize the model parameter matrix, set the initial values of the weight coefficients to 0.1, and the initial value of the covariance matrix to a diagonal matrix with diagonal elements all being 100.

[0132] Input the processed delayed data sequence into the recursive least squares algorithm model. For the i-th data point in the sequence, calculate the predicted value at time i based on the known data at times i-1 and i-2. For example, if the normalized delay at time i-2 is 1.2 and at time i-1 is 0.8, and the weight coefficients are 0.25 and 0.35 respectively, then the predicted value at time i is 0.25×1.2 + 0.35×0.8 = 0.58.

[0133] Calculate the prediction error, which is the difference between the actual observed value and the predicted value. If the actual normalized delay at time i is 0.9, then the prediction error is 0.9 - 0.58 = 0.32. Calculate the gain matrix based on the prediction error. The gain matrix is used to adjust the sensitivity of the model to new data. The typical range of gain matrix element values is from 0.01 to 0.2.

[0134] Update the parameter estimates using the gain matrix and the prediction error. If the current weight coefficients are [0.25, 0.35], and the product of the gain matrix and the prediction error is [0.015, 0.01], then the updated weight coefficients are [0.265, 0.36]. Update the covariance matrix, reducing the diagonal element values to improve the model stability. The typical updated diagonal elements decrease from the initial 100 to about 5 - 10.

[0135] The model parameters are optimized through multiple iterations. When the change amplitude of the parameters is less than a preset threshold (such as 0.001) or the maximum number of iterations (such as 1000 times) is reached, the iteration stops. The final model parameters reflect the temporal correlation of the video stream delay data. For example, if the final weight coefficients are [0.42, 0.31], it indicates the degree of correlation between the current delay and the delays at the previous two times.

[0136] Based on the calculation results of the recursive least squares algorithm model, analyze the statistical laws of the video stream delay data. Calculate the mean (such as 50.5 milliseconds), variance (such as 28.09 milliseconds²), skewness (such as 0.15), and kurtosis (such as 2.87) of the video stream delay data. Skewness reflects the asymmetry of the distribution, and kurtosis reflects the peakedness of the distribution.

[0137] Select the optimal probability distribution function to fit the video stream delay data. Common distribution functions include normal distribution, lognormal distribution, gamma distribution, etc. Based on the characteristics of the measured data, if the kurtosis is close to 3 and the skewness is close to 0, then select the normal distribution; if the distribution is right-skewed (skewness is positive), then consider the gamma distribution or the lognormal distribution.

[0138] For the selected distribution function, estimate its parameter values. For example, for the normal distribution, the parameters are the mean and the standard deviation; for the gamma distribution, the parameters are the shape parameter and the scale parameter. Use the maximum likelihood estimation method to determine the parameter values. For example, for the gamma distribution, estimate the shape parameter to be 8.2 and the scale parameter to be 6.1.

[0139] Construct a probability distribution model to quantitatively characterize the statistical features of video stream delay data. Calculate the probability values at specific delay thresholds. For example, the probability that the delay is less than 60 milliseconds is 96.5%, and the probability that the delay is less than 70 milliseconds is 99.3%. These statistical features have important reference values for video stream quality assessment, network planning, and buffer strategy optimization.

[0140] The probability distribution model can also be used to generate delay prediction intervals. For example, based on the constructed gamma distribution model, it can be obtained that 95% of the delay values will fall within the interval [40.2 milliseconds, 62.8 milliseconds]. Such prediction intervals have practical values for application scenarios such as adaptive buffer size adjustment and user experience quality assessment.

[0141] In the second aspect of the embodiments of the present invention, A crane remote control image delay detection system is provided, including: A first unit for transmitting real-time video stream data to an edge computing server local to the crane; the edge computing server performs spatio-temporal segmentation on the real-time video stream data by using a dual-channel parallel preprocessing mechanism to obtain a video frame sequence with timestamps; uses an attention mechanism to extract features from the video frame sequence to obtain the three-dimensional spatial coordinates and motion parameters of the crane; based on the three-dimensional spatial coordinates and motion parameters, calculates a motion state vector through an adaptive Kalman filtering algorithm; A second unit for inputting the motion state vector into a multi-modal deep learning model. The multi-modal deep learning model extracts the dynamic features of the video stream through a temporal graph convolutional network, deeply fuses the dynamic features with the network transmission quality indicators collected in real time, and generates a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, combines a model predictive control algorithm to calculate an optimal adjustment strategy for video coding parameters; optimizes the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmits the optimized video stream to a remote control terminal through an intelligent routing mechanism; A third unit for calculating the time difference between the arrival time and the sending time of the video stream, modeling the time difference by using a recursive least squares algorithm to obtain a probability distribution model; performing error analysis on the probability distribution model and the video stream delay prediction curve, and generating a delay compensation control strategy through a hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

[0142] In the third aspect of the embodiments of the present invention, An electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0143] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0144] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for remotely controlling and detecting image delay of a crane, characterized in that, Including: Transmitting real-time video stream data to the edge computing server local to the crane; The edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; uses an attention mechanism to extract features from the video frame sequence to obtain the three-dimensional space coordinates and motion parameters of the crane; based on the three-dimensional space coordinates and motion parameters, calculates the motion state vector through an adaptive Kalman filtering algorithm; Inputting the motion state vector into a multi-modal deep learning model, the multi-modal deep learning model extracts the dynamic features of the video stream through a temporal graph convolutional network, deeply fuses the dynamic features with the network transmission quality indicators collected in real time, and generates a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, combines the model predictive control algorithm to calculate the optimal adjustment strategy of the video coding parameters; Optimizes the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmits the optimized video stream to the remote control terminal through an intelligent routing mechanism; Calculates the time difference between the arrival time and the sending time of the video stream, uses the recursive least squares algorithm to model the time difference to obtain a probability distribution model; performs error analysis on the probability distribution model and the video stream delay prediction curve, and generates a delay compensation control strategy through a hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

2. The method according to claim 1, wherein The edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with timestamps; uses an attention mechanism to extract features from the video frame sequence to obtain the three-dimensional space coordinates and motion parameters of the crane; Based on the three-dimensional space coordinates and motion parameters, the motion state vector calculated through the adaptive Kalman filtering algorithm includes: The edge computing server uses a dual-channel parallel preprocessing mechanism to process the real-time video stream data; the time channel establishes a temporal feature vector for the real-time video stream data, and the space channel divides the real-time video stream data into multiple regions of interest and generates spatial feature vectors; fuses the temporal feature vector and the spatial feature vector to generate a video frame sequence with timestamps; Constructs a multi-head attention network for feature extraction on the video frame sequence, and the multi-head attention network includes a spatial attention module and a channel attention module; based on the multi-head attention network, calculates the mapping relationship of feature points in three-dimensional space to obtain the three-dimensional space coordinates and motion parameters of the crane; Constructs an adaptive Kalman filtering model based on the three-dimensional space coordinates and motion parameters, calculates the measurement reliability index based on the adaptive Kalman filtering model and dynamically adjusts the measurement noise matrix; analyzes the state prediction error based on the adaptive Kalman filtering model to obtain a new process noise matrix; Input the measurement noise matrix and the process noise matrix into the state prediction equation, and the state prediction equation uses the state transition matrix to predict the motion state of the crane; input the prediction result into the measurement update equation, and the measurement update equation corrects the predicted state based on the adaptive Kalman filter model; calculate the error covariance according to the corrected state, and optimize the state estimation result based on the error covariance to output the motion state vector including position parameters, speed parameters, and acceleration parameters.

3. The method according to claim 2, characterized in that, Construct an adaptive Kalman filter model based on three-dimensional space coordinates and motion parameters; the adaptive Kalman filter model establishes a measurement noise statistical model, calculates the measurement reliability index based on the measurement noise statistical model, and dynamically adjusts the measurement noise matrix according to the measurement reliability index; The adaptive Kalman filter model analyzes the state prediction error to obtain the system uncertainty parameters, and updates the process noise matrix based on the system uncertainty parameters, including: Construct an adaptive Kalman filter model based on the three-dimensional space coordinates and motion parameters of the crane; input the three-dimensional space coordinates and motion parameters into the adaptive Kalman filter model to generate an observation data sequence; The adaptive Kalman filter model establishes a measurement noise statistical model based on the observation data sequence, and the measurement noise statistical model calculates the residual vector and variance distribution of the observation data; constructs a measurement reliability index based on the residual vector and variance distribution, and the measurement reliability index uses an exponential function form to characterize the credibility of the observation data; sets an adaptive weight factor according to the credibility, and dynamically adjusts the measurement noise matrix based on the adaptive weight factor; The adaptive Kalman filter model performs error analysis on the state prediction result and the actual observation value, and calculates the statistical characteristics of the innovation sequence; evaluates the system dynamic characteristics based on the statistical characteristics of the innovation sequence to obtain a parameter vector characterizing the system uncertainty; calculates the deviation degree of the system model according to the parameter vector, and sets a fading factor based on the deviation degree; uses the fading factor to recursively update the process noise matrix.

4. The method according to claim 1, wherein Based on the change trend of the video stream delay prediction curve, combine the model predictive control algorithm to calculate the optimal adjustment strategy of the video coding parameters; Optimize the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmit the optimized video stream to the remote control terminal through the intelligent routing mechanism, including: Perform dynamic analysis on the video stream delay prediction curve, extract the fluctuation trend, change gradient, and mutation characteristics from the video stream delay prediction curve. The fluctuation trend characterizes the stability of the network state, the change gradient reflects the change rate of the network performance, and the mutation characteristics indicate the drastic change points of the network state; input the fluctuation trend, change gradient, and mutation characteristics into the model predictive control algorithm based on the sliding time window. The model predictive control algorithm outputs the optimal adjustment scheme of the coding parameters by constructing a coupling constraint model of video quality and transmission delay; Adjust the fluctuation trend, change gradient, and mutation characteristics based on the optimal adjustment scheme. When the fluctuation trend indicates a stable network state, increase the key frame interval to improve the encoding efficiency; when the mutation characteristics indicate a drastic fluctuation in network performance, decrease the key frame interval to enhance the anti-interference ability; adjust the quantization parameter according to the positive or negative direction of the change gradient. When the change gradient is positive, increase the quantization step size to reduce the bit rate, and when the change gradient is negative, decrease the quantization step size to improve the quality; calculate the real-time adjustment amount of the target bit rate based on the amplitude of the fluctuation trend. Update the adjusted fluctuation trend, change gradient, and mutation characteristics to the video encoder to generate an optimized video bitstream; calculate the path priority based on the intelligent routing mechanism combined with the delay status of the transmission node and the link bandwidth resources; send the optimized video bitstream to the remote control terminal through the transmission path with the highest path priority.

5. The method according to claim 4, characterized in that, Update the adjusted fluctuation trend, change gradient, and mutation characteristics to the video encoder to generate an optimized video bitstream; based on the network state information of the delay prediction curve, select the optimal transmission path for the optimized video bitstream through the intelligent routing mechanism. The intelligent routing mechanism combined with the delay status of the transmission node and the link bandwidth resources to calculate the path priority includes: Input the fluctuation trend, change gradient, and mutation characteristics into the video encoder for update; the video encoder processes the video bitstream based on the updated encoding parameters to generate an optimized video bitstream. Collect the processing delay generated when the network transmission node processes the optimized video bitstream, and record the queuing delay of the data packet in the node buffer; add the processing delay of a single node to the queuing delay to obtain the total node delay; accumulate the total node delays of each node along the transmission path in sequence to obtain the node delay status of the transmission path; divide the node delay status by the preset delay threshold to obtain the node delay score. Monitor the network link status carrying the optimized video bitstream, and obtain the total bandwidth capacity and the occupied bandwidth of the link; calculate the available bandwidth value of each link; select the link with the smallest available bandwidth in the transmission path as the path bandwidth resource; divide the path bandwidth resource by the current bandwidth requirement of the video bitstream to obtain the bandwidth resource score. Weight the node delay score and the bandwidth resource score. Normalize the node delay score and multiply it by the weight coefficient of 0.7, normalize the bandwidth resource score and multiply it by the weight coefficient of 0.3, and add the two to obtain the path priority score.

6. The method according to claim 1, wherein Calculate the time difference between the arrival time and the sending time of the video stream, use the recursive least squares algorithm to model the time difference to obtain a probability distribution model; perform error analysis on the probability distribution model and the video stream delay prediction curve, and generate a delay compensation control strategy through the hierarchical reinforcement learning algorithm; the remote control terminal performs adaptive time compensation on the operation instructions of the crane according to the delay compensation control strategy, including: Calculate the time difference between the arrival time and the sending time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, where the probability distribution model characterizes the statistical characteristics of the video stream delay data; Generate a video stream delay prediction curve according to the probability distribution model, where the video stream delay prediction curve reflects the change trend of the video stream delay data; compare and analyze the change trend with the actually measured delay data to calculate the prediction error; construct an error evaluation function based on the prediction error; Use the error evaluation function as the optimization objective of the hierarchical reinforcement learning algorithm, and use the hierarchical reinforcement learning algorithm to generate a delay compensation control strategy; the delay compensation control strategy dynamically adjusts the execution timing of the control command based on the prediction result of the video stream delay data; The remote control terminal receives the operation command of the operator for the crane, and calculates the time compensation value of the operation command according to the delay compensation control strategy; perform adaptive delay compensation on the operation command based on the time compensation value.

7. The method according to claim 6, wherein Calculate the time difference between the arrival time and the sending time of the video stream to generate video stream delay data; input the video stream delay data into the recursive least squares algorithm model to perform iterative modeling on the video stream delay data; Construct a probability distribution model based on the calculation results of the recursive least squares algorithm model, where the probability distribution model characterizes the statistical characteristics of the video stream delay data including: Calculate the time difference between the arrival time and the sending time of the video stream to generate video stream delay data; perform outlier detection and normalization processing on the video stream delay data to eliminate clock synchronization errors; Construct a recursive least squares algorithm model, input the video stream delay data into the recursive least squares algorithm model, calculate the prediction error and the gain matrix; update the parameter estimation value based on the prediction error and the gain matrix; optimize the covariance matrix according to the parameter estimation value to achieve iterative optimization of the model parameters; Analyze the statistical law of the video stream delay data based on the calculation results of the recursive least squares algorithm model, calculate the mean, variance, skewness and kurtosis of the video stream delay data; select the optimal probability distribution function to fit the video stream delay data and estimate the distribution function parameters; construct a probability distribution model based on the optimal probability distribution function, where the probability distribution model quantitatively characterizes the statistical features of the video stream delay data.

8. A remote control image delay detection system for a crane, which is used to implement the method according to any one of the preceding claims 1-7, characterized in that, Including: The first unit is used to transmit the real-time video stream data to the edge computing server local to the crane; The edge computing server uses a dual-channel parallel preprocessing mechanism to perform spatio-temporal segmentation on the real-time video stream data to obtain a video frame sequence with time stamps; uses an attention mechanism to extract features from the video frame sequence to obtain the three-dimensional spatial coordinates and motion parameters of the crane; based on the three-dimensional spatial coordinates and motion parameters, calculates the motion state vector through an adaptive Kalman filtering algorithm; A second unit for inputting a motion state vector into a multi-modal deep learning model. The multi-modal deep learning model extracts dynamic features of a video stream through a temporal graph convolutional network, deeply fuses the dynamic features with network transmission quality metrics collected in real time, and generates a video stream delay prediction curve; based on the change trend of the video stream delay prediction curve, combines a model predictive control algorithm to calculate an optimal adjustment strategy for video coding parameters; Optimizes the key frame distribution and coding parameters of the video stream according to the optimal adjustment strategy, and transmits the optimized video stream to a remote control terminal through an intelligent routing mechanism; A third unit for calculating the time difference between the arrival time and the sending time of the video stream, modeling the time difference using a recursive least squares algorithm to obtain a probability distribution model; performing error analysis on the probability distribution model and the video stream delay prediction curve, and generating a delay compensation control strategy through a hierarchical reinforcement learning algorithm; the remote control terminal adaptively compensates the operation instructions of the crane according to the delay compensation control strategy.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method of any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method of any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Delay compensation control method for remote control loader

    CN115562045A

  • Operation safety detection method and system for electric power operation site, and storage medium

    CN117075042A

  • Large-scale digital live broadcast method and system with multiple cooperative devices

    CN118678115A

  • Data video stream adaptive processing system and method based on streaming processing technology

    CN118694945A

  • Video coding and decoding acceleration method and system based on learnable task perception mechanism

    CN119031147A

Cited By

  • Bulk cargo storage yard automatic unloading control method and device

    CN121074865A

  • Remote control system for port container crane

    CN122079018A

  • A remote control system for a port container crane

    CN122079018B