A tower crane hook safety early warning method, device and equipment based on machine vision
By optimizing hook image processing using generative adversarial networks and spatiotemporal graph convolutional networks, the problem of poor image quality of suspended objects under complex working conditions was solved. This enabled high-precision prediction of the motion state of suspended objects and accurate delineation of dangerous areas, thereby improving the reliability of tower crane safety early warning and construction efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUCHANG UNIV OF TECH
- Filing Date
- 2025-06-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing tower crane hook safety early warning methods suffer from poor image quality and inaccurate contour recognition under complex working conditions, making it difficult to assess the stress state and dynamic shape changes of the suspended object. This results in inaccurate delineation of dangerous areas, affecting construction efficiency and safety.
A generative adversarial network model is used for super-resolution reconstruction and contour enhancement. By combining adaptive mesh contour and spatiotemporal graph convolutional network, the trajectory of the suspended object is optimized by constructing a loss function and physical constraints, and the dangerous area on the ground is accurately delineated.
Improving the clarity and detail of suspended objects in complex environments, accurately capturing the edges and key stress points of suspended objects, enabling continuous dynamic prediction of the movement state of suspended objects, accurately delineating dangerous areas on the ground, and improving construction safety and efficiency.
Smart Images

Figure CN120841394B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of safety early warning technology, and more specifically, to a method, device and equipment for safety early warning of tower crane hooks based on machine vision. Background Technology
[0002] In the current construction industry, tower cranes, as critical vertical transportation equipment, are of paramount importance for operational safety. However, the complex environment of construction sites, dense populations, and overlapping work processes create significant safety hazards during tower crane operations. Therefore, to eliminate these hazards, it is crucial to quickly identify safe zones around tower cranes and establish an early warning mechanism.
[0003] Traditional tower crane safety monitoring mainly relies on manual observation and some basic sensor warnings. However, the reliability of manual monitoring decreases under complex working conditions such as low light, dust, rain, snow, or partial obstruction of the load. Currently, some tower crane safety warning methods attempt to use machine vision technology to extract the outline of the load and identify dangerous areas. Existing tower crane hook safety warning methods collect images and environmental data to determine the movement trajectory of the hook, thereby achieving all-weather, high-precision safety monitoring of hook operations.
[0004] However, existing tower crane hook safety early warning methods also have some problems. First, under complex working conditions, the quality of the acquired images of the suspended object is poor, leading to inaccurate contour recognition. Second, the contour extraction accuracy for key areas of the suspended object is insufficient, making it difficult to accurately assess its stress state and potential risks. Third, the ability to predict the dynamic shape changes of the suspended object throughout the lifting process is limited, especially for irregular or flexible suspended objects. In addition, the delineation of dangerous areas is often not precise enough, easily resulting in areas that are too large or too small, affecting construction efficiency and safety. Summary of the Invention
[0005] To address at least one deficiency or improvement need in the prior art, the present invention provides a tower crane hook safety early warning method, device, and equipment based on machine vision. This invention solves the problems of poor image quality of the hoisted object acquired under complex working conditions, insufficient accuracy in extracting the contour of key areas, difficulty in accurately assessing its stress state and potential risks, lack of dynamic shape change prediction capability throughout the hoisting process, and inaccurate delineation of dangerous areas, which affects construction efficiency and safety.
[0006] To achieve the above objectives, according to a first aspect of the present invention, a machine vision-based tower crane hook safety early warning method is provided, comprising:
[0007] A generative adversarial network model is constructed to perform super-resolution reconstruction and contour enhancement on the acquired initial tower crane hook image to obtain the target tower crane hook image;
[0008] Calculate the key feature indicators of the target tower crane hook image, and adjust the sampling density based on the key feature indicators to form an adaptive grid contour, thus obtaining a dynamically changing graph structure;
[0009] The graph structure is input into the constructed spatiotemporal graph convolutional network, and a loss function is constructed by combining the actual characteristics of the suspended object to predict the dynamic shape changes of the suspended object;
[0010] By constructing physical constraints to optimize the prediction results of dynamic morphological changes, the predicted trajectory of the suspended object is obtained, and dangerous areas on the ground are delineated and marked for early warning.
[0011] In one possible implementation, the step of calculating key feature indicators of the target tower crane hook image and adjusting the sampling density based on the key feature indicators to form an adaptive grid contour, thereby obtaining a dynamically changing graph structure, further includes:
[0012] Edge detection is performed on the target tower crane hook image to obtain the outline of the suspended object, and an initial set of mesh outline points is constructed based on the outline of the suspended object;
[0013] Calculate the curvature and displacement gradient of each initial mesh profile point based on the initial mesh profile point set;
[0014] The sampling density of each initial mesh contour point is adjusted based on curvature and displacement gradient;
[0015] An adaptive mesh profile is generated based on the adjusted sampling density, and a dynamically changing graph structure is constructed.
[0016] In one possible implementation, the step of generating an adaptive mesh profile based on the adjusted sampling density and constructing a dynamically changing graph structure further includes:
[0017] Based on the adjusted sampling density, sampling is performed along the outline of the suspended object according to the cumulative arc length to obtain several outline sampling points;
[0018] An adaptive mesh profile is generated by calculating the position of each profile sampling point using an interpolation algorithm, and then the adaptive mesh profile is converted into a dynamically changing graph structure.
[0019] In one possible implementation, the step of inputting the graph structure into the constructed spatiotemporal graph convolutional network and constructing a loss function based on the actual characteristics of the suspended object to predict the dynamic morphological changes of the suspended object further includes:
[0020] Extracting node features from graph structures using spatiotemporal graph convolutional networks;
[0021] The prediction results are obtained by mapping node features to multiple future time steps.
[0022] Construct a loss function to calculate the predicted loss between the predicted results and the actual characteristics of the suspended object;
[0023] Predict the dynamic shape changes of the suspended object based on the total training loss value.
[0024] In one possible implementation, the spatiotemporal graph convolutional network includes spatial graph convolutional layers, temporal convolutional layers, and a spatiotemporal fusion layer; extracting node features from the graph structure through the spatiotemporal graph convolutional network further includes:
[0025] Spatial features of nodes are extracted from nodes in a graph structure using spatial graph convolutional layers.
[0026] The temporal features of spatial features are calculated by using temporal convolutional layers;
[0027] A spatiotemporal fusion layer is used to fuse spatial and temporal features to obtain a node feature matrix.
[0028] In one possible implementation, the prediction result includes predicted position, predicted contour, and predicted trajectory; the actual suspended object features include actual position, actual contour, and actual trajectory; the construction of the loss function to calculate the prediction loss between the prediction result and the actual suspended object features further includes:
[0029] Construct a location loss function to calculate the location loss value between the predicted location and the actual location;
[0030] Construct a contour loss function to calculate the contour loss value between the predicted contour and the actual contour;
[0031] Construct a trajectory loss function to calculate the trajectory loss value between the predicted trajectory and the actual trajectory;
[0032] The prediction loss is calculated using the position loss value, contour loss value, and trajectory loss value.
[0033] In one possible implementation, the method of constructing physical constraints to optimize the prediction results of dynamic morphological changes to obtain the predicted trajectory of the suspended object, and delineating and marking dangerous areas on the ground for early warning, further includes:
[0034] Based on physical laws, rigid constraint loss, energy conservation constraint loss, and collision avoidance constraint loss are constructed, and the comprehensive physical constraint loss is calculated.
[0035] The prediction results of dynamic morphological changes are iteratively optimized using comprehensive physical constraint loss until the preset conditions are met, thus obtaining the predicted trajectory of the suspended object.
[0036] Based on the predicted trajectory of the suspended object's movement, dangerous areas on the ground are delineated, and warnings are issued by using preset laser markers to indicate these dangerous areas.
[0037] In one possible implementation, the iterative optimization of the prediction results of dynamic morphological changes using comprehensive physical constraint loss until a preset condition is met to obtain the predicted trajectory of the suspended object's motion target further includes:
[0038] The prediction results of dynamic morphological changes are used as the initial trajectory for initialization.
[0039] The total loss value is calculated based on the combined physical constraint loss and the predicted loss.
[0040] The trajectory of the suspended object is updated by calculating the current trajectory gradient based on the total loss value;
[0041] The iteration optimization stops when the maximum number of iterations is reached, and the predicted trajectory of the suspended object is obtained.
[0042] According to a second aspect of the present invention, a machine vision-based tower crane hook safety early warning device is also provided, for performing the above-described machine vision-based tower crane hook safety early warning method, comprising:
[0043] The image processing module is configured to construct a generative adversarial network model to perform super-resolution reconstruction and contour enhancement processing on the acquired initial tower crane hook image to obtain the target tower crane hook image;
[0044] The grid contour module is configured to calculate key feature indicators of the target tower crane hook image and adjust the sampling density based on the key feature indicators to form an adaptive grid contour, resulting in a dynamically changing graph structure.
[0045] The morphology prediction module is configured to input the graph structure into the constructed spatiotemporal graph convolutional network and combine it with the actual features of the suspended object to construct a loss function to predict the dynamic morphological changes of the suspended object.
[0046] The safety early warning module is configured to construct physical constraints to optimize the prediction results of dynamic shape changes, obtain the predicted trajectory of the suspended object's movement target, delineate and mark dangerous areas on the ground for early warning.
[0047] According to a third aspect of the present invention, a tower crane hook safety early warning device based on machine vision is also provided, including at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit performs the steps of the above-described tower crane hook safety early warning method based on machine vision.
[0048] Compared with the prior art, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0049] This invention provides a machine vision-based tower crane hook safety early warning method. By using a generative adversarial network (GAN) model to perform super-resolution reconstruction and contour enhancement on the initial hook image, it effectively improves the clarity and detail of images of suspended objects acquired under complex construction environments (such as strong light, dust, and occlusion), solving the problem of feature extraction difficulties caused by image blurring in traditional methods. Based on the high-quality image after super-resolution reconstruction, key feature indicators are calculated, and the sampling density is dynamically adjusted through an adaptive mesh contour. This accurately captures the spatial distribution characteristics of the edges of the suspended object and key stress points, overcoming the contour extraction error of traditional fixed-threshold segmentation methods when the shape changes, and improving the accuracy of key area identification. The dynamic graph structure constructed by the adaptive mesh contour is input into a spatiotemporal graph convolutional network, and a loss function is constructed by combining the physical characteristics of the suspended object. This effectively models the dynamic shape changes of the suspended object during hoisting, breaking through the limitations of traditional static analysis methods and achieving continuous and dynamic prediction of the motion state of the suspended object. By introducing physical constraints to optimize the dynamic shape prediction results, and combining mechanical laws with data-driven models, the reliability of the predicted trajectory of the suspended object is improved, thereby accurately delineating the range of dangerous areas on the ground and avoiding the problem of blind spots or false alarms caused by trajectory deviations in traditional methods. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a flowchart illustrating an embodiment of a machine vision-based tower crane hook safety early warning method provided by the present invention.
[0052] Figure 2 Provided by the present invention Figure 1 A flowchart illustrating an embodiment of step 2;
[0053] Figure 3 Provided by the present invention Figure 1 A flowchart illustrating an embodiment of step 3;
[0054] Figure 4 Provided by the present invention Figure 1 A flowchart illustrating an embodiment of step 4;
[0055] Figure 5 A block diagram of an embodiment of the tower crane hook safety early warning device based on machine vision provided by the present invention;
[0056] Figure 6This is a schematic diagram of the structure of a tower crane hook safety early warning device based on machine vision provided in an embodiment of the present invention. Detailed Implementation
[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0058] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0059] This invention provides a method, device, and equipment for safety early warning of tower crane hooks based on machine vision, which will be described in detail below.
[0060] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of a machine vision-based tower crane hook safety early warning method provided by the present invention. In a specific embodiment of the present invention, a machine vision-based tower crane hook safety early warning method is disclosed, comprising:
[0061] Step 1: Construct a generative adversarial network model to perform super-resolution reconstruction and contour enhancement on the acquired initial tower crane hook image to obtain the target tower crane hook image;
[0062] Step 2: Calculate the key feature indicators of the target tower crane hook image, and adjust the sampling density based on the key feature indicators to form an adaptive grid contour, thereby obtaining a dynamically changing graph structure;
[0063] Step 3: Input the graph structure into the constructed spatiotemporal graph convolutional network, and combine it with the actual characteristics of the suspended object to construct a loss function to predict the dynamic shape changes of the suspended object;
[0064] Step 4: Optimize the prediction results of dynamic shape changes by constructing physical constraints to obtain the predicted trajectory of the suspended object, delineate and mark dangerous areas on the ground for early warning.
[0065] In the above embodiments, by constructing a generative adversarial network (GAN) model to perform super-resolution reconstruction and contour enhancement on the initial tower crane hook image, the problem of image quality degradation in complex construction environments is effectively solved. This not only improves image clarity and enhances detail, but also specifically strengthens the edge features and contour information of the hook and the suspended load. The introduction of the GAN enables stable image quality output even under real-world challenges such as changes in lighting, dust interference, and viewpoint occlusion, improving robustness in harsh environments.
[0066] By calculating key feature indicators of the target image and dynamically adjusting the sampling density based on these indicators to form an adaptive mesh contour, a refined description of the suspended object's shape is achieved. The feature extraction strategy can be automatically optimized according to the complexity of the suspended object's shape, ensuring both the ability to capture details in key areas and avoiding redundant consumption of computational resources. The dynamically changing graph structure design enables real-time response to changes in the suspended object's shape, effectively overcoming the limitations of traditional fixed mesh methods in handling shape variations.
[0067] By inputting the graph structure constructed from the adaptive mesh contour into a spatiotemporal graph convolutional network and combining it with the actual physical characteristics of the suspended object to construct a specialized loss function, efficient prediction of the dynamic morphological changes of the suspended object is achieved. The design of the spatiotemporal graph convolutional network integrates information from both spatial and temporal dimensions, enabling it to simultaneously capture the spatial distribution patterns and temporal motion trends of the suspended object. Through customized loss function design, it can more accurately learn the motion patterns of the suspended object, improving the physical rationality and accuracy of the prediction results and providing a more reliable basis for safety early warning.
[0068] By introducing physical constraints to optimize the prediction results of dynamic morphological changes, the reliability of the predicted trajectory is further improved. Combining mechanical principles with a data-driven model ensures that the prediction results not only conform to data patterns but also satisfy basic physical laws, greatly enhancing the credibility of the prediction. Based on the optimized predicted trajectory, dangerous areas on the ground can be accurately delineated and real-time warnings can be issued, providing clear safety guidance for construction personnel and effectively reducing the risk of accidents.
[0069] Compared with existing technologies, this embodiment provides a machine vision-based tower crane hook safety early warning method. By using a generative adversarial network model to perform super-resolution reconstruction and contour enhancement on the initial hook image, it effectively improves the clarity and detail of images of suspended objects acquired under complex construction environments (such as strong light, dust, and occlusion), solving the problem of feature extraction difficulties caused by image blurring in traditional methods. Based on the high-quality image after super-resolution reconstruction, key feature indicators are calculated, and the sampling density is dynamically adjusted through adaptive mesh contours. This accurately captures the spatial distribution characteristics of the edges of the suspended object and key stress points, overcoming the contour extraction error of traditional fixed-threshold segmentation methods when the shape changes, and improving the accuracy of key area identification. The dynamic graph structure constructed by the adaptive mesh contours is input into a spatiotemporal graph convolutional network, and a loss function is constructed by combining the physical features of the suspended object. This effectively models the dynamic shape changes of the suspended object during hoisting, breaking through the limitations of traditional static analysis methods and achieving continuous and dynamic prediction of the motion state of the suspended object. By introducing physical constraints to optimize the dynamic shape prediction results, and combining mechanical laws with data-driven models, the reliability of the predicted trajectory of the suspended object is improved, thereby accurately delineating the range of dangerous areas on the ground and avoiding the problem of blind spots or false alarms caused by trajectory deviations in traditional methods.
[0070] An enhanced generative adversarial network model optimized for images of suspended objects is constructed, consisting of a generator G and a discriminator D. The generator G adopts the UNet architecture and consists of an encoder and a decoder with skip connections in between. The discriminator D adopts the PatchGAN structure to evaluate the realism of the generated images.
[0071] The loss function of the generator G is defined as:
[0072] L G =λ adv ·L adv +λ content ·L content +λ edge ·L edge ;
[0073] Where L G L represents the total loss function of the generator. adv To combat the loss and make the generated images more realistic; L content To avoid content loss, the generated image must be consistent with the content of the original image; L edge For edge loss, special emphasis is placed on preserving and enhancing the outline of the suspended object; λ adv , λ content and λ edge These represent the weight coefficients for adversarial loss, content loss, and edge loss, respectively.
[0074] Optionally, in one specific embodiment, the weighting coefficient can be set to: λ adv =0.2, λ content =0.6 and λ edge =0.4. This configuration is particularly suitable for recognizing large steel components under complex working conditions. For example, when hoisting H-beams in dusty conditions, increasing the weight of edge loss helps maintain the accuracy of the beam's outline. In another alternative implementation, for hoisting scenarios involving small components, the weighting coefficient can be adjusted to: λ adv =0.3, λ content =0.5 and λ edge =0.2.
[0075] In addition, edge loss L edge The calculation formula is:
[0076]
[0077] Where L edge This represents edge loss, measuring the difference between the edges of the generated image and the actual image. This is an operator that sums the values of all pixels. and are the gradients of the i-th pixel in the generated image and the actual image, respectively, where N is the total number of pixels in the image, and ||·||1 represents the L1 norm.
[0078] In some implementations, different edge detection operators can be used to calculate gradients for different lighting conditions at the construction site. For example, the Sobel operator can be used under sufficient lighting conditions, while the Laplacian operator or the Canny edge detection gradient calculation method can be used under low lighting conditions to improve the robustness of edge detection.
[0079] In this embodiment of the application, the training process uses a dataset consisting of paired target tower crane hook images and artificially degraded low-quality images (obtained by adding blur, noise, contrast reduction, local occlusion, etc.) to train the network. During training, the parameters of the generator G and the discriminator D are updated alternately:
[0080] Update discriminator D and minimize it:
[0081] L D =L real +L fake ;
[0082] Where L D L represents the total loss function of the discriminator. real and L fake The classification losses are for the actual image and the generated image, respectively.
[0083] Update generator G, minimize L G .
[0084] In one specific embodiment, the training dataset may contain 2000 pairs of images of suspended objects, with 80% used for training and 20% for validation. Optionally, the artificial degradation methods for low-quality images may be combined according to the characteristics of actual working conditions, for example: for simulated rain and snow weather: adding Gaussian noise (standard deviation σ = 0.1) and motion blur (kernel size of 5×5 pixels); for simulated dusty environments: reducing contrast (reducing by 30% to 50%) and adding Gaussian blur (kernel size of 7×7 pixels); for simulated low-light environments: reducing overall brightness (reducing by 40% to 60%) and adding salt-and-pepper noise (noise ratio of 0.05); for simulated occlusion: randomly occluding 10% to 30% of the image area, prioritizing non-critical areas.
[0085] In addition, the training parameters can be set as follows:
[0086] Batch size: 16 to 32; Learning rate: Initially 0.0002, using the Adam optimizer, decaying to 0.8 times the original rate every 50 epochs; Number of training epochs: Usually 100-200 epochs depending on the performance of the validation set; Early stopping strategy: Stop training if the loss on the validation set does not decrease after 10 consecutive epochs.
[0087] According to the embodiments of this application, the low-quality suspended object image I low The image is input into the trained generator G to obtain a high-quality enhanced image.
[0088] I enhanced =G(I low );
[0089] Among them I low This represents a low-quality image of a suspended object acquired in real time; I enhanced G(I) represents the enhanced high-quality image; low ) represents the generator G on the input image I low The processing results.
[0090] In cases of partial occlusion, the model infers the possible shape of the occluded part based on the learned features of the hanging object, thus achieving continuity and integrity of the contour.
[0091] Please see Figure 2 , Figure 2 Provided by the present invention Figure 1A flowchart illustrating one embodiment of step 2 of the present invention shows that, in some embodiments of the invention, key feature indicators of the target tower crane hook image are calculated, and the sampling density is adjusted based on the key feature indicators to form an adaptive grid contour, resulting in a dynamically changing graph structure. The method further includes:
[0092] Step 201: Perform edge detection on the target tower crane hook image to obtain the outline of the suspended object, and construct an initial mesh outline point set based on the outline of the suspended object;
[0093] Step 202: Calculate the curvature and displacement gradient of each initial mesh profile point based on the initial mesh profile point set;
[0094] Step 203: Adjust the sampling density of each initial mesh profile point according to the curvature and displacement gradient;
[0095] Step 204: Generate an adaptive mesh profile based on the adjusted sampling density and construct a dynamically changing graph structure.
[0096] In the above embodiment, the target tower crane hook image I enhanced An initial contour is obtained using an edge detection algorithm (such as the Canny algorithm), and then a contour tracking algorithm is used to extract the closed contour point set of the suspended object. After obtaining the initial mesh contour point set, a model with a uniform sampling density ρ is constructed. basse A coarse grid is used to represent the profile as a set of initial grid profile points, where the distance between adjacent points is approximately equal.
[0097] For each point q on the initial mesh profile point set i Calculate the following two key characteristic indicators:
[0098] The formula for calculating curvature is:
[0099]
[0100] Where k(q) i ) indicates the degree of curvature of the profile at that point, q i Let x be the i-th point on the contour. i and y i q i x and y coordinates; x′ i 、x″ i y′ i y″ i q i The first derivative of the x-coordinate, the second derivative of the x-coordinate, the first derivative of the y-coordinate, and the second derivative of the y-coordinate.
[0101] The formula for calculating the displacement gradient is:
[0102]
[0103] in This represents the rate of change of displacement at that point compared to the previous moment; Let ||·|| denote the gradient operator; ||·|| denotes the magnitude of the vector. and Point q i Partial derivatives of displacement with respect to time in the x and y directions.
[0104] Based on the calculated curvature and displacement gradient, for each point q i Adaptive adjustment of sampling density ρ(q) i The formula is:
[0105]
[0106] Where ρ(q) i ) represents point q i Adaptive sampling density at ρ base Based on the sampling density, α k and k(q) represents the curvature influence coefficient and the displacement gradient influence coefficient, respectively; i ) represents point q i Curvature at that point; Representing point q i The displacement gradient at that location.
[0107] It should be understood that this formula ensures a higher sampling density in areas with large curvature or large displacement gradients (i.e., key areas such as connecting nodes and cantilevered parts).
[0108] In a preferred embodiment, different parameter configurations can be selected according to different types of suspended structures:
[0109] For example, in a certain H-beam hoisting application, when the beam was being lifted from the ground, ρ was selected. base = 12 pixels / dot, α k =2.8, This configuration allows for a sampling density of approximately one point every 3-4 pixels at the flange-web junction of the H-beam, while maintaining a lower density of one point every 10-12 pixels in the straight flange section. This ensures accurate contour extraction of critical areas while reducing computational burden.
[0110] Optionally, this embodiment may also employ an adaptive adjustment mechanism to dynamically adjust parameters based on the overall contour complexity calculated in real time. For example, when the average curvature of the detected contour exceeds a preset threshold (e.g., 0.05), α is automatically reduced. k To prevent oversampling, the value can be increased appropriately when the variance of the detected displacement gradient is large. Value, strengthen attention to areas of intense exercise.
[0111] In some embodiments of the present invention, generating an adaptive mesh profile based on the adjusted sampling density and constructing a dynamically changing graph structure further includes:
[0112] Based on the adjusted sampling density, sampling is performed along the outline of the suspended object according to the cumulative arc length to obtain several outline sampling points;
[0113] An adaptive mesh profile is generated by calculating the position of each profile sampling point using an interpolation algorithm, and then the adaptive mesh profile is converted into a dynamically changing graph structure.
[0114] In the above embodiments, the new sampling point position is calculated along the contour according to the cumulative arc length. In areas with high sampling density, the sampling point spacing is small. For each new sampling point, its precise position is calculated by an interpolation algorithm (such as cubic spline interpolation).
[0115] In this embodiment of the application, the adaptive mesh profile G is... adaptive Convert to a graph structure:
[0116] G = (V, E);
[0117] Where G represents the graph structure of the outline; V represents the set of nodes in the graph; and E represents the set of edges in the graph.
[0118] The node set V is represented as:
[0119] V = v1, v2, ..., v k ;
[0120] Among them, v1, v2, v k These represent the 1st, 2nd, and kth nodes in the diagram, respectively; k is the total number of sampling points for the suspended object's outline.
[0121] Edge set E is represented as:
[0122] E = e ij ;
[0123] Where e ij Represents node v i and v j The edge between, v i and v j Let i and j represent the i-th and j-th nodes, respectively.
[0124] It should be understood that, in order to construct a reasonable edge set E, this application adopts the following strategy:
[0125] Adjacent connection: Connects adjacent sampling points on the contour, that is, for each node v i , create with v i-1 and v i+1Connections (circular connections between first and last nodes); Skip connections: establish skip connections with a certain step size s (e.g., s=3 or s=5), i.e., node v i With v i±s Establish connections to capture a wider range of structural information; key point connections: for key points with large curvature or displacement gradients, increase connections with other key points to enhance the model's ability to perceive key regions.
[0126] In addition, each node v i eigenvector f i Includes the following information:
[0127] Normalized spatial coordinates (x i y i ); curvature k(r) i ) and its rate of change within a certain range; displacement gradient And its time variation trend; local shape descriptors (such as local curvature distribution, edge intensity, etc.).
[0128] Please see Figure 3 , Figure 3 Provided by the present invention Figure 1 A flowchart illustrating one embodiment of step 3 of the present invention. In some embodiments of the present invention, the graph structure is input into the constructed spatiotemporal graph convolutional network, and a loss function is constructed in conjunction with the actual characteristics of the suspended object to predict the dynamic morphological changes of the suspended object. The method further includes:
[0129] Step 301: Extract node features from the graph structure using a spatiotemporal graph convolutional network;
[0130] Step 302: Map the node features to multiple future time steps to obtain the prediction results;
[0131] Step 303: Construct a loss function to calculate the predicted loss between the predicted results and the actual characteristics of the suspended object;
[0132] Step 304: Predict the dynamic shape changes of the suspended object based on the total training loss value.
[0133] In the above embodiments, the spatiotemporal graph convolutional network extracts node features from the graph structure constructed by the adaptive mesh contour, which can effectively capture the spatial relationships and interactions between key nodes of the suspended object, while also considering continuous changes in the time dimension. Compared with traditional convolutional neural networks or recurrent neural networks, spatiotemporal graph convolutional networks can better model the topological connections between different parts of the suspended object, and are particularly suitable for handling complex motion patterns such as deformation, rotation, and swaying that may occur during the hoisting process.
[0134] By mapping the extracted node features to multiple future time steps to obtain prediction results, the system achieves short-term and medium-term prediction capabilities for the motion of suspended objects. It can not only predict the position and shape of the suspended object at the next moment but also predict its motion trend over a longer future time range, providing a longer reaction time window for safety warnings. By considering continuous change patterns in the time dimension, it can more accurately capture the dynamic characteristics of the suspended object, such as its motion inertia, acceleration changes, and possible swing periods, thereby improving the consistency and reliability of the prediction results and effectively avoiding the error accumulation problem that may occur with single-step prediction.
[0135] By constructing a specialized loss function to calculate the prediction loss between the predicted results and the actual characteristics of the hoisted object, the model training process is ensured to closely align with the needs of actual hoisting tasks. It not only considers the difference between the predicted and actual positions but also incorporates the kinematic characteristics and physical constraints of the hoisted object, enabling the model to learn motion patterns that better match real-world hoisting scenarios. By comprehensively considering various error factors, the loss function guides the model to focus on features and patterns crucial for safety warnings, improving the model's prediction accuracy in key areas while suppressing secondary errors that might affect safety judgments, making the prediction results more consistent with actual engineering requirements.
[0136] By predicting the dynamic morphological changes of suspended objects based on the total training loss value, end-to-end optimization of model parameters was achieved. This allows the network to automatically balance the learning objectives between different prediction time steps and different feature dimensions, improving the model's generalization ability and robustness. Through iterative optimization, the model can gradually correct prediction biases, adapt to the motion characteristics of different types of suspended objects, and maintain stable prediction performance when facing new lifting scenarios. It can continuously improve prediction quality, and with the accumulation of training data and the increase in scenario diversity, prediction accuracy will be further improved, providing increasingly reliable support for safety early warning.
[0137] In some embodiments of the present invention, the spatiotemporal graph convolutional network includes a spatial graph convolutional layer, a temporal convolutional layer, and a spatiotemporal fusion layer; extracting node features from the graph structure using the spatiotemporal graph convolutional network further includes:
[0138] Spatial features of nodes are extracted from nodes in a graph structure using spatial graph convolutional layers.
[0139] The temporal features of spatial features are calculated by using temporal convolutional layers;
[0140] A spatiotemporal fusion layer is used to fuse spatial and temporal features to obtain a node feature matrix.
[0141] In the above embodiments, a graph convolutional neural network containing temporal and spatial dimensions is constructed, called a Spatio-Temporal Graph Convolutional Network (ST-GCN). The network structure mainly includes:
[0142] Spatial graph convolutional layer: Extracts spatial features from nodes in the graph, using the following formula:
[0143]
[0144] in H represents the spatial features of the nodes in the (l+1)th layer after spatial graph convolution; (l) A represents the node feature matrix of the l-th layer; A represents the adjacency matrix, which represents the connection relationships between nodes in the graph; D represents the degree matrix, with the diagonal elements being the degree (number of connections) of each node. This represents the degree matrix raised to the power of -1 / 2, used for normalization; σ represents the weight matrix of the l-th spatial graph convolution; σ represents the activation function.
[0145] Temporal convolutional layers: capture the changes in node features over time, with the following formula:
[0146]
[0147] in σ represents the temporal features of the nodes in the (l+1)th layer after temporal convolution; σ represents the activation function; Conv1D represents a one-dimensional convolution operation that extracts features in the temporal dimension. This represents the node feature matrix after convolution of the l-th spatial graph; This represents the weight matrix of the l-th temporal convolution layer;
[0148] Spatiotemporal fusion layer: fuses spatial and temporal features, using the following formula:
[0149]
[0150] Where H (l+1) α represents the node feature matrix of the (l+1)th layer after fusion. fuse H represents the fusion weight parameter, which controls the relative importance of spatial and temporal features, and its value ranges from [0, 1]. (l) This represents the node feature matrix of the l-th layer.
[0151] In one specific embodiment, the spatiotemporal graph convolutional network can be configured as follows:
[0152] Network depth: 6 to 8 spatiotemporal graph convolutional blocks; Feature dimension of each spatiotemporal graph convolutional layer: 32-64 (first layer) increasing to 128-256 (last layer); Temporal convolutional kernel size: 3 or 5, covering adjacent time steps; Fusion parameter α fuse The value can be set to 0.6-0.7, slightly biased towards spatial features; the prediction head MLP structure consists of two fully connected layers, with a hidden layer dimension of 256 and an output layer dimension of 2×T (predicting the x and y coordinates at T time steps).
[0153] Optionally, different network variants can be used for different types of suspended objects. For example, in a real tower crane hoisting project, the system predicts the process of lifting a 12-meter-long steel pipe from horizontal to vertical. A 7-layer spatiotemporal graph convolutional network is used, with an initial feature dimension of 48 and a final feature dimension of 192, and a fusion parameter α. fuse =0.65, and the temporal convolution kernel size is 5.
[0154] In some embodiments of the present invention, the prediction result includes predicted position, predicted contour, and predicted trajectory; the actual suspended object features include actual position, actual contour, and actual trajectory; constructing a loss function to calculate the prediction loss between the prediction result and the actual suspended object features further includes:
[0155] Construct a location loss function to calculate the location loss value between the predicted location and the actual location;
[0156] Construct a contour loss function to calculate the contour loss value between the predicted contour and the actual contour;
[0157] Construct a trajectory loss function to calculate the trajectory loss value between the predicted trajectory and the actual trajectory;
[0158] The prediction loss is calculated using the position loss value, contour loss value, and trajectory loss value.
[0159] In the above embodiments, the loss function includes the following parts:
[0160] Location prediction loss: measures the difference between the predicted location and the actual location.
[0161]
[0162] Where L position This represents the location prediction loss, which measures the difference between the predicted location and the actual location; T represents the total number of time steps in the prediction. This indicates that the average is taken over all time steps; This represents the summation of all predicted time steps; k is the total number of sampling points for the suspended object's profile. This represents summing over all nodes; Represents a node The predicted position at time step t+i1; Represents a node At the actual position of time step t+i1; This represents the square of the Euclidean distance.
[0163] Shape consistency loss: Ensures that the predicted profile shape is consistent with the physical properties of the suspended object.
[0164]
[0165] Where L shape This represents the shape consistency loss, ensuring that the predicted contour is similar in shape to the actual contour; T represents the total number of time steps in the prediction. This indicates that the average is taken over all time steps; This represents the summation over all predicted time steps; This represents the predicted contour area at time step t+i2; This represents the actual contour area at time step t+i2; λ represents the absolute difference between the area ratio and 1, measuring the area change; λ represents the weighting coefficient, balancing area loss and shape distribution loss; PSD(·) represents the power spectral density function, describing the shape characteristics of the contour; KL(·,·) represents the KL divergence, measuring the difference between two probability distributions.
[0166] Smoothing loss: Ensures the smoothness of the predicted trajectory and avoids unreasonable jitter.
[0167]
[0168] Where L smooth T represents the smoothing loss, ensuring the smoothness of the predicted trajectory; T-1 represents the number of time steps considered (acceleration needs to be calculated in three consecutive time steps); This indicates that the average time step is taken over all considered time steps; This represents the summation over all considered time steps; k is the total number of sampling points for the suspended object's profile. This indicates that the average is taken over all nodes; This represents summing over all nodes; Represents a node The displacement from time step t+i3 to t+i3+1; Represents a node The displacement from time step t+i3-1 to t+i3; It represents the change in velocity, i.e., acceleration; This represents the square of the Euclidean distance, i.e., the square of the L2 norm.
[0169] The final training loss is the weighted sum of the above three factors:
[0170] L total =w1·Lposition +w2·L shape +w3·L smooth ;
[0171] Where L total w1 represents the total training loss; w2 represents the weight coefficient of the position prediction loss; w3 represents the weight coefficient of the shape consistency loss; and w4 represents the weight coefficient of the smoothness loss.
[0172] Please see Figure 4 , Figure 4 Provided by the present invention Figure 1 A flowchart illustrating one embodiment of step 4 of the present invention. In some embodiments of the present invention, the predicted trajectory of the suspended object is obtained by optimizing the prediction results of dynamic shape changes based on the physical constraints, and the ground danger zone is delineated and marked for early warning. The method also includes:
[0173] Step 401: Construct rigid constraint loss, energy conservation constraint loss, and collision avoidance constraint loss based on physical laws, and calculate the comprehensive physical constraint loss;
[0174] Step 402: Iteratively optimize the prediction results of dynamic morphological changes using the comprehensive physical constraint loss until the preset conditions are met, and obtain the predicted trajectory of the suspended object's motion target.
[0175] Step 403: Delineate the dangerous areas on the ground based on the predicted trajectory of the suspended object's movement, and issue early warnings by using preset laser markers to mark the dangerous areas on the ground.
[0176] In the above embodiments, a physical constraint loss function is constructed to ensure that the prediction results conform to the basic laws of the real physical world:
[0177] rigid constraint loss L rigid This refers to the degree of shape deformation constraint applied to rigid objects such as metal components. The calculation formula is as follows:
[0178]
[0179] Where L rigid N represents the loss due to rigid constraints. e e represents the total number of edges; ij Let E represent the edge connecting nodes i and j; let E represent the set of all edges. and These represent the Euclidean distances between nodes i and j at time steps t and t-1, respectively. Represents the square of the Euclidean distance; This represents summing the set of all edges;
[0180] It should be understood that this constraint ensures that the relative distance between points inside the rigid member changes little during the movement process.
[0181] Energy conservation loss L energy This ensures that the predicted trajectory conforms to the principle of energy conservation; the calculation formula is as follows:
[0182]
[0183] Where L engrgy This represents the loss due to energy conservation. and Let represent the kinetic energy at time steps t and t-1, respectively. and Let represent the potential energy at time steps t and t-1, respectively; (·) represents the absolute value.
[0184]
[0185] in The kinetic energy represents the time step t; m represents the estimated mass of the suspended object; k is the total number of sampling points on the profile of the suspended object. This represents the velocity of node i4 at time step t; Represents the square of the Euclidean distance; This represents summing over all nodes;
[0186]
[0187] in The potential energy represents time step t; m represents the estimated mass of the suspended object; g represents the gravitational acceleration; and k is the total number of sampling points on the profile of the suspended object. This represents the vertical height coordinate of node i5 at time step t; This represents summing over all nodes;
[0188] Collision avoidance loss L collision This means that based on the scene segmentation results, we ensure that the predicted outline of the suspended object does not collide with the ground engineering structure.
[0189]
[0190] Where L collision This represents the collision avoidance loss; k is the total number of sampling points for the suspended object's outline; (x,y,c,p) represents a pixel in the scene, its category, and confidence level, where x, y, c, and p represent the pixel's horizontal coordinate, vertical coordinate, category label, and confidence level, respectively; S scene Represents a semantic segmentation graph of the scene; This represents the position of node i6 at time step t; φ(·) represents the distance weighting function, used to calculate the collision penalty value based on the distance between the node and scene objects; ||·||2 represents the Euclidean distance, used to calculate the distance between the node position and pixels in the scene; This represents summing over all nodes; This represents the summation of all pixels in the scene.
[0191] Combined physical constraint loss: The above three physical constraints are combined into a comprehensive loss function:
[0192] L physics =λ rigid ·L rigid +λ energy ·L energy +λ collision ·L collision ;
[0193] Where L physics λ represents the overall physical constraint loss; rigid The weighting coefficient representing the loss due to rigid constraints; λ energy The weighting coefficient representing the energy conservation loss; λ collision L represents the weighting coefficient for collision avoidance loss. rigid L represents the loss due to rigid constraints. energy L represents the energy conservation loss; collision This indicates a collision to avoid damage;
[0194] In a preferred embodiment, the distance weighting function φ(d) can be defined as:
[0195]
[0196] Where φ(d) represents the distance weighting function; d represents the current distance; d safe This indicates the safe distance threshold; if indicates a conditional judgment.
[0197] Optionally, different parameter configurations can be used for different types of suspended loads. For example, in a practical application, for the hoisting of a 22-meter span truss, the system used λ rigid =0.7, λ energy =0.25, λ collision =0.9 and d safe The system features a parameter configuration of 1.3 meters. During hoisting, even when the truss approaches a supporting column, the system maintains the overall rigidity of the truss and avoids collisions with the column. The minimum safe distance for the predicted trajectory remains above 1.1 meters, effectively ensuring construction safety.
[0198] In some embodiments of the present invention, the prediction results of dynamic morphological changes are iteratively optimized using comprehensive physical constraint loss until a preset condition is met to obtain the predicted trajectory of the suspended object's motion target, and the method further includes:
[0199] The prediction results of dynamic morphological changes are used as the initial trajectory for initialization.
[0200] The total loss value is calculated based on the combined physical constraint loss and the predicted loss.
[0201] The trajectory of the suspended object is updated by calculating the current trajectory gradient based on the total loss value;
[0202] The iteration optimization stops when the maximum number of iterations is reached, and the predicted trajectory of the suspended object is obtained.
[0203] In the above embodiments, the predicted trajectory is optimized using a physical constraint loss function to make it more consistent with physical laws:
[0204] Trajectory initialization: The prediction results of the spatiotemporal graph convolutional network are used as the initial trajectory.
[0205] G pred,t+1 G pred,t+2 , ..., G pred,t+T ;
[0206] Among them G pred,t+1 G pred,t+2 G pred,t+T These represent the predicted contour meshes at time t+1, t+2, and t+T, respectively, where T represents the total time step of the prediction.
[0207] Iterative optimization process: Minimize the overall loss function using the gradient descent algorithm.
[0208] L = L pred +L physics ;
[0209] Where L represents the total loss function; L pred Indicates predicted loss; L phvsics Indicates physical constraint loss;
[0210] Calculate the loss function value L for the current trajectory, and calculate the gradient of the loss function with respect to the trajectory. Then update the trajectory:
[0211]
[0212] Where η lr The learning rate; ← represents the gradient of the loss function with respect to the trajectory G; ← indicates an assignment operation;
[0213] Repeat the above steps until convergence or the maximum number of iterations is reached.
[0214] Constraint Satisfaction Check: Perform physical constraint checks on the optimized trajectory to ensure that the following conditions are met:
[0215] Rigid constraints: The relative distance within the components is within the allowable range; Energy conservation: Energy changes are within the physically reasonable range; Collision-free: A safe distance is maintained from ground engineering structures.
[0216] To better implement the machine vision-based tower crane hook safety early warning method in this embodiment of the invention, based on the machine vision-based tower crane hook safety early warning method, please refer to the corresponding... Figure 5 , Figure 5 This is a block diagram of an embodiment of the tower crane hook safety early warning device based on machine vision provided by the present invention. The embodiment of the present invention provides a tower crane hook safety early warning device 500 based on machine vision, comprising:
[0217] Image processing module 510 is configured to construct a generative adversarial network model to perform super-resolution reconstruction and contour enhancement processing on the acquired initial tower crane hook image to obtain the target tower crane hook image;
[0218] The mesh contour module 520 is configured to calculate key feature indicators of the target tower crane hook image and adjust the sampling density based on the key feature indicators to form an adaptive mesh contour, resulting in a dynamically changing graph structure.
[0219] The morphology prediction module 530 is configured to input the graph structure into the constructed spatiotemporal graph convolutional network and combine the actual features of the suspended object to construct a loss function to predict the dynamic morphological changes of the suspended object.
[0220] The safety early warning module 540 is configured to construct physical constraints to optimize the prediction results of dynamic shape changes, obtain the predicted trajectory of the suspended object's movement target, delineate and mark dangerous areas on the ground for early warning.
[0221] It should be noted that the device 500 provided in the above embodiments can implement the technical solutions described in the above method embodiments. The specific implementation principles of the above modules or units can be found in the corresponding content in the above method embodiments, and will not be repeated here.
[0222] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a tower crane hook safety early warning device based on machine vision provided in an embodiment of the present invention. Based on the above-described tower crane hook safety early warning method based on machine vision, the present invention also provides a tower crane hook safety early warning device based on machine vision. This machine vision-based tower crane hook safety early warning device can be a computing device such as a mobile terminal, desktop computer, laptop, handheld computer, or server. The machine vision-based tower crane hook safety early warning device 600 includes a processor 610, a memory 620, and a display 630. Figure 6Only some components of the machine vision-based tower crane hook safety warning device are shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.
[0223] In some embodiments, the memory 620 can be an internal storage unit of the machine vision-based tower crane hook safety warning device 600, such as a hard disk or memory of the machine vision-based tower crane hook safety warning device 600. In other embodiments, the memory 620 can also be an external storage device of the machine vision-based tower crane hook safety warning device 600, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, FlashCard, etc., equipped on the machine vision-based tower crane hook safety warning device 600. Furthermore, the memory 620 can include both internal and external storage units of the machine vision-based tower crane hook safety warning device 600. The memory 620 is used to store application software and various types of data installed on the machine vision-based tower crane hook safety warning device 600, such as the program code of the machine vision-based tower crane hook safety warning device 600. The memory 620 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 620 stores a machine vision-based tower crane hook safety warning program 640, which can be executed by the processor 610 to realize a machine vision-based tower crane hook safety warning method according to various embodiments of this application.
[0224] In some embodiments, processor 610 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in memory 620 or process data, such as executing a machine vision-based tower crane hook safety early warning method.
[0225] In some embodiments, display 630 may be an LED display, a liquid crystal display, a touch-screen liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 630 is used to display information from the machine vision-based tower crane hook safety warning device 600 and to display a visual user interface. Components 610-630 of the machine vision-based tower crane hook safety warning device 600 communicate with each other via a bus.
[0226] In one embodiment, when the processor 610 executes the machine vision-based tower crane hook safety warning program 640 in the memory 620, it implements the steps of the machine vision-based tower crane hook safety warning method described above.
[0227] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
[0228] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0229] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A machine vision-based safety early warning method for tower crane hooks, characterized in that, include: A generative adversarial network model is constructed to perform super-resolution reconstruction and contour enhancement on the acquired initial tower crane hook image to obtain the target tower crane hook image; Calculate the key feature indicators of the target tower crane hook image, and adjust the sampling density based on the key feature indicators to form an adaptive grid contour, thus obtaining a dynamically changing graph structure; The graph structure is input into a constructed spatiotemporal graph convolutional network, and a loss function is constructed based on the actual features of the suspended object to predict the dynamic morphological changes of the suspended object. Specifically, the spatiotemporal graph convolutional network extracts node features from the graph structure, maps the node features to multiple future time steps to obtain prediction results, constructs a loss function to calculate the prediction loss between the prediction results and the actual features of the suspended object, and predicts the dynamic morphological changes of the suspended object based on the total training loss value. The spatiotemporal graph convolutional network includes spatial graph convolutional layers, temporal convolutional layers, and a spatiotemporal fusion layer. Extracting node features from the graph structure through the spatiotemporal graph convolutional network includes: extracting spatial features of nodes from the nodes in the graph structure through the spatial graph convolutional layer, and calculating the spatial features through the temporal convolutional layer. The temporal features after temporal convolution are fused with spatial and temporal features using a spatiotemporal fusion layer to obtain a node feature matrix. The prediction result includes predicted position, predicted contour, and predicted trajectory. The actual suspended object features include actual position, actual contour, and actual trajectory. The calculation of the prediction loss between the prediction result and the actual suspended object features using a loss function includes: constructing a position loss function to calculate the position loss value between the predicted position and the actual position; constructing a contour loss function to calculate the contour loss value between the predicted contour and the actual contour; and constructing a trajectory loss function to calculate the trajectory loss value between the predicted trajectory and the actual trajectory. The position loss value, contour loss value, and trajectory loss value are then used to calculate the prediction loss. By constructing physical constraints to optimize the prediction results of dynamic morphological changes, the predicted trajectory of the suspended object is obtained, and dangerous areas on the ground are delineated and marked for early warning.
2. The tower crane hook safety early warning method based on machine vision as described in claim 1, characterized in that, The calculation of key feature indicators of the target tower crane hook image, and the adjustment of sampling density based on these key feature indicators to form an adaptive grid contour, resulting in a dynamically changing graph structure, includes: Edge detection is performed on the target tower crane hook image to obtain the outline of the suspended object, and an initial set of mesh outline points is constructed based on the outline of the suspended object; Calculate the curvature and displacement gradient of each initial mesh profile point based on the initial mesh profile point set; The sampling density of each initial mesh contour point is adjusted based on curvature and displacement gradient; An adaptive mesh profile is generated based on the adjusted sampling density, and a dynamically changing graph structure is constructed.
3. The tower crane hook safety early warning method based on machine vision as described in claim 2, characterized in that, The process of generating an adaptive mesh profile based on the adjusted sampling density and constructing a dynamically changing graph structure includes: Based on the adjusted sampling density, sampling is performed along the outline of the suspended object according to the cumulative arc length to obtain several outline sampling points; An adaptive mesh profile is generated by calculating the position of each profile sampling point using an interpolation algorithm, and then the adaptive mesh profile is converted into a dynamically changing graph structure.
4. The tower crane hook safety early warning method based on machine vision as described in claim 1, characterized in that, The method of constructing physical constraints to optimize the prediction results of dynamic morphological changes yields the predicted trajectory of the suspended object, and delineates and marks dangerous areas on the ground for early warning, including: Based on physical laws, rigid constraint loss, energy conservation constraint loss, and collision avoidance constraint loss are constructed, and the comprehensive physical constraint loss is calculated. The prediction results of dynamic morphological changes are iteratively optimized using comprehensive physical constraint loss until the preset conditions are met, thus obtaining the predicted trajectory of the suspended object. Based on the predicted trajectory of the suspended object's movement, dangerous areas on the ground are delineated, and warnings are issued by using preset laser markers to indicate these dangerous areas.
5. A tower crane hook safety early warning method based on machine vision as described in claim 4, characterized in that, The method of iteratively optimizing the prediction results of dynamic morphological changes using comprehensive physical constraint loss until a preset condition is met to obtain the predicted trajectory of the suspended object's motion target includes: The prediction results of dynamic morphological changes are used as the initial trajectory for initialization. The total loss value is calculated based on the combined physical constraint loss and the predicted loss. The trajectory of the suspended object is updated by calculating the current trajectory gradient based on the total loss value; The iteration optimization stops when the maximum number of iterations is reached, and the predicted trajectory of the suspended object is obtained.
6. A tower crane hook safety early warning device based on machine vision, characterized in that, A tower crane hook safety early warning method based on machine vision as described in any one of claims 1-5 includes: The image processing module is configured to construct a generative adversarial network model to perform super-resolution reconstruction and contour enhancement processing on the acquired initial tower crane hook image to obtain the target tower crane hook image; The grid contour module is configured to calculate key feature indicators of the target tower crane hook image and adjust the sampling density based on the key feature indicators to form an adaptive grid contour, resulting in a dynamically changing graph structure. The morphology prediction module is configured to input the graph structure into a constructed spatiotemporal graph convolutional network and combine it with the actual features of the suspended object to construct a loss function to predict the dynamic morphological changes of the suspended object. Specifically, it extracts node features from the graph structure through the spatiotemporal graph convolutional network, maps the node features to multiple future time steps to obtain prediction results, constructs a loss function to calculate the prediction loss between the prediction results and the actual features of the suspended object, and predicts the dynamic morphological changes of the suspended object based on the total training loss value. The spatiotemporal graph convolutional network includes spatial graph convolutional layers, temporal convolutional layers, and a spatiotemporal fusion layer. Extracting node features from the graph structure through the spatiotemporal graph convolutional network includes: extracting spatial features of nodes from the graph structure through the spatial graph convolutional layer, and extracting spatial features through the temporal convolutional layer... The spatial features are calculated as temporal features after temporal convolution. A spatiotemporal fusion layer is used to fuse the spatial and temporal features to obtain a node feature matrix. The prediction result includes predicted position, predicted contour, and predicted trajectory. The actual suspended object features include actual position, actual contour, and actual trajectory. The loss function is constructed to calculate the prediction loss between the prediction result and the actual suspended object features. This includes: constructing a position loss function to calculate the position loss value between the predicted position and the actual position; constructing a contour loss function to calculate the contour loss value between the predicted contour and the actual contour; and constructing a trajectory loss function to calculate the trajectory loss value between the predicted trajectory and the actual trajectory. The position loss value, contour loss value, and trajectory loss value are then used to calculate the prediction loss. The safety early warning module is configured to construct physical constraints to optimize the prediction results of dynamic shape changes, obtain the predicted trajectory of the suspended object's movement target, delineate and mark dangerous areas on the ground for early warning.
7. A tower crane hook safety early warning device based on machine vision, characterized in that, It includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit performs the steps of the machine vision-based tower crane hook safety early warning method according to any one of claims 1-5.
Citation Information
Patent Citations
Intelligent wardrobe based on RFID (Radio Frequency Identification) technology and control method thereof
CN110384336A
Early warning method and device for tower crane
CN110745704A