Type identification method, device and equipment for hoisting object of tower crane and storage medium

By acquiring tower crane status information and using image processing technology, valid lifting operations are determined, and video preprocessing and feature fusion are performed. This solves the problem of low accuracy in tower crane object identification and achieves efficient and accurate object type identification.

CN121747007APending Publication Date: 2026-03-27CHINA CONSTR THIRD ENG BUREAU GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional tower crane hoisting operations rely on manual visual observation, resulting in low accuracy in identifying hoisted objects and making it difficult to meet the demands of high-frequency hoisting tasks in complex construction scenarios.

Method used

By acquiring tower crane status information to determine valid lifting operations, performing operation video preprocessing, and using image enhancement technology and feature fusion methods to screen candidate regions and analyze cases, the location and type of the target object are accurately located.

Benefits of technology

It improves the accuracy and efficiency of tower crane object identification, reduces positioning deviations caused by blind spots and visual fatigue, and adapts to high-frequency lifting tasks in complex construction environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747007A_ABST
    Figure CN121747007A_ABST
Patent Text Reader

Abstract

The invention discloses a type identification method, device and equipment for a hoisted object of a tower crane and a storage medium, and the method comprises the following steps: determining an effective hoisting time of the tower crane according to state information of the tower crane, and obtaining an operation video of the tower crane under the effective hoisting time; carrying out preprocessing and hanging object identification on the operation video to obtain a candidate region set corresponding to each video frame in the operation video; performing detection and identification processing on each candidate region in the candidate region set, and performing instance analysis processing on each candidate region according to a processing result obtained by processing to obtain instance features of instances contained in the corresponding video frame; and analyzing and processing the instance characteristics, determining the target hoisting object and the position information of the target hoisting object in the instance, and determining the hoisting object type of the target hoisting object under the effective hoisting times according to the position information. And the accuracy and the efficiency of identifying the hoisting object of the tower crane during the operation of the tower crane are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology for tower crane operations, and in particular to a method, apparatus, equipment and storage medium for identifying the type of objects hoisted by a tower crane. Background Technology

[0002] Tower cranes, as core equipment in modern intelligent construction systems, profoundly impact the digital management capabilities of the entire construction process due to their level of intelligence. In construction scenarios where industrialization and informatization are deeply integrated, tower cranes are not only key execution units for material transportation but also crucial hubs for dynamic data interaction on construction sites. Lifting target recognition technology, as a core component of tower crane intelligent upgrades, directly relates to key areas such as construction efficiency optimization, safety risk prevention and control, and resource collaborative management. Especially in complex construction scenarios, the accurate identification of hoisted object characteristics is a fundamental prerequisite for ensuring component installation accuracy, avoiding collision risks, and achieving dynamic path planning. It plays an irreplaceable supporting role in building an integrated intelligent construction closed loop of "perception-decision-execution."

[0003] Traditional hoisting operations rely primarily on operators' visual observation and experience-based decision-making, which has several inherent limitations. Spatially, the geometric constraints between the fixed viewpoint of the tower crane operator's cab and the dynamic trajectory of the load make it difficult for manual visual assessment to accurately determine the spatial orientation of the hook and the target lifting point. Blind spots caused by multiple obstacles further exacerbate the risk of positioning errors. Temporally, operator fatigue and distraction during extended periods of operation reduce the real-time performance of target tracking, making it difficult to adapt to the fast-paced demands of high-frequency hoisting tasks.

[0004] Therefore, there is an urgent need for a processing method to improve the accuracy of identifying objects hoisted by tower cranes. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, device, and storage medium for identifying the type of objects hoisted by tower cranes, in order to solve the technical problem of low accuracy in identifying objects hoisted by tower cranes in related technologies.

[0006] In a first aspect, embodiments of this application provide a method for identifying the type of load lifted by a tower crane, including: The effective lifting operations of the tower crane are determined based on the tower crane's status information, and the operation video of the tower crane under the effective lifting operations is obtained; The operation video is preprocessed and the suspended object is identified to obtain a set of candidate regions corresponding to each video frame in the operation video; Each candidate region in the candidate region set is detected and identified, and each candidate region is analyzed based on the processing results to obtain the instance features of the instance contained in the corresponding video frame. The instance features are analyzed and processed to determine the target object and its location information in the instance, and the type of the target object is determined based on the location information in the effective lifting operations.

[0007] Secondly, embodiments of this application provide a tower crane lifting load type identification device, including: The video acquisition module is used to determine the effective lifting operations of the tower crane based on the tower crane's status information, and to acquire the operation video of the tower crane under the effective lifting operations; The video processing module is used to preprocess the operation video and identify the suspended object to obtain a set of candidate regions corresponding to each video frame in the operation video; The instance analysis module is used to detect and identify each candidate region in the candidate region set, and perform instance analysis on each candidate region based on the processing results to obtain the instance features of the instance contained in the corresponding video frame. The type determination module is used to analyze and process the instance features, determine the target object and its position information in the instance, and determine the type of the target object under the effective lifting operations based on the position information.

[0008] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the tower crane lifting type identification method described in any of the above claims.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the tower crane lifting type identification method described in any of the preceding claims.

[0010] This application provides a method, apparatus, device, and storage medium for identifying the type of objects lifted by a tower crane. When identifying and analyzing objects lifted by a tower crane during operation, the method determines the effective lifting cycles of the tower crane based on its status information and acquires operational videos of the tower crane under the effective lifting cycles. The operational videos are preprocessed and the objects are identified to obtain a set of candidate regions corresponding to each video frame. Each candidate region in the set is detected and identified, and based on the processing results, each candidate region undergoes instance analysis to obtain instance features of the instances contained in the corresponding video frames. The instance features are analyzed to determine the target object and its location information within the instance, and the type of object lifted by the target object under the effective lifting cycles is determined based on the location information. By combining the tower crane's status information, the system automatically analyzes and processes the loads lifted by the tower crane. It focuses on analyzing data from effective lifting operations, thereby improving the efficiency of load analysis and identification. Finally, through a multi-level processing pipeline of "candidate area screening → detection and identification → instance analysis → target determination," it can gradually filter out background interference and accurately locate the target load and determine its type, thus improving the accuracy and efficiency of load identification during tower crane operations. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating one step of the tower crane lifting type identification method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating one step of obtaining a data processing task according to an embodiment of this application; Figure 3 This is a flowchart illustrating the steps for commonality analysis and processing provided in an embodiment of this application; Figure 4 This is a schematic diagram of an architecture of a data processing system provided in an embodiment of this application; Figure 5 This is a schematic diagram of a tower crane lifting type identification device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 7 This is another structural schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0013] It should be understood that the steps described in the method embodiments disclosed in this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0015] In related technologies, traditional hoisting operations mainly rely on the operator's visual observation and experience-based decision-making, which has several inherent limitations. From a spatial perspective, the geometric constraints between the fixed viewpoint of the tower crane's cab and the dynamic trajectory of the load make it difficult for manual visual inspection to accurately determine the spatial pose relationship between the hook and the target lifting point. Blind spots caused by multiple obstacles further exacerbate the risk of positioning errors. From a temporal perspective, operator fatigue and distraction during long working hours reduce the real-time performance of target tracking, making it difficult to adapt to the fast-paced demands of high-frequency hoisting tasks.

[0016] To address the technical problems existing in related technologies, this application provides a method for identifying the type of load lifted by a tower crane. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart illustrating one step of the tower crane lifting type identification method provided in this application embodiment, which includes steps 101 to 104.

[0017] Step 101: Determine the effective lifting count of the tower crane based on the tower crane's status information, and obtain the operation video of the tower crane under the effective lifting count.

[0018] In one embodiment, when identifying the load being lifted by the tower crane, it is first necessary to determine whether the current state of the tower crane is a valid lifting operation, meaning that the tower crane is currently in operation. Therefore, when performing type identification processing on the load being lifted by the tower crane, the tower crane's status information is obtained. Upon determining that the current lifting operation is a valid lifting operation based on the tower crane's status information, the operation video of the tower crane during that valid lifting operation is acquired. A valid lifting operation is the process in which the tower crane is performing a lifting operation and the hook is loaded.

[0019] For example, the tower crane's status information includes time-series monitoring data of the crane's lifting weight and lifting height. The lifting height data refers to the height of the crane's hook above the ground, and the lifting weight data refers to the weight of the load being lifted during the lifting process. By analyzing and processing the tower crane's status information, it is determined whether the tower crane is currently in a valid lifting phase.

[0020] Furthermore, when analyzing and processing effective lifting operations, operational videos of the tower crane under each effective lifting operation will be obtained. Specifically, refer to... Figure 2 , Figure 2 This is a flowchart illustrating the steps for obtaining a valid lifting operation video according to an embodiment of this application, wherein the steps include steps 201 to 204.

[0021] Step 201: Obtain the status information of the tower crane, and generate a lifting height time series and a lifting weight time series based on the status information. The status information includes the lifting height and lifting weight of the tower crane, where the lifting height is the height of the tower crane's hook. Step 202: Compare the hoisting height time series with the corresponding hoisting height threshold, and obtain the first time period based on the comparison results; Step 203: Compare the lifting weight time series with the corresponding lifting weight threshold, and obtain the second time period based on the comparison results; Step 204: Align the first time period and the second time period to obtain the effective lifting counts of the tower crane, and acquire the operation video of the tower crane within the time period of the effective lifting counts.

[0022] Specifically, when performing tower crane lifting type identification, the tower crane's status information, including lifting height and lifting weight, is first acquired. This status information is used to generate lifting height and lifting weight time series. The lifting height time series is then compared with a preset lifting height threshold to determine the time period where the hook experiences a significant height change as the first time period. Simultaneously, the lifting weight time series is compared with a preset lifting weight threshold to determine the time period where the lifting weight exceeds the threshold and persists for a certain duration as the second time period. The intersection of the first and second time periods is then performed to obtain the overlapping time interval, which is the effective lifting time period for the tower crane. Based on this, operation videos recorded by the tower crane monitoring equipment within this time period are retrieved as input data for subsequent lifting type identification, ensuring that the acquired video content corresponds to a genuine and valid lifting operation process.

[0023] In practical applications, determining whether a lifting operation is valid involves analyzing the lifting height and weight of the tower crane hook. Specifically, when judging independent lifting operations based on lifting weight and height data, the original lifting weight and height data need to be smoothed. Due to factors such as mechanical vibration (frequency range 5-20Hz), hydraulic system pressure pulsation (amplitude ±3%FS), and electromagnetic interference (intensity ≤30V / m) in the operating environment, the original lifting height and weight data generally contain high-frequency noise (signal-to-noise ratio ≤40dB), pulse-type outliers (deviation from the mean ≥3σ), and signal baseline drift (drift rate 0.05% / min). In this case, a discrete Kalman filter algorithm can be used to construct a state-space model, establishing state equations and observation equations that include the dynamic characteristics of the lifting system. Specifically, the hook height... and lifting heavy loads It is modeled as a linear dynamic system with Gaussian white noise.

[0024] Equations of state: ; Observation equation: ; Wherein, the state vector Includes position, velocity, and acceleration components, process noise. ~ and observation noise ~ The covariance matrix was calibrated using Allan variance analysis. Specifically, considering the non-stationary nature of hoisting operations, a time-varying parameter adaptive mechanism was designed. When a sudden change in lifting weight (Δm / Δt ≥ 5% rated load / s) is detected, the predicted covariance matrix is ​​automatically adjusted. Update weight coefficients In the prediction phase, the state and error covariance at the current moment are predicted based on the state equation. In the update phase, the Kalman gain is calculated based on the observed values, and the state estimate and error covariance are updated. The prediction and update steps are repeated, and the entire hoisting height and load time series data are filtered and smoothed to eliminate abnormal data caused by factors such as swaying and sensor errors during the hoisting process, generating hoisting height and load time series curves.

[0025] Next, the crane is judged to be in the lifting state by the lifting load time series curve. If the lifting load on the lifting load time series curve is high... The lifting weight is greater than the lifting weight threshold when the hook is unloaded. If the crane is in the hoisting state, it is in the lifting state; otherwise, it is in the non-hoisting state. Further, when the crane is in the hoisting state, the smoothed lifting height time series curve is used to determine whether the crane's hoisting is a valid operation. If the lifting height on the lifting height time series curve is... Greater than the lifting height threshold when the tower crane begins hoisting If the lifting operation is successful, the tower crane's lifting status is considered a valid lifting operation; otherwise, the tower crane's lifting status is considered an invalid lifting operation.

[0026] Alternatively, analysis can be performed based on the first and second time periods to determine whether a lifting operation is valid. This is done by analyzing whether there is a time overlap between the first and second time periods; if so, it is considered a valid lifting operation; otherwise, it is considered invalid. Next, when acquiring the operation video within the time period of the valid lifting operation, the first and second time periods are time-aligned to obtain the overlapping time period, which is the time period in which the valid lifting operation occurred. Then, the pre-collected operation video within that time period is acquired. Specifically, this can include: time-aligning the first and second time periods to determine if they overlap; if an overlapping time period is determined, the lifting operations included in the overlapping time period are identified as valid lifting operations of the tower crane, and the operation video of the tower crane within the overlapping time period is acquired.

[0027] Step 102: Preprocess the operation video and identify the suspended object to obtain a set of candidate regions corresponding to each video frame in the operation video.

[0028] In one embodiment, after obtaining the operation video of the tower crane during effective lifting operations, preprocessing and load identification processing are performed on the operation video to obtain a set of candidate regions corresponding to each video frame in the operation video. During preprocessing, video preprocessing algorithms can be used to enhance the image of the operation video, improving the robustness of the recognition algorithm under complex conditions such as low-light environments and camera shake. Preprocessing includes digital image stabilization, low-light enhancement, and overexposure compression of the operation video. Before performing digital image stabilization, low-light enhancement, and overexposure compression, the operation video can also be assessed for image shake, low light, and overexposure.

[0029] For example, when preprocessing the acquired operation video, the process may include determining whether there is video image jitter. If jitter is present, a digital anti-shake algorithm is used to eliminate the jitter in the hoisting video, thereby improving the detection capability of foreground moving targets. It may also determine whether the operation video was captured in a low-light environment. If it is determined to have been captured in a low-light environment, nonlinear brightness restoration can be performed to improve the brightness and contrast of the hoisting video image. If it is determined to have been captured in an overexposed environment, dynamic range compression can be applied to the overexposed frames to reduce the brightness of the hoisting video image.

[0030] When preprocessing the acquired video footage, jitter detection can also be performed. The motion vector distribution of feature points between adjacent frames is calculated using optical flow analysis, and the horizontal baseline fluctuation in the image is detected by combining Hough transform. If the baseline offset exceeds 5 pixels within 10 consecutive frames and the motion vectors show irregular divergence, it is determined that there is significant jitter. This triggers an electronic image stabilization algorithm based on feature point matching, which compensates for motion in subsequent frames by constructing an affine transformation matrix, reducing the inter-frame displacement variance, and thus improving the detection accuracy of the suspended object recognition method in swaying acquisition scenarios.

[0031] When preprocessing the acquired work videos, environmental illumination analysis can also be performed. By analyzing the brightness distribution of the video histogram (when the Y channel value of 80% of the pixels is below 40), it is determined to be a low-light environment. At this time, the nonlinear brightness recovery module based on Retinex theory is activated. First, guided filtering is used to separate the illumination component and the reflection component. Adaptive Gamma correction (γ=2.4) is performed on the V channel in the HSV color space. At the same time, the contrast-limited histogram equalization of the H and S channels is performed through the CLAHE algorithm.

[0032] When preprocessing the acquired work video, pixel brightness values ​​can also be analyzed. If it is detected that more than 15% of the pixels in the work video have a brightness value of 255 (RGB three-channel overexposure), a multi-scale dynamic range compression process can be started. Under the Laplacian pyramid decomposition framework, S-curves are used to enhance the local contrast of the high-frequency layer, and adaptive brightness attenuation based on image entropy is applied to the low-frequency layer.

[0033] Furthermore, after completing the preprocessing of the video, when obtaining the candidate region set corresponding to each video frame, specifically, when obtaining the candidate region set, it is possible to... Figure 3 , Figure 3 This is a flowchart illustrating a step for obtaining a set of candidate regions according to an embodiment of this application, wherein the step includes steps 301 to 303.

[0034] Step 301: Perform foreground detection on the operation video to obtain the motion image of the tower crane hook; Step 302: Extract features from the motion image to obtain temporal motion features and spatial texture features, and perform feature fusion processing on the temporal motion features and spatial texture features to obtain the motion saliency image corresponding to the motion image. Step 303: Perform multi-scale sliding window cropping on the motion saliency image to obtain a set of candidate regions corresponding to the motion image.

[0035] Foreground detection refers to the process of separating moving objects from a video sequence. This can be achieved using background subtraction or optical flow methods, with the aim of highlighting the dynamic trajectory of the hook. Temporal motion features can be understood as features describing the motion patterns of objects and can be extracted using optical flow features. Spatial texture features can be understood as features describing the surface texture of objects and can be extracted using SIFT or HOG features. Feature fusion processing refers to integrating features from different dimensions, which can be achieved using weighted averaging or feature concatenation. Multi-scale sliding window cropping refers to sliding a window across the image at different scales, which can be achieved using image pyramids or multi-resolution grids. In other words, during processing, a dynamic background model is constructed from the preprocessed work video. Moving targets are initially separated using background subtraction. A spatiotemporal joint attention mechanism is introduced to calculate a motion saliency map, suppressing static interference such as scaffolding and material stacking. Multi-scale sliding window cropping is then performed on the salient regions of the motion saliency map to generate a corresponding set of candidate regions.

[0036] Specifically, the process first involves foreground detection of the operation video to obtain motion images of the hook, thereby effectively separating dynamic targets from static background interference. Then, feature extraction is performed on the motion images, simultaneously acquiring temporal motion features and spatial texture features. Feature fusion is then used to generate motion saliency images, enhancing the salient areas of the suspended object. Finally, the motion saliency images are cropped using a multi-scale sliding window, and the candidate region range is adaptively determined based on the saliency intensity distribution.

[0037] For example, foreground detection uses a Gaussian mixture model for background modeling, and then performs feature extraction. Temporal motion features can be obtained through optical flow calculation, while spatial texture features can be extracted through local binary mode descriptors. Then, during feature fusion, feature vector concatenation can be used to obtain a motion saliency image. Finally, multi-scale sliding window cropping is performed on the motion saliency image to obtain a set of candidate regions corresponding to each video frame.

[0038] In the actual processing, when processing the preprocessed work video to obtain a candidate region set, a dynamic background model is constructed on the work video, and the moving targets are initially separated using the background subtraction method. A spatiotemporal joint attention mechanism is introduced to calculate the motion saliency map to suppress static interference such as scaffolding and material stacking. Furthermore, for the generated candidate region set, the salient regions of the motion saliency map are subjected to multi-scale sliding window clipping.

[0039] In practice, for each pixel in each video frame of the assignment video, a Gaussian mixture model is built. This model consists of multiple Gaussian distributions to describe the color (or grayscale value) distribution of the pixel. Each Gaussian distribution has its own parameters such as weights, mean, and variance. ; in, It is the mean of a Gaussian distribution. It's the variance, the color value of a pixel. Represented by a weighted sum of multiple Gaussian distributions: ; in, It is the first The weights are distributed in a Gaussian manner, satisfying... , It represents the number of Gaussian distributions.

[0040] For each pixel, when a new video frame arrives, the model parameters are updated based on the matching of the pixel value with the existing Gaussian distribution. If the pixel value matches a Gaussian distribution, the mean and variance of that distribution are updated, specifically in the following manner:

[0041]

[0042] in, It is the learning rate, which controls the speed at which the model is updated.

[0043] A Gaussian mixture model is used to label the foreground mask. If a pixel value matches all Gaussian distributions below a set threshold, the pixel is labeled as a foreground pixel; otherwise, it is labeled as a background pixel. Foreground pixels within the same video frame constitute the foreground mask.

[0044] Next, the foreground pixel regions marked in the previous step are further processed using opening and closing morphological operations to remove some isolated noise points or small false detection areas, fill in small holes in the foreground, and enhance the integrity and connectivity of the foreground object.

[0045] To further extract the motion features of the suspended object from complex backgrounds while suppressing static interferences such as scaffolding and material stacking, a spatiotemporal joint attention mechanism is used to fuse temporal motion features and spatial texture features to calculate the motion saliency map of the hoisting video image. Since the optical flow field directly reflects pixel-level motion information, it can capture the minute displacements of the suspended object. At the same time, the sparse optical flow method has high computational efficiency and is suitable for real-time processing.

[0046] First, optical flow field calculations are performed on a sequence of three consecutive images. The sparse optical flow method is used to calculate the motion vectors between adjacent frames. Calculate optical flow amplitude Subsequently, Normalization to In the time dimension, a 3D convolution operation is performed on three consecutive frames of images, according to... Extract spatiotemporal features. In the spatial dimension, use the Sobel operator to calculate the edge response of the current frame. Highlighting the outline of the suspended object. Spatiotemporal feature fusion enhances the saliency of moving targets while suppressing interference from static backgrounds (such as scaffolding). Sobel edge detection helps distinguish the suspended object from rigid structures (such as steel reinforcement piles), avoiding false detections. Combining spatial and temporal features, a spatiotemporal joint attention weight is constructed:

[0047] The spatiotemporal joint attention weights constructed above are motion saliency maps, i.e., motion saliency images. When the attention weights are below a threshold... At that time, the location was subject to static background interference. Among them, the threshold... It can be set to 0.75. Perform multi-scale sliding window clipping on the salient regions (not less than 0.75) of the motion saliency map to generate a set of candidate regions for the suspended object.

[0048] Step 103: Detect and identify each candidate region in the candidate region set, and perform instance analysis on each candidate region based on the processing results to obtain the instance features of the instance contained in the corresponding video frame.

[0049] In one embodiment, after obtaining the candidate region set corresponding to each video frame, instance analysis processing is performed on each candidate region in the candidate region set to obtain the instance features of all instances contained in the corresponding video frame. Here, an instance refers to an independent object unit in the video frame that can be semantically segmented, such as a hook, a suspended object, or a robotic arm, etc. The suspended object can be a material such as steel bars, steel pipes, steel strips, channel steel, and precast concrete components.

[0050] When processing the candidate region set, instance features of each video frame are obtained by identifying and analyzing each candidate region. The instance features can be the centroid coordinates of each instance. The specific location of the instance in the corresponding video frame is determined by the instance features. Then, the type of the identified instance is combined with spatial analysis to determine the type of object being lifted by the tower crane.

[0051] Furthermore, when processing to obtain instance features, the process may include: using a pre-set construction site building material recognition model to perform hook detection and building material classification and recognition on each candidate region in the candidate region set, to obtain the candidate hanging object region corresponding to the candidate region containing the hook; performing instance segmentation and centroid calculation on the candidate hanging object region based on the instance segmentation model, to obtain the instances contained in the candidate hanging object region and the centroid coordinates of each instance.

[0052] Specifically, during processing, a pre-trained, pre-defined construction site building material recognition model is used for identification. By detecting and recognizing hooks, the presence of hooks in candidate regions is determined, and candidate regions containing hooks are marked as candidate hanging object regions. Subsequently, an instance segmentation model is used to perform pixel-level segmentation of the candidate hanging object regions, extracting the mask information of each hanging object instance and calculating its centroid coordinates and bounding box. Combined with temporal context information, the motion trajectories of the same instances in adjacent frames are associated to achieve tracking and feature consistency matching of the hanging object in continuous video frames, thereby accurately obtaining the type of hanging object and its spatial position changes.

[0053] For example, the aforementioned pre-built construction site material recognition model is constructed based on a deep neural network and can be trained using an improved lightweight YOLO model on a diverse and high-quality dataset of labeled objects in large-scale construction site scenarios. During training, a dataset of labeled objects is pre-acquired, which includes the following features: Object type coverage: encompassing 10 major categories of construction materials and tools, including hooks, reinforcing bars, steel pipes, wheelbarrows, channel steel, precast concrete components, timber, and wooden formwork, with no fewer than 5000 samples per category. Regarding lighting conditions, it includes four scenarios: strong light (noon, illuminance > 80000 lux), cloudy (illuminance 20000-50000 lux), dusk (illuminance 5000-10000 lux), and nighttime construction site lighting (illuminance < 1000 lux). Regarding weather conditions, the samples included sunny days (visibility > 5 km), foggy days (visibility 0.5-1 km), and light rain (rainfall < 2.5 mm / h), with a ratio of 5:3:2. Regarding the movement of the suspended object, the samples included four types of dynamic samples: stationary hoisting (speed 0 m / s), lifting and lowering (speed 0.2–1.5 m / s), horizontal movement (speed 0.5–2 m / s), and rotational oscillation (angular velocity 5–15° / s). Regarding spatial distribution characteristics, the samples included three compositions: foreground (suspended object occupies > 50% of the image area), midground (20%–50%), and background (5%–20%), with a ratio of 4:3:3.

[0054] The improved YOLO model structure uses the YOLO v8 model as its backbone. A dynamic small object detection branch is designed on top of this backbone. A detection head consisting of three sets of dilated convolutions with dilation rates [2, 4, 6] is embedded in the third layer of the FPN, specifically designed to identify slender objects such as steel bars and pipes with an aspect ratio > 5:1. An illumination invariance module is added to the backbone model, and a Retinex decomposition layer is inserted before the classification head to decompose the input image into a reflection component R and an illumination component L. The R component is used as the classification feature input. Furthermore, the loss function is optimized by using Focal Loss (α=0.8, γ=2) to balance positive and negative samples, and a weight coefficient of 1.5 is set for slender objects.

[0055] During training, a combination of four enhancement strategies was employed: geometric transformation, color perturbation, noise injection, and motion blur. Geometric transformations included random rotation, horizontal flipping, and random cropping. Random rotation angles were within ±15°, the probability of horizontal flipping was set to 0.5, and the proportion of random cropping ranged from 0.8 to 1.2. Color perturbation included brightness adjustment, saturation adjustment, and contrast adjustment. Brightness adjustment ranged from ±30%, saturation adjustment from ±40%, and contrast adjustment from ±25%. Noise injection included salt-and-pepper noise and Gaussian noise. The density of salt-and-pepper noise was set to 0.01, and the variance σ of Gaussian noise was set to 0.01. Motion blur was achieved by randomly generating motion directions within the 0–360° range, with the blur kernel size set to 5–15 pixels.

[0056] After training, a pre-built instance segmentation deep neural network is used for instance segmentation during subsequent use. The pre-built instance segmentation deep neural network is used to segment the detected hook and candidate hanging object regions, and the centroid coordinates of each instance are calculated. The instance segmentation network is implemented using a modified Mask R-CNN. The implementation steps include parallel computation of LBP features (radius 2 pixels, 8 neighborhoods) and HOG features (cell size 8×8) after the conv3_x layer of ResNet-50, and concatenation with the P3 layer features output by FPN for feature fusion. K=3 dynamic convolutional kernels are used to dynamically generate mask weights for the category based on the ROI features to predict the dynamic mask. Monte Carlo Dropout is used, performing 50 forward propagations during the testing phase, and the probability mean of the output mask is taken. The centroid coordinates of the segmented instance masks are calculated.

[0057] Step 104: Analyze and process the instance features, determine the target object and its location information in the instance, and determine the type of the target object to be lifted in the effective lifting cycles based on the location information.

[0058] In one embodiment, after the calculation and processing of the centroid coordinates of the instance are completed, the target object and its position information are determined in the determined instance, and then the type of object to be lifted by the tower crane in the current effective lifting session is determined.

[0059] For example, when further analyzing and processing the obtained instance, it is necessary to first determine that the instance is the object lifted by the tower crane in that effective lifting cycle, and then determine the position of the object and match it with the type obtained by category identification in order to determine the type of the object being lifted in that effective lifting cycle.

[0060] Furthermore, when processing, one can refer to... Figure 4 , Figure 4 This is a flowchart illustrating a step for obtaining the type of suspended object according to an embodiment of this application, wherein the step includes steps 401 to 403.

[0061] Step 401: The candidate suspended object region is filtered by trajectory continuity to obtain intermediate instances by filtering out instances of discontinuous motion. Step 402: Perform motion similarity matching between the intermediate instance and the motion of the hook in the corresponding video frame to determine the target region to which the intermediate instance belongs in the candidate hanging object region; Step 403: Spatial matching is performed between the target region and the recognition result obtained from classification and recognition, and the type of the target object contained in the video frame to which the determined target region belongs is taken as the type of the target object under the effective lifting count.

[0062] Trajectory continuity refers to the consistency of the motion trajectory of the analyzed instance in the time series. It can be achieved by using a trajectory prediction model based on Kalman filtering or an optical flow motion continuity detection algorithm. By identifying abrupt changes or discontinuities in the motion trajectory, it effectively filters out non-continuous instances such as background interference objects, ensuring that intermediate instances retain the inherent stable motion characteristics of the suspended object. Motion similarity matching can be understood as quantifying the dynamic correlation between intermediate instances and the motion trajectory of the hook. It can be achieved by using Euclidean distance measurement or dynamic time warping algorithm to accurately match the target area that moves synchronously with the hook. Spatial matching refers to calibrating the consistency between the target area and the classification recognition result in spatial geometry. It can be achieved by using bounding box overlap rate calculation or centroid coordinate matching method to ensure the accuracy of the suspended object type determination.

[0063] Specifically, firstly, based on the physical characteristic that the suspended object must maintain a stable trajectory during hoisting, trajectory continuity analysis filters out interference instances with abrupt or intermittent trajectories, retaining intermediate instances with continuous motion characteristics. Next, leveraging the strong correlation between the stability of the hook's trajectory and the object's motion, motion similarity matching calculates the similarity between the intermediate instances and the hook's motion, accurately determining the target area moving synchronously with the hook. Finally, based on the principle of spatial consistency, the target area is geometrically aligned with the classification and recognition results, directly using the real-time recognition result of the video frame to which the target area belongs as the basis for determining the object type. This phased processing mechanism forms a complete technical chain from interference filtering and precise positioning to reliable recognition. Each step is progressive and mutually supportive, effectively solving the recognition deviation problem caused by interference instances and positioning offsets in dynamic construction environments.

[0064] In fact, when performing verification based on trajectory continuity, continuous... The displacement of the candidate suspended object's core is compared with a displacement threshold. The condition for determining that the candidate suspended object is a static disturbance is as follows: ; in, This represents the target displacement in adjacent frames. , These are the threshold values ​​for the mean and variance, respectively. If the mean displacement of the suspended object's center of gravity is less than the displacement threshold and the variance of the suspended object's center of gravity is less than the displacement variance, then the candidate suspended object is determined to be a static disturbance.

[0065] In motion similarity-based matching, the similarity of the center of mass of the identified tower crane hook and the candidate object's center of mass is measured according to the shape of the trajectory. The time axes of the hook and the candidate object's trajectories are aligned, and the minimum path distance is calculated. ; in, , These are the trajectory coordinates of the hook and the candidate object in the image, respectively. To normalize the path, non-linear alignment of the hook and candidate object trajectories along the time axis is allowed. The minimum path distance between the motion trajectory of each candidate segmentation instance and the hook's motion trajectory is calculated using a trajectory shape similarity metric. The candidate segmentation instance with the smallest path distance is selected as the final object region.

[0066] Based on the above processing method, the risk of false detection caused by dynamic interference at the construction site can be eliminated more effectively, the movement relationship between the hoisted object and the hook can be accurately correlated, and the type of hoisted object can be reliably determined, thereby improving the robustness and accuracy of tower crane hoisted object identification and ensuring the safety and efficiency of the construction process.

[0067] In addition, when determining the type of object to be lifted, the type of object lifted by the tower crane hook in the current independent valid lifting operation is determined. Since there are several video frames in the operation video corresponding to the valid lifting operation, after determining the type of object lifted by the tower crane in each video frame based on the above processing method, a further overall judgment will be made on the operation video. Specifically, the object types associated with all video frames can be judged based on a voting mechanism.

[0068] Therefore, when determining the type of the target object to be suspended, the process also includes: spatially matching the target area with the recognition results obtained from classification to obtain the type of the object to be suspended in each video frame; and making a judgment based on the majority voting mechanism according to the type of the object to be suspended in each video frame, and determining the type of the object to be suspended with the most votes as the type of the target object to be suspended in the effective number of suspensions.

[0069] In other words, during the judgment process, spatial matching of the target area with the classification and recognition results ensures that the position information of the target suspended object in each video frame strictly corresponds to the recognition result, thereby obtaining continuous single-frame suspended object type data in the time series. Then, based on this, the suspended object types of all video frames within the valid lifting sessions are statistically aggregated using a majority voting mechanism. Utilizing the physical characteristic that the type of the suspended object remains unchanged throughout the entire lifting process, abnormal recognition results from interfering frames are automatically identified and eliminated. Finally, the suspended object type is determined by vote counting. This process effectively integrates spatial position matching and temporal statistics, forming a dual suppression mechanism against dynamic interference: the spatial matching stage eliminates mismatches caused by positional deviations, and the temporal voting stage overcomes recognition fluctuations caused by instantaneous environmental interference. The two work together to ensure a stable and consistent suspended object type determination result within the valid lifting session time period.

[0070] In summary, the above embodiments provide a method for identifying the type of objects lifted by a tower crane. When identifying and analyzing the objects lifted by the tower crane during operation, the effective lifting cycles of the tower crane are determined based on the crane's status information, and operation videos of the tower crane under the effective lifting cycles are obtained. The operation videos are preprocessed and the objects are identified to obtain a set of candidate regions corresponding to each video frame in the operation video. Each candidate region in the candidate region set is detected and identified, and based on the processing results, instance analysis is performed on each candidate region to obtain instance features of the instances contained in the corresponding video frames. The instance features are analyzed to determine the target object and its location information within the instance, and the type of object lifted by the target object under the effective lifting cycles is determined based on the location information. By combining the tower crane's status information, the system automatically analyzes and processes the loads lifted by the tower crane. It focuses on analyzing data from effective lifting operations, thereby improving the efficiency of load analysis and identification. Finally, through a multi-level processing pipeline of "candidate area screening → detection and identification → instance analysis → target determination," it can gradually filter out background interference and accurately locate the target load and determine its type, thus improving the accuracy and efficiency of load identification during tower crane operations.

[0071] Based on the method described in the above embodiments, this embodiment will further describe it from the perspective of the tower crane hoisting type identification device. The tower crane hoisting type identification device can be implemented as an independent entity or integrated into an electronic device, such as a terminal, which may include a mobile phone, tablet computer, etc.

[0072] Please see Figure 5 , Figure 5 This is a schematic diagram of a tower crane load type identification device provided in an embodiment of this application, such as... Figure 5 As shown in the embodiment of this application, the tower crane lifting load type identification device 500 includes: The video acquisition module 501 is used to determine the effective lifting times of the tower crane based on the tower crane's status information, and to acquire the operation video of the tower crane under the effective lifting times; The video processing module 502 is used to preprocess the operation video and identify the suspended object to obtain a set of candidate regions corresponding to each video frame in the operation video; The instance analysis module 503 is used to detect and identify each candidate region in the candidate region set, and perform instance analysis on each candidate region based on the processing results to obtain the instance features of the instance contained in the corresponding video frame. The type determination module 504 is used to analyze and process the instance features, determine the target object and its location information in the instance, and determine the type of the target object under the effective lifting count based on the location information.

[0073] In one embodiment, the video acquisition module 501 is further configured to: The status information of the tower crane is obtained, and the lifting height time series and lifting weight time series are generated based on the status information. The status information includes the lifting height and lifting weight of the tower crane, and the lifting height is the height of the tower crane hook. The hoisting height time series is compared with the corresponding hoisting height threshold, and the first time period is obtained based on the comparison results. The time series of suspended loads is compared with the corresponding suspended load thresholds, and the second time period is obtained based on the comparison results. The first and second time periods are time-aligned to obtain the effective lifting counts of the tower crane, and the operation video of the tower crane within the time period of the effective lifting counts is obtained.

[0074] In one embodiment, the video acquisition module 501 is further configured to: The first time period and the second time period are aligned to determine whether the first time period and the second time period overlap. When overlapping time periods are identified, the number of lifting operations included in the overlapping time periods are determined as valid lifting operations of the tower crane, and the operation video of the tower crane within the overlapping time periods is obtained.

[0075] In one embodiment, the video processing module 502 is further configured to: Foreground detection is performed on the operation video to obtain the motion image of the tower crane's hook; Feature extraction is performed on the motion image to obtain the temporal motion features and spatial texture features. The temporal motion features and spatial texture features are then fused to obtain the motion saliency image corresponding to the motion image. Multi-scale sliding window cropping is performed on the motion saliency image to obtain a set of candidate regions corresponding to the motion image.

[0076] In one embodiment, the instance analysis module 503 is further configured to: A pre-set construction site building material recognition model is used to perform hook detection and building material classification and recognition on each candidate region in the candidate region set, so as to obtain the candidate hanging object region corresponding to the candidate region containing hooks in the candidate region set; The candidate suspended object region is segmented and centroids are calculated based on the instance segmentation model to obtain the instances contained in the candidate suspended object region and the centroid coordinates of each instance.

[0077] In one embodiment, the type determination module 504 is further configured to: The candidate suspended object region is filtered by using trajectory continuity to filter out instances of discontinuous motion and obtain intermediate instances. The intermediate instance is matched with the motion of the hook in its corresponding video frame to determine the target region to which the intermediate instance belongs in the candidate hanging object region; Spatial matching is performed between the target region and the recognition results obtained from classification, and the type of the target object contained in the video frame to which the target region belongs is taken as the type of the target object under effective lifting operations.

[0078] In one embodiment, the type determination module 504 is further configured to: Spatial matching is performed between the target area and the recognition results obtained from classification to obtain the type of hanging object corresponding to each video frame; Based on the type of object to be suspended for each video frame, a majority voting mechanism is used to determine the type of object to be suspended for the target object in valid suspension attempts.

[0079] Additionally, please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a mobile terminal such as a smartphone or tablet computer. Figure 6 As shown, the electronic device 600 includes a processor 601 and a memory 602. The processor 601 and the memory 602 are electrically connected.

[0080] The processor 601 is the control center of the electronic device 600. It connects various parts of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 602 and calling the data stored in the memory 602, it performs various functions of the electronic device 600 and processes data, thereby monitoring the electronic device 600 as a whole.

[0081] In this embodiment, the processor 601 in the electronic device 600 will load the instructions corresponding to the processes of one or more applications into the memory 602 according to the following steps, and the processor 601 will run the applications stored in the memory 602 to realize any step in the tower crane lifting object type identification method provided in the above embodiment.

[0082] The electronic device 600 can implement the steps of any embodiment of the tower crane lifting type identification method provided in the embodiments of this application. Therefore, it can achieve the beneficial effects that any tower crane lifting type identification method provided in the embodiments of this application can achieve. For details, please refer to the previous embodiments, which will not be repeated here.

[0083] Please see Figure 7 , Figure 7 This is another structural schematic diagram of the electronic device provided in the embodiments of this application, such as... Figure 7 As shown, Figure 7 A specific structural block diagram of an electronic device provided in an embodiment of this application is shown. This electronic device can be used to implement the tower crane lifting object type identification method provided in the above embodiments. The electronic device 700 can be a mobile terminal such as a smartphone or laptop computer.

[0084] RF circuit 710 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 710 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity modules (SIM cards), memory, etc. RF circuit 710 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.

[0085] The memory 720 can be used to store software programs and modules, such as the program instructions / modules corresponding to the tower crane lifting type identification method in the above embodiment. The processor 780 executes various functional applications and the tower crane lifting type identification method by running the software programs and modules stored in the memory 720.

[0086] Memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, memory 720 may further include memory remotely located relative to processor 780, which can be connected to electronic device 700 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0087] The input unit 730 can be used to receive uploaded digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 780, and can receive and execute commands sent by the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0088] Display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic device 700. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 740 may include display panel 741, optionally configured as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms. Further, touch-sensitive surface 731 may cover display panel 741. When touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to processor 780 to determine the type of touch event. Subsequently, processor 780 provides corresponding visual output on display panel 741 according to the type of touch event. Although in the figures, touch-sensitive surface 731 and display panel 741 are implemented as two separate components to achieve input and output functions, in some embodiments, touch-sensitive surface 731 and display panel 741 can be integrated to achieve input and output functions.

[0089] The electronic device 700 may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can generate an interruption when the flip is closed or shut down. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the electronic device 700, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0090] Audio circuitry 760, speaker 761, and microphone 762 provide an audio interface between the user and electronic device 700. Audio circuitry 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. Conversely, microphone 762 converts collected sound signals into electrical signals, which are then received by audio circuitry 760, converted back into audio data, and processed by processor 780. The audio data is then transmitted via RF circuitry 710 to, for example, another terminal, or output to memory 720 for further processing. Audio circuitry 760 may also include an earphone jack to facilitate communication between peripheral headphones and electronic device 700.

[0091] Electronic device 700, through transmission module 770 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 770 is shown in the figure, it is understood that it is not an essential component of electronic device 700 and can be omitted as needed without changing the essence of the invention.

[0092] The processor 780 is the control center of the electronic device 700. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 700 by running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, thereby providing overall monitoring of the electronic device. Optionally, the processor 780 may include one or more processing cores; in some embodiments, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.

[0093] The electronic device 700 also includes a power supply 790 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to the processor 780 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 790 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0094] Although not shown, the electronic device 700 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, one or more of which are stored in the memory and configured to be executed by one or more processors to implement any step of the tower crane lifting object type identification method provided in the above embodiment.

[0095] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.

[0096] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, this application provides a storage medium storing multiple instructions that, when executed by a processor, can implement any step in the tower crane lifting object type identification method provided in the above embodiments.

[0097] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0098] Since the instructions stored in the storage medium can execute the steps in any embodiment of the tower crane lifting type identification method provided in this application, the beneficial effects that any tower crane lifting type identification method provided in this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0099] The foregoing has provided a detailed description of a method, apparatus, electronic device, and storage medium for identifying the type of load lifted by a tower crane, as provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application. Moreover, those skilled in the art can make several improvements and modifications without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.

Claims

1. A method for identifying the type of load lifted by a tower crane, characterized in that, include: The effective lifting operations of the tower crane are determined based on the tower crane's status information, and the operation video of the tower crane under the effective lifting operations is obtained; The operation video is preprocessed and the suspended object is identified to obtain a set of candidate regions corresponding to each video frame in the operation video; Each candidate region in the candidate region set is detected and identified, and each candidate region is analyzed based on the processing results to obtain the instance features of the instance contained in the corresponding video frame. The instance features are analyzed and processed to determine the target object and its location information in the instance, and the type of the target object is determined based on the location information in the effective lifting operations.

2. The method for identifying the type of load lifted by a tower crane as described in claim 1, characterized in that, The step of determining the effective lifting operations of the tower crane based on its status information and acquiring the operational video of the tower crane during the effective lifting operations includes: The status information of the tower crane is obtained, and a lifting height time series and a lifting weight time series are generated based on the status information. The status information includes the lifting height and lifting weight of the tower crane, and the lifting height is the height of the hook of the tower crane. The hoisting height time series is compared with the corresponding hoisting height threshold, and the first time period is obtained based on the comparison result. The lifting weight time series is compared with the corresponding lifting weight threshold, and a second time period is obtained based on the comparison result. The first time period and the second time period are time-aligned to obtain the effective lifting counts of the tower crane, and the operation video of the tower crane within the time period of the effective lifting counts is obtained.

3. The method for identifying the type of load lifted by a tower crane as described in claim 2, characterized in that, The step of aligning the first time period and the second time period to obtain the effective lifting operations of the tower crane, and acquiring the operation video of the tower crane within the time period of the effective lifting operations, includes: The first time period and the second time period are time-aligned to determine whether the first time period and the second time period overlap. When an overlapping time period is determined, the number of lifting operations included in the overlapping time period is determined as the valid lifting operations of the tower crane, and the operation video of the tower crane within the overlapping time period is acquired.

4. The method for identifying the type of load lifted by a tower crane as described in claim 1, characterized in that, The preprocessing and object identification of the operation video yields a candidate region set for each video frame in the operation video, including: Foreground detection is performed on the operation video to obtain a motion image of the tower crane's hook; Feature extraction is performed on the motion image to obtain temporal motion features and spatial texture features. The temporal motion features and spatial texture features are then fused to obtain a motion saliency image corresponding to the motion image. The motion saliency image is cropped using a multi-scale sliding window to obtain a set of candidate regions corresponding to the motion image.

5. The method for identifying the type of load lifted by a tower crane as described in claim 1, characterized in that, The process of detecting and identifying each candidate region in the candidate region set, and performing instance analysis on each candidate region based on the processing results to obtain instance features of the instances contained in the corresponding video frames, includes: A pre-set construction site building material recognition model is used to perform hook detection and building material classification and recognition on each candidate region in the candidate region set, so as to obtain the candidate hanging object region corresponding to the candidate region containing hooks in the candidate region set; The candidate suspended object region is segmented and its centroid is calculated based on the instance segmentation model to obtain the instances contained in the candidate suspended object region and the centroid coordinates of each instance.

6. The method for identifying the type of load lifted by a tower crane as described in claim 5, characterized in that, The step of analyzing and processing the instance features, determining the target object and its location information in the instance, and determining the type of the target object to be lifted in the effective lifting operation based on the location information includes: The candidate suspended object region is filtered using trajectory continuity to obtain intermediate instances by filtering out instances of discontinuous motion. The intermediate instance is matched with the motion of the hook in the corresponding video frame to determine the target area to which the intermediate instance belongs in the candidate hanging object area; Spatially match the target region with the recognition result obtained from classification and identification, and use the type of the target object contained in the video frame to which the target region belongs as the type of the target object under the effective lifting operation.

7. The method for identifying the type of load lifted by a tower crane as described in claim 6, characterized in that, The step of spatially matching the target region with the recognition result obtained through classification and identification, and taking the type of the target object contained in the video frame to which the determined target region belongs as the type of the target object under the effective lifting operation, further includes: Spatially match the target area with the recognition results obtained from classification to obtain the type of hanging object corresponding to each video frame; Based on the type of object to be suspended for each video frame, a majority voting mechanism is used to determine the type of object to be suspended for the target object under the effective number of suspensions.

8. A device for identifying the type of load lifted by a tower crane, characterized in that, include: The video acquisition module is used to determine the effective lifting operations of the tower crane based on the tower crane's status information, and to acquire the operation video of the tower crane under the effective lifting operations; The video processing module is used to preprocess the operation video and identify the suspended object to obtain a set of candidate regions corresponding to each video frame in the operation video; The instance analysis module is used to detect and identify each candidate region in the candidate region set, and perform instance analysis on each candidate region based on the processing results to obtain the instance features of the instance contained in the corresponding video frame. The type determination module is used to analyze and process the instance features, determine the target object and its position information in the instance, and determine the type of the target object under the effective lifting operations based on the position information.

9. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.