Intelligent tower crane remote operation method and system based on vision and laser radar fusion recognition

Through the fusion recognition method of vision and lidar, a three-dimensional model of high-precision tower crane is generated, which solves the problem that a single sensor in traditional tower crane operation is difficult to fully reflect the environment, and achieves the accuracy and safety improvement of tower crane operation.

CN119858860BActive Publication Date: 2025-08-12SHANGHAI CONSTR NO 5 GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510235668.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-08-12
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Traditional tower crane operation relies on a single sensor, which is difficult to comprehensively and accurately reflect the complex construction environment, increasing the difficulty of operator operation.

Method used

The visual and lidar fusion recognition method is adopted to collect point cloud data through lidars at multiple preset positions, and the visual camera collects environmental images, and data fusion is carried out to generate a high-precision three-dimensional model of the tower rig, and the target position information is obtained by combining Beidou positioning to realize real-time operation status recognition and remote operation.

Benefits of technology

It improves the accuracy and safety of tower crane operation, reduces the operator's information processing burden, and improves construction efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119858860B_ABST
    Figure CN119858860B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides an intelligent tower crane remote operation method and system based on vision and laser radar fusion recognition, belonging to the field of tower crane operation control technology. The method includes: fusing first point cloud data, second point cloud data, and third point cloud data to obtain target point cloud data; fusing an environmental image and the target point cloud data to obtain an initial three-dimensional tower crane model of a target tower crane; obtaining target position information corresponding to a target object in the initial three-dimensional tower crane model based on Beidou positioning and adding the target position information to the initial three-dimensional tower crane model to obtain a target three-dimensional tower crane model; performing real-time operation status recognition based on the target three-dimensional tower crane model to obtain a target operation status; sending the target three-dimensional tower crane model and the target operation status to a terminal device so that the terminal device displays the target three-dimensional tower crane model and the target operation status to a target user; obtaining a control instruction sent by the terminal device and controlling the target tower crane to perform a tower crane operation according to the control instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tower crane operation control, and in particular to an intelligent tower crane remote operation method and system based on vision and laser radar fusion recognition. Background Art

[0002] Tower cranes are essential equipment on construction sites, primarily used for vertical transportation in high-rise buildings. The safety and precision of their operation directly impact construction efficiency and personnel safety. However, traditional tower crane operation typically relies on direct observation and control from the ground or in a control room. This approach not only requires high operator experience and skills, but is also subject to limitations in vision, environmental interference, and operational precision, making it prone to operational errors and potentially leading to safety accidents.

[0003] Several auxiliary technologies have been developed to improve the safety and efficiency of tower crane operations. However, most of these technologies rely solely on vision or lidar, lacking effective data fusion methods. This makes it difficult to fully and accurately reflect the real-world environment surrounding the tower crane. Especially in complex and changing construction environments, data from a single sensor often has limitations and cannot fully capture the various factors that may affect operational safety. There is an urgent need for an intelligent tower crane remote operation method that can fuse multi-source sensor data, provide a comprehensive three-dimensional environmental model, and provide real-time status monitoring. Summary of the Invention

[0004] The main purpose of the embodiments of the present invention is to provide an intelligent tower crane remote operation method and system based on vision and lidar fusion recognition, aiming to solve the problem that most related technologies use vision or lidar alone, lack effective data fusion means, and are difficult to fully and accurately reflect the real environment around the tower crane, thereby increasing the difficulty of operation for operators.

[0005] In a first aspect, an embodiment of the present invention provides an intelligent tower crane remote operation method based on vision and laser radar fusion recognition, comprising:

[0006] Using a first laser radar at a first preset position to collect first point cloud data corresponding to the surrounding environment of the target tower crane, and using a second laser radar at a second preset position to collect second point cloud data corresponding to the surrounding environment of the target tower crane;

[0007] Using a third laser radar at a third preset position to collect third point cloud data corresponding to the surrounding environment of the target tower crane, and using a visual camera at the third preset position to collect an environmental image corresponding to the surrounding environment of the target tower crane;

[0008] Performing data fusion on the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data;

[0009] Fusing the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane;

[0010] Obtaining target position information corresponding to a target object in the initial tower crane three-dimensional model according to Beidou positioning, and adding the target position information to the initial tower crane three-dimensional model to obtain a target tower crane three-dimensional model;

[0011] Performing real-time operating state recognition based on the target tower crane three-dimensional model to obtain a target operating state corresponding to the target tower crane;

[0012] Sending the target tower crane three-dimensional model and the target operating status to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and the target operating status to a target user;

[0013] Obtain the control instruction sent by the terminal device, and control the target tower crane to perform tower crane operations according to the control instruction.

[0014] In a second aspect, an embodiment of the present invention provides an intelligent tower crane remote control system based on vision and laser radar fusion recognition, including:

[0015] A first acquisition module is configured to acquire first point cloud data corresponding to the surrounding environment of a target tower crane using a first laser radar at a first preset position, and to acquire second point cloud data corresponding to the surrounding environment of the target tower crane using a second laser radar at a second preset position;

[0016] A second acquisition module is configured to acquire third point cloud data corresponding to the surrounding environment of the target tower crane using a third laser radar at a third preset position and to acquire an environmental image corresponding to the surrounding environment of the target tower crane using a visual camera at the third preset position;

[0017] A first fusion module is used to perform data fusion on the first point cloud data, the second point cloud data and the third point cloud data to obtain target point cloud data;

[0018] A second fusion module is used to fuse the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane;

[0019] A model adjustment module is used to obtain target position information corresponding to the target object in the initial tower crane three-dimensional model according to Beidou positioning, and add the target position information to the initial tower crane three-dimensional model to obtain a target tower crane three-dimensional model;

[0020] A state recognition module is used to perform real-time operation state recognition based on the three-dimensional model of the target tower crane to obtain a target operation state corresponding to the target tower crane;

[0021] a data sending module, configured to send the target tower crane three-dimensional model and the target operating status to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and the target operating status to a target user;

[0022] The instruction execution module is used to obtain the control instruction sent by the terminal device and control the target tower crane to perform tower crane operations according to the control instruction.

[0023] An embodiment of the present invention provides an intelligent tower crane remote operation method and system based on vision and laser radar fusion recognition, the method comprising: using a first laser radar at a first preset position to collect first point cloud data corresponding to the surrounding environment of the target tower crane, and using a second laser radar at a second preset position to collect second point cloud data corresponding to the surrounding environment of the target tower crane; using a third laser radar at a third preset position to collect third point cloud data corresponding to the surrounding environment of the target tower crane and using a visual camera at a third preset position to collect an environmental image corresponding to the surrounding environment of the target tower crane; performing data fusion on the first point cloud data, the second point cloud data and the third point cloud data to obtain target point cloud data; fusing the environmental image The method uses the target point cloud data to obtain an initial 3D model of the target crane corresponding to its surroundings. The method then uses Beidou positioning to obtain target position information corresponding to the target object in the initial 3D model and adds the target position information to the initial 3D model to obtain the target crane model. The method then performs real-time operating status recognition based on the target crane model to obtain the target operating status of the target crane. The target crane model and the target operating status are then transmitted to a terminal device connected to the target crane, allowing the terminal device to display the model and the target operating status to a target user. The method then receives control instructions from the terminal device and controls the target crane to perform crane operations based on the control instructions. This method utilizes laser radars installed at different preset locations to collect point cloud data from all directions and angles around the target crane, including first, second, and third point cloud data. By fusing these point cloud data, more comprehensive and accurate target point cloud data can be generated. Furthermore, by combining these point cloud data with environmental images from a visual camera, a high-precision initial 3D model of the crane can be generated, significantly improving the accuracy and comprehensiveness of environmental perception. Beidou positioning is used to obtain the target object's target position information and add it to the initial 3D tower crane model to create the target crane model. This provides the operator with real-time and accurate position information of the target object, significantly improving operational accuracy. The target crane's 3D model is used to identify its operating status in real time, determining its target operating status and helping the operator make more accurate operational decisions. The 3D model and target operating status are then transmitted to a remote terminal device, providing the operator with detailed operational guidance and environmental information, reducing their information processing burden and improving operational efficiency. Finally, based on the control instructions sent by the terminal device, the tower crane is automatically controlled to perform the lifting operation, reducing the operator's workload, improving efficiency, and shortening operation time, thereby significantly improving overall construction efficiency. This addresses the problem that most related technologies, which rely solely on vision or lidar, lack effective data fusion methods, making it difficult to fully and accurately reflect the real environment around the tower crane, thereby increasing operator difficulty. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 A flowchart of an intelligent tower crane remote operation method based on vision and laser radar fusion recognition provided by an embodiment of the present invention;

[0026] Figure 2 A schematic diagram of the module structure of an intelligent tower crane remote operation system based on vision and laser radar fusion recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0029] It should be understood that the terms used in this specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0030] Embodiments of the present invention provide a method and system for remotely operating an intelligent tower crane based on fusion recognition of vision and laser radar. This method can be applied to a terminal device, such as a tablet computer, laptop computer, desktop computer, personal digital assistant, or wearable device. The terminal device can also be a server or a server cluster.

[0031] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0032] Please refer to Figure 1 , Figure 1 A flowchart of an intelligent tower crane remote operation method based on vision and lidar fusion recognition is provided in an embodiment of the present invention.

[0033] like Figure 1 As shown, the intelligent tower crane remote operation method based on vision and laser radar fusion recognition includes steps S101 to S108.

[0034] Step S101: Use a first laser radar at a first preset position to collect first point cloud data corresponding to the surrounding environment of a target tower crane, and use a second laser radar at a second preset position to collect second point cloud data corresponding to the surrounding environment of the target tower crane.

[0035] Exemplarily, the first and second preset positions are fixedly located at either the base or tip of the boom of the target tower crane, thereby ensuring that the LiDAR can be securely installed and acquire stable environmental data. The base of the boom serves as the core support portion of the tower crane, and the LiDAR installed there can provide comprehensive basic environmental information about the tower crane; while the LiDAR at the tip of the boom can capture environmental data at greater distances and higher angles, which is particularly important for environmental perception during high-altitude operations. Therefore, when the first preset position is the base of the boom of the target tower crane, the second preset position is the tip of the boom of the target tower crane. When the first preset position is the tip of the boom of the target tower crane, the second preset position is the base of the boom of the target tower crane.

[0036] For example, the first preset location is the base of the target tower crane's boom, so a first laser radar is installed at the first preset location, thereby obtaining first point cloud data corresponding to the target tower crane's surroundings captured by the first laser radar. The second preset location is the tip of the target tower crane's boom, so a second laser radar is installed at the second preset location, thereby obtaining second point cloud data corresponding to the target tower crane's surroundings captured by the second laser radar.

[0037] Step S102: using a third laser radar at a third preset position to collect third point cloud data corresponding to the surrounding environment of the target tower crane, and using a visual camera at the third preset position to collect an environmental image corresponding to the surrounding environment of the target tower crane.

[0038] For example, the third preset position is flexibly set at the movable position of the target tower crane's boom. In essence, the third preset position corresponds to the position of the flexible movable component of the tower crane's boom. This configuration allows the LiDAR to adjust its position in response to the boom's movement during operation, capturing changes in the boom's surroundings in real time. For example, when the boom rotates or extends, the LiDAR at the third preset position can adjust its angle and direction at will, ensuring continuous and accurate environmental data under various operating conditions.

[0039] Exemplarily, a third laser radar is installed at a third preset position, thereby obtaining third point cloud data corresponding to the surrounding environment of the target tower crane collected by the third laser radar.

[0040] For example, through this combination of fixed and flexible setup, the LiDAR sensors at the first, second, and third preset locations work together to collect comprehensive, multi-angle environmental data around the tower crane. The fixed LiDAR sensors provide basic and stable perception data, while the mobile LiDAR sensors dynamically adjust based on the crane's operating status, providing wider coverage and more comprehensive data. This design not only improves tower crane operation safety but also provides operators with more accurate, real-time environmental information, enhancing overall operational efficiency.

[0041] Exemplarily, the visual camera is set at the third preset position at an angle that can obtain the surrounding environment of the target tower base, so that the visual camera collects an environmental image corresponding to the surrounding environment of the target tower crane.

[0042] Step S103: performing data fusion on the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data.

[0043] Exemplarily, the collected first point cloud data, second point cloud data, and third point cloud data are subjected to denoising processing to remove invalid points and noise points caused by environmental interference, equipment noise, etc., to ensure the purity of the data.

[0044] For example, the denoised first, second, and third point cloud data are initially registered, aligning the point cloud data collected at different locations to the same coordinate system for subsequent precise registration and fusion. A precise registration algorithm, such as the Iterative Closest Point (ICP) algorithm or the Normal Distribution Transform (NDT) algorithm, is then used to precisely register the first, second, and third point cloud data, ensuring precise spatial positioning and alignment of the point cloud data. The aligned first, second, and third point cloud data are then reassembled and optimized to generate the target point cloud data. The target point cloud data should contain comprehensive information from the first, second, and third point cloud data, covering the entire area surrounding the target tower crane.

[0045] In some embodiments, the data fusion of the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data includes: performing data clustering on the first point cloud data to obtain a first clustering result and a first target type corresponding to the first clustering result; performing data clustering on the second point cloud data to obtain a second clustering result and a second target type corresponding to the second clustering result; performing data clustering on the third point cloud data to obtain a third clustering result and a third target type corresponding to the third clustering result; using a target recognition model to perform target recognition on the environment image to obtain a corresponding target recognition type; and using the first target type to identify the target according to the target recognition type. The first clustering result is filtered to obtain a first point cloud cluster; the second clustering result is filtered using the second target type according to the target recognition type to obtain a second point cloud cluster; the third clustering result is filtered using the third target type according to the target recognition type to obtain a third point cloud cluster; the first point cloud cluster and the third point cloud cluster are data aligned to obtain a first alignment result, and the second point cloud cluster and the third point cloud cluster are data aligned to obtain a second alignment result; the first point cloud data, the second point cloud data and the third point cloud data are data fused according to the first alignment result and the second alignment result to obtain the target point cloud data.

[0046] Exemplarily, a clustering algorithm, such as K-means, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or Mean Shift, is selected to cluster the first point cloud data to obtain a first clustering result, thereby classifying the first clustering result and determining a first target type corresponding to each cluster cluster.

[0047] Exemplarily, the above clustering steps are repeated for the second point cloud data and the third point cloud data, respectively, to obtain a second clustering result and a third clustering result, and determine the corresponding second target type and third target type.

[0048] For example, a target recognition model, such as a deep learning model (e.g., YOLO or Faster R-CNN), is used to identify targets in the environment image and obtain the corresponding target recognition type. The target recognition type should match the target type in the clustering result of the point cloud data.

[0049] For example, based on the target recognition type, the first clustering result is screened for a first target type that matches the target recognition type to obtain a first point cloud cluster. Based on the target recognition type, the second clustering result is screened for a second target type that matches the target recognition type to obtain a second point cloud cluster. Based on the target recognition type, the third clustering result is screened for a third target type that matches the target recognition type to obtain a third point cloud cluster.

[0050] For example, a point cloud registration algorithm, such as ICP or NDT, is used to align the first point cloud cluster and the third point cloud cluster to obtain a first alignment result. The first and third point cloud clusters are aligned in the same coordinate system. The same registration algorithm is used to align the second and third point cloud clusters to obtain a second alignment result. The second and third point cloud clusters are aligned in the same coordinate system. The coordinate system corresponding to the third point cloud cluster is used as the target coordinate system in both the first and second alignment results.

[0051] For example, the first, second, and third point cloud data are fused based on a feature point matching algorithm in combination with the first and second alignment results to obtain target point cloud data. The target point cloud data should include comprehensive information from the first, second, and third point cloud data, and cover the entire area surrounding the target tower crane.

[0052] In some embodiments, the use of a target recognition model to perform target recognition on the environmental image to obtain a corresponding target recognition type includes: using the input layer of the target recognition model to perform a preprocessing operation on the environmental image to obtain a preprocessed image; using the feature extraction layer of the target recognition model to perform feature extraction on the preprocessed image to obtain initial feature information; using the variable convolutional layer of the target recognition model to add a learnable offset and an adjustment factor at the sampling position to adaptively perform feature sampling on the initial feature information to obtain target feature information; and using the target recognition layer of the target recognition model to perform target recognition on the target feature information to obtain the target recognition type.

[0053] Exemplarily, the environment image is input to the input layer of the object recognition model. Preprocessing operations are performed on the environment image to improve image quality and model recognition performance. Preprocessing operations include, but are not limited to, image scaling and normalization.

[0054] For example, the feature extraction layer of the object recognition model is used to extract features from the preprocessed image. This layer typically includes multiple layers of convolution operations to extract low-level and high-level features from the image. The features extracted by the convolution operation are called initial feature information. These features typically contain the basic structure and local details of the image.

[0055] For example, a variable convolutional layer in an object recognition model adds a learnable offset and adjustment factor to the sampling position. By learning these offsets and adjustment factors, the variable convolutional layer adaptively adjusts the position and shape of the convolution kernel to accommodate different image features and structures. Based on these offsets and adjustment factors, the initial feature information is sampled to obtain target feature information. This adaptively adjusted target feature information better captures key information and structure in the image.

[0056] For example, the target recognition model uses the target recognition layer to identify the target feature information. The target recognition layer typically includes a fully connected layer and a classifier to convert the feature information into a target recognition type. The target recognition layer then classifies the target feature information to obtain the final target recognition type. The target recognition type should match the predefined target category.

[0057] Specifically, the input layer of the target recognition model is used to preprocess the environmental image, and the feature extraction layer, variable convolution layer and target recognition layer are used to extract features and recognize targets in the image, so as to obtain accurate target recognition types and provide reliable environmental perception information for subsequent point cloud data processing and tower crane operation.

[0058] Step S104: fusing the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane.

[0059] Exemplarily, image target detection is performed on the environment image to obtain a first target position corresponding to the environment image and a first target type corresponding to the first target position. 3D target detection is performed on the target point cloud data to obtain a second target position corresponding to the target point cloud data and a second target type corresponding to the second target position.

[0060] For example, target coordinate point detection is performed on the second target position to obtain a minimum rectangular frame that can cover the target object corresponding to the second target position, and the third target position corresponding to the minimum rectangular frame is obtained. In addition, if the visual camera and the third preset position are in the same position, it can be determined that the third target position and the first target position are in the same coordinate system.

[0061] Exemplarily, the degree of overlap between the first target position and the third target position is calculated, and then when the degree of overlap is greater than a preset value, the first target type and the second target type are compared to see whether they are the same. Thus, when the degree of overlap between the first target position and the third target position is greater than the preset value and the first target type and the second target type are the same, the target type corresponding to the associated point cloud data corresponding to the second target position is determined; when the degree of overlap between the first target position and the third target position is less than or equal to the preset value or the first target type and the second target type are different, the target type corresponding to the associated point cloud data corresponding to the second target position cannot be determined, and target detection is then re-performed on the target point cloud data and the environmental image until the target type corresponding to the associated point cloud data corresponding to the second target position is determined.

[0062] Exemplarily, an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane is determined according to the target type and the target point cloud data.

[0063] In some embodiments, the fusing of the environmental image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane includes: performing image enhancement on the environmental image to obtain a corresponding target image; performing edge recognition on the target image to obtain target edge information corresponding to the target image, and performing target recognition on the target edge information according to a convolutional neural network to obtain a first recognition result corresponding to the target image; performing target recognition on the target point cloud data to obtain a second recognition result corresponding to the target point cloud data; and performing data fusion on the environmental image and the target point cloud data according to the first recognition result and the second recognition result to obtain the initial three-dimensional tower crane model.

[0064] For example, image enhancement is performed on the environment image to improve image quality and target recognition accuracy. Common image enhancement methods include, but are not limited to, contrast enhancement, brightness adjustment, denoising, and sharpening. The target image is thus obtained.

[0065] Exemplarily, Canny edge detection, Sobel operator, Laplacian operator, etc. are used to perform edge recognition on the target image to extract target edge information in the target image.

[0066] For example, a convolutional neural network (CNN) is used to identify target edges. Through multiple layers of convolution and pooling, the CNN extracts high-level features from the image. It then uses fully connected layers and a classifier to classify the target, obtaining a first recognition result corresponding to the environment image. This first recognition result includes the category and location of the target in the environment image. Specifically, the first recognition result includes the first target location and the first target type corresponding to the first target location.

[0067] Exemplarily, a second recognition result is obtained by performing target recognition on the target point cloud data using an object recognition algorithm, such as a deep learning model (e.g., PointNet or PointNet++). The second recognition result includes the category and location information of the target in the target point cloud data. Specifically, the second recognition result includes the second target location and the second target type corresponding to the second target location.

[0068] For example, target coordinate point detection is performed on the second target position to obtain a minimum rectangular frame that can cover the target object corresponding to the second target position, and the third target position corresponding to the minimum rectangular frame is obtained. In addition, if the visual camera and the third preset position are in the same position, it can be determined that the third target position and the first target position are in the same coordinate system.

[0069] Exemplarily, the degree of overlap between the first target position and the third target position is calculated, and then when the degree of overlap is greater than a preset value, the first target type and the second target type are compared to see whether they are the same. Thus, when the degree of overlap between the first target position and the third target position is greater than the preset value and the first target type and the second target type are the same, the target type corresponding to the associated point cloud data corresponding to the second target position is determined; when the degree of overlap between the first target position and the third target position is less than or equal to the preset value or the first target type and the second target type are different, the target type corresponding to the associated point cloud data corresponding to the second target position cannot be determined, and target detection is then re-performed on the target point cloud data and the environmental image until the target type corresponding to the associated point cloud data corresponding to the second target position is determined.

[0070] Exemplarily, an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane is determined according to the target type and the target point cloud data.

[0071] In some embodiments, the image enhancement of the environmental image to obtain the corresponding target image includes: performing brightness feature extraction on the environmental image to obtain a corresponding first brightness feature and performing saturation feature extraction on the environmental image to obtain a corresponding first saturation feature; performing data enhancement on the first brightness feature to obtain a second brightness feature, and using a gamma correction algorithm to correct the second brightness feature to obtain a third brightness feature; determining a target window, and obtaining a window brightness feature corresponding to the third brightness feature under the target window; performing regional histogram equalization on the window brightness feature to obtain a corresponding target brightness feature, and recombining according to the target brightness feature to obtain a fourth brightness feature; performing bilateral filtering on the first saturation feature to obtain a second saturation feature, and performing a stretching transformation on the second saturation feature to obtain a third saturation feature; obtaining the corresponding target image based on the fourth brightness feature and the third saturation feature.

[0072] Exemplarily, the environment image is converted into the HSV space to perform brightness feature extraction to obtain a first brightness feature corresponding to the environment image, and saturation feature extraction is performed to obtain a first saturation feature corresponding to the environment image.

[0073] For example, a data enhancement operation, such as contrast enhancement, brightness adjustment, etc., is performed on the first brightness feature to obtain the second brightness feature. Data enhancement can improve the contrast and detail information of the environment image.

[0074] For example, a gamma correction algorithm is applied to the second brightness feature to perform correction processing to obtain the third brightness feature. Gamma correction can adjust the brightness and contrast of an image to improve the visual effect of the image.

[0075] For example, a target window is determined, and the window brightness feature of the third brightness feature under the target window is extracted. The window brightness feature only contains brightness information within the target window. Regional histogram equalization is then performed on the window brightness feature to make the brightness features within the target window more balanced, thereby obtaining the target brightness feature. Regional histogram equalization can enhance the contrast and details of the target region. Finally, the third brightness feature is recombined based on the target brightness feature to obtain the fourth brightness feature.

[0076] For example, bilateral filtering is performed on the first saturation feature to obtain a second saturation feature. Bilateral filtering can smooth an image while preserving edge information, improving image smoothness and detail quality. The second saturation feature is then subjected to a stretch transformation to obtain a third saturation feature. The stretch transformation can adjust the saturation range of the image, improving color saturation and visual quality.

[0077] Exemplarily, the environmental image is converted into the HSV space to perform hue feature extraction to obtain the hue feature corresponding to the environmental image, and then feature fusion is performed based on the fourth brightness feature, the third saturation feature and the hue feature to obtain a new HSV image, and then the new HSV image is converted into the RGB space to obtain the target image.

[0078] In some embodiments, the edge recognition of the target image to obtain target edge information corresponding to the target image includes: performing a wavelet transform on the target image to obtain a first transform result of the target image in the horizontal and vertical directions; performing a wavelet transform on the target image to obtain a second transform result of the target image in the horizontal direction and performing a wavelet transform on the target image to obtain a third transform result of the target image in the vertical direction; determining a target gradient corresponding to the target image based on the first transform result, the second transform result, and the third transform result; determining a segmentation threshold corresponding to the target image based on the target gradient, and determining a target edge point corresponding to the target image based on the segmentation threshold; and performing data clustering on the target edge points to obtain target edge information corresponding to the target image.

[0079] Exemplarily, a wavelet transform is performed on a target image to obtain first transform results for the target image in the horizontal and vertical directions. The first transform results include low-frequency information of the target image in both the horizontal and vertical directions. A wavelet transform is performed on the target image to obtain a second transform result for the target image in the horizontal direction. The second transform result primarily includes high-frequency information of the target image in the horizontal direction. A wavelet transform is performed on the target image to obtain a third transform result for the target image in the vertical direction. The third transform result primarily includes high-frequency information of the target image in the vertical direction.

[0080] Exemplarily, a first optical component in the horizontal direction and a second optical component in the vertical direction of the target image are obtained, and then the first optical component and the second optical component are squared, and then the square root of the sum is taken to obtain the target optical component. The second transformation result and the third transformation result are then summed, and the summed result is multiplied by the first transformation result to obtain the target transformation result. The target optical component and the target transformation result are then divided to obtain the target gradient at the corresponding position of the target image.

[0081] Exemplarily, a gradient calculation is performed on the target image to obtain the gradient information of each pixel in the target image. The gradient information of each pixel is ratio-calculated with the gradient information of its adjacent pixels. The purpose of the ratio calculation is to highlight the areas in the image where the gradient changes significantly, so as to better identify edge points. After the ratio calculation, a non-maximum suppression operation is further performed. The purpose of the non-maximum suppression operation is to retain only the pixels with the largest local gradient, thereby reducing the number of edge points and making the edges clearer and more accurate. After the non-maximum suppression operation, the pixels with the largest local gradient are called maximum points. These maximum points may be potential edge points, but further processing is required to determine the true edges. The arg function is calculated on the maximum points to obtain the segmentation threshold. The arg function is usually used to calculate the relative position of the extreme points in the gradient direction, so as to determine the appropriate segmentation threshold. The segmentation threshold is used to distinguish edge points from non-edge points.

[0082] For example, the target image is edge-identified based on the segmentation threshold. When the pixel value corresponding to the target image is greater than or equal to the segmentation threshold, the pixel value is determined as the target edge point corresponding to the target image. The target edge point accurately depicts the outline and boundary of the target object.

[0083] For example, data clustering is performed on the extracted target edge points using K-means clustering, DBSCAN clustering, hierarchical clustering, etc. By clustering the target edge points, target edge information corresponding to the target image is obtained. The target edge information includes the target's outline, structure, and shape information.

[0084] In some embodiments, the data clustering of the target edge points to obtain the target edge information corresponding to the target image includes: determining an initial cluster center and obtaining distance information between the target edge point and the initial cluster center; determining the target probability that the target edge point belongs to the initial cluster center based on the distance information; calculating the first membership corresponding to the target edge point belonging to the initial cluster center, and determining the cluster target value corresponding to the initial cluster center based on the distance information and the first membership; updating the first membership based on the cluster target value and the target probability to obtain a second membership; updating the initial cluster center based on the second membership to obtain a target cluster center; performing data clustering on the target edge points based on the target cluster center to obtain the target edge information corresponding to the target image; wherein the cluster target value and the second membership are obtained according to the following formula:

[0085]

[0086]

[0087] in, represents the clustering target value, cer represents the number of the initial cluster centers, and m represents the number of target edge points. represents the segmentation threshold, represents the pixel value corresponding to the jth target edge point, represents the number of pixel values corresponding to the j-th target edge point, Indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the pixel value corresponding to the i-th initial cluster center; The second degree of membership indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the probability that the j-th target edge point belongs to the target corresponding to the i-th initial cluster center, Indicates the minimum value of the target probability that the j-th target edge point belongs to the i-th initial cluster center.

[0088] For example, initial cluster centers are determined from the target edge points. The initial cluster centers can be determined by random selection, K-means++ initialization, or other initialization methods. The number of initial cluster centers is pre-set based on actual needs.

[0089] Exemplarily, the distance information between each target edge point and each initial cluster center is calculated using Euclidean distance, Manhattan distance, or other distance measurement methods to quantify the similarity between the target edge point and the initial cluster center.

[0090] For example, the calculated distance information is used to determine the target probability of each target edge point belonging to each initial cluster center. The target probability can be calculated using a Gaussian distribution function or other probability distribution function, so that target edge points with closer distances have a higher probability of belonging to the corresponding cluster center.

[0091] For example, the first membership degree of each target edge point to each initial cluster center is calculated. The first membership degree can be calculated by target probability or other membership functions, and represents the degree of belonging of the target edge point to each cluster center.

[0092] For example, the clustering target value corresponding to each initial cluster center is calculated based on the distance information and the first membership degree. The clustering target value can reflect the clustering effect of the initial cluster center. The clustering target value is obtained according to the following formula:

[0093]

[0094] in, Represents the clustering target value, cer represents the number of initial cluster centers, and m represents the number of target edge points. represents the segmentation threshold, Represents the pixel value corresponding to the j-th target edge point, Indicates the number of pixel values corresponding to the j-th target edge point, Indicates the first membership degree of the j-th target edge point to the i-th initial cluster center, Represents the pixel value corresponding to the i-th initial cluster center.

[0095] For example, by calculating the cluster target value, the clustering effect of the initial cluster centers can be quantified. The cluster target value reflects the performance of each initial cluster center during the clustering process and provides an objective indicator for evaluating the clustering effect. The cluster target value is calculated based on distance information and the first membership degree, which helps optimize the position of the initial cluster centers. By adjusting the cluster target value, the cluster centers can be brought closer to the actual cluster center positions, thereby improving clustering accuracy. The calculation formula of the cluster target value comprehensively considers the pixel values of the target edge points, the segmentation threshold, the first membership degree, and the pixel values of the initial cluster centers. These factors work together to enable the cluster target value to more accurately reflect the clustering effect, thereby improving clustering accuracy. By calculating the cluster target value, the stability of the clustering process can be enhanced. As a dynamically adjusted parameter, the cluster target value can provide feedback in each iteration, ensuring that the clustering process proceeds in a more optimal direction and reducing the instability of the clustering results.

[0096] For example, the first membership is updated according to the clustering target value and the target probability to obtain the second membership. The second membership is obtained according to the following formula:

[0097]

[0098] in, Represents the clustering target value, cer represents the number of initial cluster centers, and m represents the number of target edge points. Indicates the second membership degree of the jth target edge point to the ith initial cluster center, Indicates the target probability that the jth target edge point belongs to the target corresponding to the i-th initial cluster center, Indicates the minimum value of the target probability that the j-th target edge point belongs to the i-th initial cluster center.

[0099] Exemplarily, the first membership is updated by the clustering target value and target probability, and the obtained second membership can more accurately reflect the degree of attribution of the target edge point to the cluster center. Compared with the first membership, the second membership integrates more information, making the calculation of the membership more accurate. The calculation of the second membership takes into account the clustering target value and target probability, which helps to optimize the clustering results. By adjusting the membership, the cluster center is brought closer to the actual cluster center position, thereby improving the accuracy and stability of clustering. The update process of the second membership dynamically adjusts the degree of attribution of the target edge point to the cluster center, making the clustering process more stable. By updating the membership, feedback can be provided in each iteration to ensure that the clustering process proceeds in a more optimal direction.

[0100] Exemplarily, the initial cluster center is updated based on the second membership degree to obtain a target cluster center. The updating method is to perform a weighted average based on the contribution of the target edge point to each cluster center, so that the cluster center is closer to the actual cluster center location. Using the updated target cluster center, data clustering is performed on the target edge points until the target cluster center no longer changes. Thus, data clustering of the target edge points based on the target cluster center obtains target edge information corresponding to the target image.

[0101] Step S105: obtaining target position information corresponding to the target object in the initial tower crane three-dimensional model according to Beidou positioning, and adding the target position information to the initial tower crane three-dimensional model to obtain a target tower crane three-dimensional model.

[0102] For example, the target position information of target objects (such as hooks and booms) in the initial 3D tower crane model is obtained based on Beidou positioning. The relative positional relationship between the remaining objects in the initial 3D tower crane model and the target object is then determined. Based on the relative positional relationship and the target positional information, the associated positional relationship corresponding to the remaining objects in the initial 3D tower crane model is determined. The associated positional relationship and the target positional information are then added to the initial 3D tower crane model to obtain the target 3D tower crane model. In other words, the target 3D tower crane model now contains the actual positional information corresponding to each target object.

[0103] In some embodiments, obtaining the target position information corresponding to the target object in the initial tower crane three-dimensional model based on Beidou positioning includes: obtaining the target hanging object corresponding to the target tower crane from the initial tower crane three-dimensional model; obtaining the hanging object position information corresponding to the target hanging object using the Beidou positioning; and calculating the position of the target object in the initial tower crane three-dimensional model based on the hanging object position information to obtain the target position information corresponding to the target object.

[0104] For example, a target load corresponding to the target crane is identified from the initial 3D crane model. The target load typically refers to an object being lifted by the crane, such as construction materials or mechanical equipment. The target load is determined using identifiers or feature points in the initial 3D crane model.

[0105] For example, the BeiDou positioning system is activated to ensure its normal operation. The BeiDou positioning system determines the target hanging object's location information by receiving satellite signals. The hanging object's location information generally includes longitude, latitude, and altitude to form a three-dimensional coordinate point.

[0106] Exemplarily, the relative position between the target object and the target hanging object in the initial tower crane three-dimensional model is calculated to ensure that the position of the target object is consistent with the position of the hanging object, and then the target position information of the target object in the Beidou positioning system is determined based on the relative position and the hanging object position information.

[0107] Specifically, the target crane's corresponding load is obtained from the initial 3D crane model. The BeiDou positioning system is used to obtain the load's position information. This information is then used to calculate the position of the target object in the initial 3D crane model, ultimately obtaining the target position information for the target object. This process improves the accuracy and practicality of the initial 3D crane model, providing reliable data support for crane operation and monitoring.

[0108] Step S106: performing real-time operating status recognition based on the target tower crane three-dimensional model to obtain a target operating status corresponding to the target tower crane.

[0109] For example, a 3D model of the target tower crane, already acquired and containing target position information, is loaded. The target tower crane 3D model is a crucial component of the digital twin of the tower crane, accurately reflecting the crane's geometry and the spatial relationships between its components. The 3D model should already accurately embed the real-time position information of target objects (such as the hook and boom), laying the foundation for subsequent operational status identification and analysis.

[0110] For example, key operating state parameters are extracted from the three-dimensional model of the target tower crane. These parameters are quantitative indicators that describe the tower crane's operating state at a specific moment. They primarily include: boom angle, hook height, load position, crane speed, and the spatial relationship between the load and surrounding objects. After obtaining the operating state parameters at each moment, the rate of change of these parameters between adjacent moments is calculated. The rate of change describes how quickly the tower crane's operating state changes over time. These parameters primarily include: boom angle change rate, hook height change rate, load movement speed, and the corresponding degree of change in the spatial relationship between the load and surrounding objects. The calculated rate of change of operating state parameters is then input into a pre-trained state prediction model. This model utilizes a machine learning algorithm or deep learning network to accurately predict the tower crane's operating state at the next moment by learning from a large amount of historical data. The model inputs rate of change data, including boom angle change rate, hook height change rate, load movement speed, and other data. The output is the target operating state of the tower crane, which may include, but is not limited to, safe operation, abnormality, and so on.

[0111] In some embodiments, the target tower crane three-dimensional model includes a first tower crane three-dimensional model corresponding to a first moment and a second tower crane three-dimensional model corresponding to a second moment adjacent to the first moment, and the real-time operation status identification based on the target tower crane three-dimensional model to obtain the target operation status corresponding to the target tower crane includes: obtaining first feature information corresponding to the target hoist from the first tower crane three-dimensional model and obtaining second feature information corresponding to the target hoist from the second tower crane three-dimensional model; determining the change angle corresponding to each associated position in the target hoist according to the first feature information and the second feature information; and determining the maximum movement angle corresponding to the target hoist according to the change angle; performing difference calculation based on the change angle and the maximum movement angle to determine the angle difference corresponding to the associated position in the target hoist; determining the change weight corresponding to the associated position according to the change angle and the angle difference, and determining the target change value corresponding to the target tower crane according to the change weight, the first feature information and the second feature information; and determining the target operation status corresponding to the target tower crane according to the target change value.

[0112] Exemplarily, the target tower crane three-dimensional model includes a first tower crane three-dimensional model corresponding to a first moment and a second tower crane three-dimensional model corresponding to a second moment adjacent to the first moment.

[0113] For example, a first three-dimensional model of a tower crane is loaded, and first feature information of the target hanging object in the model is extracted. The first feature information may include the position, posture, shape, and relative position of the target hanging object with respect to other components of the tower crane.

[0114] For example, a second tower crane 3D model is loaded, and second feature information of the target hanging object in the model is also extracted, ensuring that the extracted second feature information is consistent with the information extracted in the first tower crane 3D model in terms of definition and representation.

[0115] For example, the first and second feature information of the target hanging object at various associated locations in the first and second 3D tower crane models are compared. The associated locations can be key points on the target hanging object or contact points with other tower crane components. By calculating the angular change of the associated locations in the two models, the change angle of the target hanging object at each associated location is determined.

[0116] For example, based on the change angles of each associated position calculated in the previous step, the maximum movement angle of the target hanging object in all associated positions is determined. This angle reflects the maximum rotation or tilt of the hanging object during the entire movement process.

[0117] For example, the difference between the change angle of each associated position and the maximum movement angle of the target hanging object is calculated to obtain the angle difference of each associated position. The angle difference reflects the size difference of each associated position relative to the maximum movement angle.

[0118] For example, the change weight of each associated position is determined based on the calculated angle difference. The change weight can be represented by the absolute value of the angle difference or other related metrics to measure the importance and contribution of the change of each associated position.

[0119] For example, a target change value for the target crane is calculated by comprehensively considering the change weight, the first feature information, and the second feature information for each associated position. The target change value is a comprehensive assessment of the overall change of the target crane, reflecting its dynamic characteristics and state changes during movement.

[0120] For example, the change information between the second feature information and the first feature information is calculated, and then a weighted sum is performed according to the change weight and the change information to obtain the target change value of the target tower crane.

[0121] Exemplarily, the calculated target change value is compared with a preset threshold. When the target change value is greater than or equal to the preset threshold, the target operating state is determined to be abnormal operation; when the target change value is less than the preset threshold, the target operating state is determined to be normal operation.

[0122] Specifically, the system acquires the characteristic information of the target load from the first and second 3D models of the tower cranes, and determines the target operating state of the target cranes. This process improves the real-time and accuracy of tower crane operations, providing reliable support for safe and efficient operation of the cranes.

[0123] Step S107: Send the target tower crane three-dimensional model and the target operating status to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and the target operating status to a target user.

[0124] Exemplarily, the target tower crane three-dimensional model and target operating status are sent to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and target operating status to the target user. Then, the target user can obtain the target tower crane three-dimensional model and target operating status according to the terminal device, and then perform tower crane operations according to the target tower crane three-dimensional model and target operating status, and determine corresponding control instructions.

[0125] Step S108: Obtain the control instruction sent by the terminal device, and control the target tower crane to perform tower crane operation according to the control instruction.

[0126] Exemplarily, the control instructions sent by the terminal device are obtained, and then the target tower crane is controlled to perform the tower crane operation according to the control instructions.

[0127] See also Figure 2 , Figure 2An intelligent tower crane remote operation system 200 based on vision and laser radar fusion recognition is provided in an embodiment of the present application. The intelligent tower crane remote operation system 200 based on vision and laser radar fusion recognition includes a first acquisition module 201, a second acquisition module 202, a first fusion module 203, a second fusion module 204, a model adjustment module 205, a state recognition module 206, a data sending module 207, and an instruction execution module 208. The first acquisition module 201 is used to use a first laser radar at a first preset position to collect first point cloud data corresponding to the surrounding environment of the target tower crane, and use a second laser radar at a second preset position to collect second point cloud data corresponding to the surrounding environment of the target tower crane; the second acquisition module 202 is used to use a third laser radar at a third preset position to collect third point cloud data corresponding to the surrounding environment of the target tower crane and use a visual camera at the third preset position to collect an environmental image corresponding to the surrounding environment of the target tower crane; the first fusion module 203 is used to collect the first point cloud data, the second point cloud data, and the second point cloud data. The cloud data and the third point cloud data are fused to obtain target point cloud data; a second fusion module 204 is used to fuse the environmental image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane; a model adjustment module 205 is used to obtain the target position information corresponding to the target object in the initial three-dimensional tower crane model according to Beidou positioning, and add the target position information to the initial three-dimensional tower crane model to obtain the target tower crane three-dimensional model; a state recognition module 206 is used to perform real-time operation state recognition according to the three-dimensional model of the target tower crane to obtain the target operation state corresponding to the target tower crane; a data sending module 207 is used to send the three-dimensional tower crane model and the target operation state to a terminal device connected to the target tower crane for communication, so that the terminal device displays the three-dimensional tower crane model and the target operation state to the target user; an instruction execution module 208 is used to obtain the control instructions sent by the terminal device, and control the target tower crane to perform tower crane operations according to the control instructions.

[0128] In some embodiments, during the process of fusing the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data, the first fusion module 203 executes:

[0129] Performing data clustering on the first point cloud data to obtain a first clustering result and a first target type corresponding to the first clustering result;

[0130] performing data clustering on the second point cloud data to obtain a second clustering result and a second target type corresponding to the second clustering result;

[0131] performing data clustering on the third point cloud data to obtain a third clustering result and a third target type corresponding to the third clustering result;

[0132] Performing target recognition on the environment image using a target recognition model to obtain a corresponding target recognition type;

[0133] Filtering the first clustering result using the first target type according to the target recognition type to obtain a first point cloud cluster;

[0134] Filtering the second clustering result using the second target type according to the target recognition type to obtain a second point cloud cluster;

[0135] Filtering the third clustering result using the third target type according to the target recognition type to obtain a third point cloud cluster;

[0136] Performing data alignment on the first point cloud cluster and the third point cloud cluster to obtain a first alignment result, and performing data alignment on the second point cloud cluster and the third point cloud cluster to obtain a second alignment result;

[0137] The target point cloud data is obtained by performing data fusion on the first point cloud data, the second point cloud data, and the third point cloud data according to the first alignment result and the second alignment result.

[0138] In some embodiments, during the process of using the target recognition model to perform target recognition on the environment image to obtain the corresponding target recognition type, the first fusion module 203 performs:

[0139] Performing a preprocessing operation on the environment image using the input layer of the target recognition model to obtain a preprocessed image;

[0140] Using the feature extraction layer of the target recognition model to perform feature extraction on the preprocessed image to obtain initial feature information;

[0141] Utilizing the variable convolutional layer of the target recognition model to add a learnable offset and an adjustment factor at the sampling position to adaptively perform feature sampling on the initial feature information to obtain target feature information;

[0142] The target recognition layer of the target recognition model is used to perform target recognition on the target feature information to obtain the target recognition type.

[0143] In some embodiments, during the process of fusing the environment image and the target point cloud data to obtain the initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane, the second fusion module 204 executes:

[0144] Performing image enhancement on the environment image to obtain a corresponding target image;

[0145] Performing edge recognition on the target image to obtain target edge information corresponding to the target image, and performing target recognition on the target edge information according to a convolutional neural network to obtain a first recognition result corresponding to the target image;

[0146] Performing target recognition on the target point cloud data to obtain a second recognition result corresponding to the target point cloud data;

[0147] The initial tower crane three-dimensional model is obtained by performing data fusion on the environment image and the target point cloud data according to the first recognition result and the second recognition result.

[0148] In some implementations, during the process of performing image enhancement on the environment image to obtain the corresponding target image, the second fusion module 204 executes:

[0149] Performing brightness feature extraction on the environment image to obtain a corresponding first brightness feature and performing saturation feature extraction on the environment image to obtain a corresponding first saturation feature;

[0150] Performing data enhancement on the first brightness feature to obtain a second brightness feature, and performing correction processing on the second brightness feature using a gamma correction algorithm to obtain a third brightness feature;

[0151] Determine a target window, and obtain a window brightness feature corresponding to the third brightness feature under the target window;

[0152] Performing regional histogram equalization processing on the window brightness feature to obtain a corresponding target brightness feature, and recombining according to the target brightness feature to obtain a fourth brightness feature;

[0153] Performing bilateral filtering on the first saturation feature to obtain a second saturation feature, and performing stretching transformation on the second saturation feature to obtain a third saturation feature;

[0154] The corresponding target image is obtained according to the fourth brightness feature and the third saturation feature.

[0155] In some implementations, during the process of performing edge recognition on the target image to obtain target edge information corresponding to the target image, the second fusion module 204 executes:

[0156] Performing wavelet transform on the target image to obtain first transform results of the target image in the horizontal direction and the vertical direction;

[0157] Performing a wavelet transform on the target image to obtain a second transform result of the target image in the horizontal direction, and performing a wavelet transform on the target image to obtain a third transform result of the target image in the vertical direction;

[0158] determining a target gradient corresponding to the target image according to the first transformation result, the second transformation result, and the third transformation result;

[0159] Determining a segmentation threshold corresponding to the target image according to the target gradient, and determining a target edge point corresponding to the target image according to the segmentation threshold;

[0160] Data clustering is performed on the target edge points to obtain target edge information corresponding to the target image.

[0161] In some implementations, during the process of performing data clustering on the target edge points to obtain target edge information corresponding to the target image, the second fusion module 204 executes:

[0162] Determine an initial cluster center, and obtain distance information between the target edge point and the initial cluster center;

[0163] Determining the target probability that the target edge point belongs to the target corresponding to the initial cluster center according to the distance information;

[0164] Calculating a first degree of membership of the target edge point to the initial cluster center, and determining a clustering target value corresponding to the initial cluster center according to the distance information and the first degree of membership;

[0165] Update the first membership degree according to the clustering target value and the target probability to obtain a second membership degree;

[0166] Update the initial cluster center according to the second membership degree to obtain a target cluster center;

[0167] Performing data clustering on the target edge points according to the target cluster center to obtain the target edge information corresponding to the target image;

[0168] The clustering target value and the second degree of membership are obtained according to the following formula:

[0169]

[0170]

[0171] in, represents the clustering target value, cer represents the number of the initial cluster centers, and m represents the number of target edge points. represents the segmentation threshold, represents the pixel value corresponding to the jth target edge point, represents the number of pixel values corresponding to the j-th target edge point, Indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the pixel value corresponding to the i-th initial cluster center; The second degree of membership indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the probability that the j-th target edge point belongs to the target corresponding to the i-th initial cluster center, Indicates the minimum value of the target probability that the j-th target edge point belongs to the i-th initial cluster center.

[0172] In some implementations, during the process of obtaining target position information corresponding to the target object in the initial tower crane three-dimensional model according to Beidou positioning, the model adjustment module 205 executes:

[0173] Obtaining a target hanging object corresponding to the target tower crane from the initial three-dimensional tower crane model;

[0174] Obtaining the hanging object position information corresponding to the target hanging object by using the Beidou positioning;

[0175] The position of the target object in the initial three-dimensional tower crane model is calculated according to the hanging object position information to obtain the target position information corresponding to the target object.

[0176] In some embodiments, the target tower crane three-dimensional model includes a first tower crane three-dimensional model corresponding to a first moment and a second tower crane three-dimensional model corresponding to a second moment adjacent to the first moment. In the process of performing real-time operating state identification based on the target tower crane three-dimensional model to obtain the target operating state corresponding to the target tower crane, the state recognition module 206 executes:

[0177] Obtain first feature information corresponding to the target hanging object from the first three-dimensional model of the tower crane and obtain second feature information corresponding to the target hanging object from the second three-dimensional model of the tower crane;

[0178] Determine a change angle corresponding to each associated position of the target hanging object according to the first characteristic information and the second characteristic information; and determine a maximum movement angle corresponding to the target hanging object according to the change angle;

[0179] Determine the angle difference corresponding to the associated position in the target hanging object by performing a difference calculation based on the change angle and the maximum movement angle;

[0180] Determining a change weight corresponding to the associated position according to the change angle and the angle difference, and determining a target change value corresponding to the target tower crane according to the change weight, the first feature information, and the second feature information;

[0181] The target operating state corresponding to the target tower crane is determined according to the target change value.

[0182] In some embodiments, the intelligent tower crane remote operation system 200 based on vision and lidar fusion recognition can be applied to terminal equipment.

[0183] It should be noted that, those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the intelligent tower crane remote operation system 200 based on vision and lidar fusion recognition described above can refer to the corresponding process in the aforementioned embodiment of the intelligent tower crane remote operation method based on vision and lidar fusion recognition, and will not be repeated here.

[0184] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the XXXX methods provided in the description of the embodiment of the present invention.

[0185] The storage medium may be an internal storage unit of the terminal device described in the aforementioned embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the terminal device.

[0186] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units in the methods, systems, and devices disclosed above may be implemented as software, firmware, hardware, or any combination thereof. In hardware embodiments, the division between functional modules / units described above does not necessarily correspond to the division between physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage media encompasses both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0187] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further limitations, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.

[0188] The serial numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope of protection of the claims.

Claims

1. An intelligent tower crane remote operation method based on vision and laser radar fusion recognition, characterized in that: The method comprises: Using a first laser radar at a first preset position to collect first point cloud data corresponding to the surrounding environment of the target tower crane, and using a second laser radar at a second preset position to collect second point cloud data corresponding to the surrounding environment of the target tower crane; Using a third laser radar at a third preset position to collect third point cloud data corresponding to the surrounding environment of the target tower crane, and using a visual camera at the third preset position to collect an environmental image corresponding to the surrounding environment of the target tower crane; Performing data fusion on the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data; Fusing the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane; Obtaining target position information corresponding to a target object in the initial three-dimensional tower crane model according to Beidou positioning, and adding the target position information to the initial three-dimensional tower crane model to obtain a target three-dimensional tower crane model; Performing real-time operating state recognition based on the target tower crane three-dimensional model to obtain a target operating state corresponding to the target tower crane; Sending the target tower crane three-dimensional model and the target operating status to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and the target operating status to a target user; Obtaining a control instruction sent by the terminal device, and controlling the target tower crane to perform a tower crane operation according to the control instruction; The fusing of the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane includes: Performing image enhancement on the environment image to obtain a corresponding target image; Performing edge recognition on the target image to obtain target edge information corresponding to the target image, and performing target recognition on the target edge information according to a convolutional neural network to obtain a first recognition result corresponding to the target image; The performing edge recognition on the target image to obtain target edge information corresponding to the target image includes: Performing wavelet transform on the target image to obtain first transform results of the target image in the horizontal direction and the vertical direction; Performing a wavelet transform on the target image to obtain a second transform result of the target image in the horizontal direction, and performing a wavelet transform on the target image to obtain a third transform result of the target image in the vertical direction; determining a target gradient corresponding to the target image according to the first transformation result, the second transformation result, and the third transformation result; Determining a segmentation threshold corresponding to the target image according to the target gradient, and determining a target edge point corresponding to the target image according to the segmentation threshold; Performing data clustering on the target edge points to obtain target edge information corresponding to the target image; The performing data clustering on the target edge points to obtain target edge information corresponding to the target image includes: Determine an initial cluster center, and obtain distance information between the target edge point and the initial cluster center; Determining the target probability that the target edge point belongs to the target corresponding to the initial cluster center according to the distance information; Calculating a first degree of membership of the target edge point to the initial cluster center, and determining a clustering target value corresponding to the initial cluster center according to the distance information and the first degree of membership; Update the first membership degree according to the clustering target value and the target probability to obtain a second membership degree; Update the initial cluster center according to the second membership degree to obtain a target cluster center; Performing data clustering on the target edge points according to the target cluster center to obtain the target edge information corresponding to the target image; The clustering target value and the second degree of membership are obtained according to the following formula: ; ; in, represents the clustering target value, cer represents the number of the initial cluster centers, and m represents the number of target edge points. represents the segmentation threshold, represents the pixel value corresponding to the jth target edge point, represents the number of pixel values corresponding to the j-th target edge point, Indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the pixel value corresponding to the i-th initial cluster center; The second degree of membership indicates that the j-th target edge point belongs to the i-th initial cluster center. represents the probability that the j-th target edge point belongs to the target corresponding to the i-th initial cluster center, Indicates the minimum value of the target probability that the j-th target edge point belongs to the i-th initial cluster center.

2. The method according to claim 1, characterized in that The step of fusing the first point cloud data, the second point cloud data, and the third point cloud data to obtain target point cloud data includes: Performing data clustering on the first point cloud data to obtain a first clustering result and a first target type corresponding to the first clustering result; performing data clustering on the second point cloud data to obtain a second clustering result and a second target type corresponding to the second clustering result; performing data clustering on the third point cloud data to obtain a third clustering result and a third target type corresponding to the third clustering result; Performing target recognition on the environment image using a target recognition model to obtain a corresponding target recognition type; Filtering the first clustering result using the first target type according to the target recognition type to obtain a first point cloud cluster; Filtering the second clustering result using the second target type according to the target recognition type to obtain a second point cloud cluster; Filtering the third clustering result using the third target type according to the target recognition type to obtain a third point cloud cluster; Performing data alignment on the first point cloud cluster and the third point cloud cluster to obtain a first alignment result, and performing data alignment on the second point cloud cluster and the third point cloud cluster to obtain a second alignment result; The target point cloud data is obtained by performing data fusion on the first point cloud data, the second point cloud data, and the third point cloud data according to the first alignment result and the second alignment result.

3. The method according to claim 2, characterized in that The performing target recognition on the environment image using the target recognition model to obtain a corresponding target recognition type includes: Performing a preprocessing operation on the environment image using the input layer of the target recognition model to obtain a preprocessed image; Using the feature extraction layer of the target recognition model to perform feature extraction on the preprocessed image to obtain initial feature information; Utilizing the variable convolutional layer of the target recognition model to add a learnable offset and an adjustment factor at the sampling position to adaptively perform feature sampling on the initial feature information to obtain target feature information; The target recognition layer of the target recognition model is used to perform target recognition on the target feature information to obtain the target recognition type.

4. The method according to claim 1, wherein The fusing of the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane includes: Performing target recognition on the target point cloud data to obtain a second recognition result corresponding to the target point cloud data; The initial tower crane three-dimensional model is obtained by performing data fusion on the environment image and the target point cloud data according to the first recognition result and the second recognition result.

5. The method according to claim 4, characterized in that The performing image enhancement on the environment image to obtain a corresponding target image includes: Performing brightness feature extraction on the environment image to obtain a corresponding first brightness feature and performing saturation feature extraction on the environment image to obtain a corresponding first saturation feature; Performing data enhancement on the first brightness feature to obtain a second brightness feature, and performing correction processing on the second brightness feature using a gamma correction algorithm to obtain a third brightness feature; Determine a target window, and obtain a window brightness feature corresponding to the third brightness feature under the target window; Performing regional histogram equalization processing on the window brightness feature to obtain a corresponding target brightness feature, and recombining according to the target brightness feature to obtain a fourth brightness feature; Performing bilateral filtering on the first saturation feature to obtain a second saturation feature, and performing stretching transformation on the second saturation feature to obtain a third saturation feature; The corresponding target image is obtained according to the fourth brightness feature and the third saturation feature.

6. The method according to claim 1, characterized in that The obtaining target position information corresponding to the target object in the initial tower crane three-dimensional model according to Beidou positioning includes: Obtaining a target hanging object corresponding to the target tower crane from the initial three-dimensional tower crane model; Obtaining the hanging object position information corresponding to the target hanging object by using the Beidou positioning; The position of the target object in the initial three-dimensional tower crane model is calculated according to the hanging object position information to obtain the target position information corresponding to the target object.

7. The method according to claim 6, characterized in that The target tower crane three-dimensional model includes a first tower crane three-dimensional model corresponding to a first moment and a second tower crane three-dimensional model corresponding to a second moment adjacent to the first moment, and performing real-time operation state identification based on the target tower crane three-dimensional model to obtain a target operation state corresponding to the target tower crane includes: Obtain first feature information corresponding to the target hanging object from the first three-dimensional model of the tower crane and obtain second feature information corresponding to the target hanging object from the second three-dimensional model of the tower crane; Determine a change angle corresponding to each associated position of the target hanging object according to the first characteristic information and the second characteristic information; and determine a maximum movement angle corresponding to the target hanging object according to the change angle; Determine the angle difference corresponding to the associated position in the target hanging object by performing a difference calculation based on the change angle and the maximum movement angle; Determining a change weight corresponding to the associated position according to the change angle and the angle difference, and determining a target change value corresponding to the target tower crane according to the change weight, the first feature information, and the second feature information; The target operating state corresponding to the target tower crane is determined according to the target change value.

8. An intelligent tower crane remote operation system based on vision and laser radar fusion recognition, applied to the intelligent tower crane remote operation method based on vision and laser radar fusion recognition as described in any one of claims 1 to 7, characterized in that: include: A first acquisition module is configured to acquire first point cloud data corresponding to the surrounding environment of a target tower crane using a first laser radar at a first preset position, and to acquire second point cloud data corresponding to the surrounding environment of the target tower crane using a second laser radar at a second preset position; A second acquisition module is configured to acquire third point cloud data corresponding to the surrounding environment of the target tower crane using a third laser radar at a third preset position and to acquire an environmental image corresponding to the surrounding environment of the target tower crane using a visual camera at the third preset position; A first fusion module is used to perform data fusion on the first point cloud data, the second point cloud data and the third point cloud data to obtain target point cloud data; A second fusion module is used to fuse the environment image and the target point cloud data to obtain an initial three-dimensional tower crane model corresponding to the surrounding environment of the target tower crane; A model adjustment module is used to obtain target position information corresponding to the target object in the initial tower crane three-dimensional model according to Beidou positioning, and add the target position information to the initial tower crane three-dimensional model to obtain a target tower crane three-dimensional model; A state recognition module is used to perform real-time operation state recognition based on the three-dimensional model of the target tower crane to obtain a target operation state corresponding to the target tower crane; a data sending module, configured to send the target tower crane three-dimensional model and the target operating status to a terminal device that is communicatively connected to the target tower crane, so that the terminal device displays the target tower crane three-dimensional model and the target operating status to a target user; The instruction execution module is used to obtain the control instruction sent by the terminal device and control the target tower crane to perform tower crane operations according to the control instruction.

Citation Information

Patent Citations

  • Tower crane hoisting object identification and collision information measurement system and method

    CN113860178A

  • Tower crane control method and system based on camera and laser radar

    CN116395567A