A tower crane object segmentation method and system based on visual and laser fusion based on prompt engineering

By combining the prompt engineering method of visual and laser fusion on the tower crane, the tower crane objects are segmented using the visual large model of supporting points and frames, the problems of missegment and real-time segmentation in the existing technology are solved, and the high accuracy and real-time segmentation effect of hanging objects is achieved.

CN118429628BActive Publication Date: 2025-05-23TALESUN TECH BUILDING INTELLIGENCE (SHENZHEN) CO LTD +2
2 Cites -1 Cited by

Patent Information

Application Number
CN202311662487.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-05-23
Estimated Expiration
2043-12-06

AI Technical Summary

Technical Problem

The existing methods of tower cranes are prone to missegment, which is not conducive to real-time segmentation, especially when sparse lidar point clouds and rich background information exist.

Method used

The visual and laser fusion method based on prompt engineering is adopted. By installing a laser radar and image acquisition device under the tower crane truck, combining the Meta DINOv2 model for training and tuning, a visual big model supporting points and frames is generated, and the point cloud data is clustered and projected, and the image data is combined for semantic segmentation, and the clustering points that are not in line with are eliminated to achieve accurate segmentation of the hanging objects.

Benefits of technology

Real-time and accurate segmentation of tower crane objects is realized, the probability of missegment is reduced, the advantages of lidar and image data are fully utilized, and the accuracy and real-timeness of segmentation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118429628B_ABST
    Figure CN118429628B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for segmenting a tower crane suspended object by visual and laser fusion based on prompt engineering, comprising the following steps: collecting laser radar point cloud data of the suspended object by laser radar, and collecting image data of the suspended object by image acquisition device; using open source data set Meta 11B to train Meta DINOv2 model, and tuning Meta DINOv2 model to obtain a visual large model of support points and frames; extracting point cloud data for generating prompt information from the laser radar point cloud data, and clustering and projecting the point cloud data to obtain point form prompt information and frame form prompt information; based on the point form prompt information, frame form prompt information, image data, and using the visual large model of support points and frames to obtain image segmentation results, according to the image segmentation results, eliminating cluster points that do not meet the requirements, and completing the segmentation of the suspended object. The present invention achieves the purpose of being able to segment suspended objects in real time and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of suspended object segmentation, and in particular to a tower crane suspended object segmentation method and system based on vision and laser fusion of prompt engineering. Background Art

[0002] In the intelligent driving of tower cranes, the accurate segmentation of the crane's hoisted objects is an important step that affects collision detection and motion planning. In the existing methods, the laser radar and camera are mainly installed under the tower crane trolley, and the viewports of the laser radar and camera are vertically downward to scan the hoisted objects in real time.

[0003] The segmentation of suspended objects based on LiDAR point clouds mainly depends on the degree of spatial aggregation of point clouds, and uses 3D convolutional neural networks to complete target segmentation. Image-based segmentation is mainly based on various 2D semantic segmentation networks. However, the LiDAR point clouds of tower cranes are relatively sparse, and it is very easy to mis-segment when there are interfering objects nearby. Especially for long objects, the Euclidean distance of the point cloud of the same suspended object may even be smaller than the distance between the target and other obstacles, and the possibility of mis-segmentation is greater. Because images have rich texture information, they can better distinguish the semantic boundaries of objects, but due to the lack of guidance from depth information, semantic segmentation is easily interfered by the ground background.

[0004] The Transformer-based visual model has a strong zero-sample migration capability and can perform complete and reliable object semantic segmentation of all objects in the image. However, the semantic segmentation based on the visual model is computationally intensive, which is not conducive to the realization of real-time segmentation of suspended objects. Summary of the invention

[0005] In order to overcome the shortcomings of the prior art, the present invention provides a tower crane hanging object segmentation method and system based on vision and laser fusion of prompt engineering, which is used to solve the technical problems that the existing hanging object segmentation method is prone to mis-segmentation and is not conducive to real-time segmentation, thereby achieving the purpose of being able to segment the hanging objects in real time and accurately.

[0006] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0007] A tower crane object segmentation method based on vision and laser fusion based on prompt engineering, comprising the following steps:

[0008] Install the laser radar and image acquisition device under the tower crane trolley, and make the viewport direction perpendicular to the ground;

[0009] Collecting laser radar point cloud data of the hanging object by the laser radar, and collecting image data of the hanging object by the image acquisition device;

[0010] The Meta DINOv2 model is trained using the open source dataset Meta 11B, and the Meta DINOv2 model is tuned to obtain a large visual model that supports points and boxes;

[0011] Extracting point cloud data for generating prompt information from the laser radar point cloud data, and clustering and projecting the point cloud data to obtain point-form prompt information and box-form prompt information;

[0012] Based on the point-form prompt information, the frame-form prompt information, and the image data, and using the support point and the large visual model of the frame, an image segmentation result is obtained, and non-conforming cluster points are eliminated according to the image segmentation result to complete the segmentation of the hanging object.

[0013] As a preferred implementation mode of the present invention, when the Meta DINOv2 model is tuned, it includes:

[0014] Acquire tower crane ground scene data, and send the tower crane ground scene data to the Meta DINO v2 model for Finetuning, and retain the prompt engineering capability of the Meta DINO v2 model;

[0015] Determine whether the tuned Meta DINOv2 model is compatible with prompt points and prompt boxes;

[0016] If yes, the tuning of the Meta DINOv2 model is completed;

[0017] If not, continue to input the tower crane ground scene data to perform finetuning on the Meta DINOv2 model.

[0018] As a preferred embodiment of the present invention, when judging whether the tuned Meta DINOv2 model is compatible with the prompt point and the prompt box, it includes:

[0019] Given a cue point and a cue box for testing, determine whether the tuned Meta DINOv2 model can obtain the features in the cue box;

[0020] If not, continue to Finetuning the Meta DINOv2 model;

[0021] If yes, then executing the prompt point compatibility determination step;

[0022] The step of determining whether the prompt point is compatible includes:

[0023] Determine whether the tuned Meta DINOv2 model can obtain the features within the box area formed by all the cue points;

[0024] If not, continue to Finetuning the Meta DINOv2 model;

[0025] If so, the similar area search step is performed.

[0026] As a preferred embodiment of the present invention, when executing the similar area search step, it includes:

[0027] Determine whether the tuned Meta DINOv2 model can search for regions that are similar to all features adjacent to the features it has acquired, and use the regions as segmentation results;

[0028] If yes, the tuning of the Meta DINOv2 model is completed;

[0029] If not, continue to Finetuning the Meta DINOv2 model.

[0030] As a preferred embodiment of the present invention, when extracting point cloud data for generating prompt information, it includes:

[0031] According to the operation parameters of the tower crane, determine the area near the height of the suspended object under the tower crane trolley, and extract point cloud data of the area near the height of the suspended object from the laser radar point cloud data;

[0032] Among them, the operating parameters include: the position of the tower crane trolley, the hook height, the boom angle, and the area adjacent to the height of the suspended object is an area where the height difference with the suspended object is less than a preset threshold.

[0033] As a preferred implementation manner of the present invention, when clustering the point cloud data, it includes:

[0034] Performing Euclidean distance clustering on the point cloud data to obtain a point cloud clustering result;

[0035] Obtaining the horizontal field of view center of the laser radar, and traversing the point cloud clustering results to determine whether the distance between each point cloud clustering result and the horizontal field of view center is less than a threshold;

[0036] The point cloud clustering results with a distance less than the threshold are retained, and the point cloud clustering results with a distance greater than or equal to the threshold are eliminated to obtain the final point cloud clustering results.

[0037] As a preferred embodiment of the present invention, when projecting the point cloud data, it includes:

[0038] Determine a two-dimensional plane for projection, and project the final point cloud clustering result onto the two-dimensional plane to obtain a plurality of projection points;

[0039] Using the plurality of projection points as the point-form prompt information;

[0040] Obtaining a minimum area circumscribed rectangle of the plurality of projection points as the frame-form prompt information;

[0041] The two-dimensional plane is perpendicular to the main optical axis of the image acquisition device, and the projection result is in the form of 0 and 1 without accumulation.

[0042] As a preferred embodiment of the present invention, when obtaining the image segmentation result, it includes:

[0043] Sending the point-form prompt information, the frame-form prompt information and the image data into the visual macro model of the support point and frame;

[0044] The features in the frame-form prompt information are obtained through the visual macro model of the support points and the frame, and the features in the frame area formed by the point-form prompt information are obtained:

[0045] And search for regions similar to all features adjacent to the feature to obtain the image segmentation result.

[0046] As a preferred implementation manner of the present invention, when eliminating non-conforming cluster points according to the image segmentation result, it includes:

[0047] According to the image segmentation result and the point cloud projection result, and in accordance with the pixel meaning correspondence, an imaging frustum is established;

[0048] The clustering points within the imaging cone are retained, and the clustering points outside the imaging cone are removed.

[0049] A tower crane object segmentation system based on vision and laser fusion based on prompt engineering, including:

[0050] Visual large model building unit: use the open source data set Meta 11B to train the Meta DINOv2 model, and tune the Meta DINOv2 model to obtain a visual large model that supports points and boxes;

[0051] A prompt information acquisition unit is used to extract point cloud data used to generate prompt information from the laser radar point cloud data, and cluster and project the point cloud data to obtain point-form prompt information and box-form prompt information;

[0052] Segmentation unit: based on the point prompt information, the frame prompt information, and the image data, and using the visual macro model of the support points and the frame, obtains an image segmentation result, removes cluster points that do not conform to the image segmentation result, and completes the segmentation of the hanging object;

[0053] Among them, the laser radar and the image acquisition device are installed under the tower crane trolley, and the viewport direction is perpendicular to the ground; the laser radar point cloud data of the suspended object is collected by the laser radar, and the image data of the suspended object is collected by the image acquisition device.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) The present invention optimizes the Meta DINOv2 model until the model is compatible with the prompt points and prompt boxes and retains the prompt engineering capability, thereby obtaining a visual macro model that supports points and boxes; in addition, the point cloud data is projected to obtain the corresponding point cloud projection results, and the image segmentation results are obtained through the visual macro model that supports points and boxes, and the point cloud projection results and the image segmentation results are combined to eliminate the cluster points that do not meet the requirements, thereby achieving accurate segmentation of the hanging objects and greatly reducing the probability of mis-segmentation;

[0056] (2) The visual big model supporting points and frames provided by the present invention can be combined with prompt engineering to perform customized image semantic segmentation; the model fully converts the lidar data into visual prompts in the form of points and surfaces, and completes semantic segmentation under the prompts, thereby making full use of lidar and images to carry out semantic segmentation of suspended objects with the support of the visual big model.

[0057] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a step diagram of the tower crane object segmentation method based on vision and laser fusion based on prompt engineering provided by the present invention;

[0059] Figure 2 is a schematic diagram of the deployment positions of the laser radar and camera of this embodiment;

[0060] Figure 3 It is a calculation flow chart of the segmentation method of this embodiment.

[0061] Explanation of the figure numbers: 1. Tower crane trolley; 2. Laser radar; 3. Camera; 4. Hook; 5. Boom; 6. Hoisted object; 7. Imaging cone. DETAILED DESCRIPTION

[0062] The method for segmenting tower crane objects by combining vision and laser fusion based on prompt engineering provided by the present invention is as follows: Figure 1 As shown, the following steps are included:

[0063] Step S1: Install the laser radar and image acquisition device under the tower crane trolley, and make the viewport direction perpendicular to the ground;

[0064] Step S2: collecting laser radar point cloud data of the hanging object by using a laser radar, and collecting image data of the hanging object by using an image acquisition device;

[0065] Step S3: Use the open source dataset Meta 11B to train the Meta DINOv2 model, and tune the MetaDINOv2 model to obtain a large visual model that supports points and boxes;

[0066] Step S4: extracting point cloud data for generating prompt information from the laser radar point cloud data, and clustering and projecting the point cloud data to obtain point-form prompt information and box-form prompt information;

[0067] Step S5: Based on the prompt information in the form of points, the prompt information in the form of frames, and the image data, and using the visual macro model of the support points and frames, the image segmentation result is obtained, and the cluster points that do not meet the requirements are eliminated according to the image segmentation result to complete the segmentation of the hanging object.

[0068] Specifically, the image acquisition device includes a camera.

[0069] In the above step S3, when tuning the Meta DINOv2 model, it includes:

[0070] Obtain the tower crane ground scene data, and send the tower crane ground scene data to the Meta DINOv2 model for Finetuning, while retaining the Meta DINOv2 model's prompt engineering capabilities;

[0071] Determine whether the tuned Meta DINOv2 model is compatible with prompt points and prompt boxes;

[0072] If yes, the tuning of Meta DINOv2 model is completed;

[0073] If not, continue to input the tower crane ground scene data to finetune the Meta DINOv2 model.

[0074] Furthermore, when judging whether the tuned Meta DINOv2 model is compatible with the prompt points and prompt boxes, it includes:

[0075] Given the cue points and cue boxes used for testing, determine whether the tuned Meta DINOv2 model can obtain the features in the cue boxes;

[0076] If not, continue Finetuning the Meta DINOv2 model;

[0077] If yes, then executing the prompt point compatibility determination step;

[0078] The prompt point compatibility determination step includes:

[0079] Determine whether the tuned Meta DINOv2 model can obtain the features within the box area formed by all the cue points;

[0080] If not, continue Finetuning the Meta DINOv2 model;

[0081] If so, the similar area search step is performed.

[0082] Furthermore, when performing the similar area search step, it includes:

[0083] Determine whether the tuned Meta DINOv2 model can search for regions that are similar to all features adjacent to the features it has acquired, and use the regions as segmentation results;

[0084] If yes, the tuning of Meta DINOv2 model is completed;

[0085] If not, continue Finetuning the Meta DINOv2 model.

[0086] Specifically, the actual data FinetuningDINOv2 in tower crane operation is used, which is compatible with point and box prompt information. That is, given a point and a box, it is possible to calculate the features within the box or the box area formed by all prompt points. That is, the image encoder of the tuned DINOv2 model is used to calculate the features in the area, using the ViT architecture, calculating the 14x14 feature embedding, and searching for all adjacent areas with similar features, and using these areas as segmentation results.

[0087] In the above step S4, when extracting the point cloud data for generating prompt information, it includes:

[0088] According to the operating parameters of the tower crane, the area near the height of the suspended object under the tower crane trolley is determined, and the point cloud data of the area near the height of the suspended object is extracted from the laser radar point cloud data;

[0089] Among them, the operating parameters include: tower crane trolley position, hook height, boom angle, and the area adjacent to the height of the hoisted object is an area where the height difference with the hoisted object is less than a preset threshold.

[0090] In the above step S4, when clustering the point cloud data, it includes:

[0091] Perform Euclidean distance clustering on the point cloud data to obtain the point cloud clustering results;

[0092] Get the center of the horizontal field of view of the laser radar, and traverse the point cloud clustering results to determine whether the distance between each point cloud clustering result and the center of the horizontal field of view is less than the threshold;

[0093] The point cloud clustering results with a distance less than the threshold are retained, and the point cloud clustering results with a distance greater than or equal to the threshold are eliminated to obtain the final point cloud clustering results.

[0094] In the above step S4, when projecting the point cloud data, it includes:

[0095] Determine a two-dimensional plane for projection, and project the final point cloud clustering result onto the two-dimensional plane to obtain a number of projection points;

[0096] Use several projection points as point-form prompt information;

[0097] Get the minimum area circumscribed rectangle of several projection points as a box-shaped prompt information;

[0098] The two-dimensional plane is perpendicular to the main optical axis of the image acquisition device, and the projection result is in the form of 0 and 1 without accumulation.

[0099] In the above step S5, when the image segmentation result is obtained, it includes:

[0100] Sending the point-form prompt information, the frame-form prompt information and the image data into the visual macro model supporting the points and frames;

[0101] The features within the box-shaped prompt information are obtained through the visual model that supports points and boxes, and the features within the box area formed by the point-shaped prompt information are obtained:

[0102] And search for areas similar to all features adjacent to the feature to obtain the image segmentation result.

[0103] In the above step S5, when eliminating the non-compliant cluster points according to the image segmentation result, it includes:

[0104] According to the image segmentation results and point cloud projection results, and in accordance with the pixel meaning correspondence, an imaging frustum is established;

[0105] The clustering points within the imaging cone are retained, and the clustering points outside the imaging cone are removed.

[0106] The visual and laser fusion tower crane object segmentation system based on prompt engineering provided by the present invention includes: a visual large model building unit, a prompt information acquisition unit and a segmentation unit.

[0107] Visual large model building unit: Use the open source dataset Meta 11B to train the Meta DINOv2 model, and tune the Meta DINOv2 model to obtain a visual large model that supports points and boxes.

[0108] Prompt information acquisition unit: used to extract point cloud data used to generate prompt information from the laser radar point cloud data, and cluster and project the point cloud data to obtain point-form prompt information and box-form prompt information.

[0109] Segmentation unit: Based on the prompt information in the form of points, the prompt information in the form of frames, and the image data, and using the visual large model of the support points and frames, the image segmentation results are obtained, and the cluster points that do not meet the requirements are eliminated according to the image segmentation results to complete the segmentation of the hanging objects.

[0110] Among them, the laser radar and image acquisition device are installed under the tower crane trolley, and the viewport direction is perpendicular to the ground; the laser radar point cloud data of the hoisted object is collected by the laser radar, and the image data of the hoisted object is collected by the image acquisition device.

[0111] The following examples are provided to further illustrate the present invention, but the scope of the present invention is not limited thereto.

[0112] Figure 2 is a schematic diagram of the deployment positions of the laser radar and camera of this embodiment, Figure 3 is a calculation flow chart of the method of this embodiment. Figure 2 and Figure 3 , the specific process of the segmentation method provided in the embodiment is described:

[0113] In step 401, the open source Meta DINOv2 model is first tuned, and some tower crane ground scene data is input for finetuning, and the prompt engineering capability of the original network is retained to obtain a large visual model of support points and boxes.

[0114] In step 402 , the tower crane operation information is read to obtain the current tower crane trolley position, hook height, and boom angle, and then the process goes to step 403 .

[0115] In step 403 , the point cloud within a height of 0.5-3 meters below the hook height is clustered using Euclidean distance clustering, and the clustering result closest to the center of the horizontal field of view of the laser radar is finally retained, and then the process goes to step 404 .

[0116] In step 404 , the clustering result is projected onto a horizontal plane, and the process proceeds to step 405 .

[0117] In step 405, the minimum area circumscribed rectangle of the projection point on the horizontal plane is calculated as prompt information in the form of a frame, and the projection point is used as prompt information in the form of a point, and the process goes to step 406.

[0118] In step 406 , the tower crane camera image is read, and based on the prompt information in the form of points and boxes, the visual large model is called for segmentation, and then the process goes to step 407 .

[0119] In step 407, an imaging cone is calculated according to the image segmentation result, the point cloud within the imaging cone is retained, and the point cloud outside the imaging cone is removed, thereby completing the segmentation of the suspended object.

[0120] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.

Claims

1. A tower crane object segmentation method based on visual and laser fusion based on prompt engineering, It is characterized in that The following steps are involved: Install the laser radar and image acquisition device under the tower crane trolley, and make the viewport direction perpendicular to the ground; Collecting laser radar point cloud data of the hanging object by the laser radar, and collecting image data of the hanging object by the image acquisition device; The Meta DINOv2 model is trained using the open source dataset Meta 11B, and the Meta DINOv2 model is tuned to obtain a large visual model that supports points and boxes; Extracting point cloud data for generating prompt information from the laser radar point cloud data, and clustering and projecting the point cloud data to obtain point-form prompt information and box-form prompt information; Based on the point-form prompt information, the frame-form prompt information, and the image data, and using the support point and the visual macro model of the frame, an image segmentation result is obtained, and non-conforming cluster points are removed according to the image segmentation result to complete the segmentation of the hanging object; Wherein, when clustering the point cloud data, it includes: Performing Euclidean distance clustering on the point cloud data to obtain a point cloud clustering result; Obtaining the horizontal field of view center of the laser radar, and traversing the point cloud clustering results to determine whether the distance between each point cloud clustering result and the horizontal field of view center is less than a threshold; The point cloud clustering results with a distance less than the threshold are retained, and the point cloud clustering results with a distance greater than or equal to the threshold are eliminated to obtain the final point cloud clustering results; Wherein, when projecting the point cloud data, it includes: Determine a two-dimensional plane for projection, and project the final point cloud clustering result onto the two-dimensional plane to obtain a plurality of projection points; Using the plurality of projection points as the point-form prompt information; Obtaining a minimum area circumscribed rectangle of the plurality of projection points as the frame-form prompt information; The two-dimensional plane is perpendicular to the main optical axis of the image acquisition device, and the projection result is in the form of 0 and 1 without accumulation.

2. According to the method for segmenting tower crane objects by combining vision and laser fusion based on prompt engineering in claim 1, It is characterized in that When tuning the Meta DINOv2 model, it includes: Acquire tower crane ground scene data, and send the tower crane ground scene data to the Meta DINOv2 model for Finetuning, and retain the prompt engineering capability of the Meta DINOv2 model; Determine whether the tuned Meta DINOv2 model is compatible with prompt points and prompt boxes; If yes, the tuning of the Meta DINOv2 model is completed; If not, continue to input the tower crane ground scene data to perform finetuning on the Meta DINOv2 model.

3. According to the method for segmenting tower crane objects by combining vision and laser fusion based on prompt engineering in claim 2, It is characterized in that When judging whether the tuned Meta DINOv2 model is compatible with prompt points and prompt boxes, it includes: Given the hint points and hint boxes for testing, determine whether the fine-tuned Meta DINOv2 model can obtain the features within the said hint boxes; If not, continue to perform Finetuning on the said Meta DINOv2 model; If so, execute the hint point compatibility judgment step; Among them, the said hint point compatibility judgment step includes: Judge whether the fine-tuned Meta DINOv2 model can obtain the features within the box area formed by all hint points; If not, continue to perform Finetuning on the said Meta DINOv2 model; If so, execute the similar area search step.

4. The vision and laser fusion tower crane load segmentation method based on prompt engineering according to claim 3, characterized in that, When executing the similar area search step, it includes: Judge whether the fine-tuned Meta DINOv2 model can search for an area similar to all the features adjacent to the features it obtains, and use the said area as the segmentation result; If so, complete the tuning of the said Meta DINOv2 model; If not, continue to perform Finetuning on the said Meta DINOv2 model.

5. The vision and laser fusion tower crane load segmentation method based on prompt engineering according to claim 1, characterized in that, When extracting the point cloud data for generating the prompt information, it includes: According to the operating parameters of the tower crane, determine the load height adjacent area below the tower crane trolley, and extract the point cloud data of the load height adjacent area from the lidar point cloud data; Among them, the said operating parameters include: the position of the tower crane trolley, the hook height, the boom angle, and the load height adjacent area is the area where the height difference from the load is less than the preset threshold.

6. The vision and laser fusion tower crane load segmentation method based on prompt engineering according to claim 1, characterized in that, When obtaining the image segmentation result, it includes: Send the point-form prompt information, the box-form prompt information, and the image data into the vision large model that supports points and boxes; Obtain the features within the box-form prompt information through the vision large model that supports points and boxes, and obtain the features within the box area formed by the point-form prompt information: And search for an area similar to all the features adjacent to the said features to obtain the said image segmentation result.

7. The vision and laser fusion tower crane load segmentation method based on prompt engineering according to claim 6, characterized in that, When removing the non-conforming clustering points according to the said image segmentation result, it includes: According to the said image segmentation result and the point cloud projection result, and corresponding according to the pixel meaning, establish an imaging frustum; Retain the clustering points within the said imaging frustum and remove the clustering points outside the said imaging frustum.

8. A vision and laser fusion tower crane load segmentation system based on prompt engineering, characterized in that, including: Visual large model building unit: use the open source data set Meta 11B to train the Meta DINOv2 model, and tune the Meta DINOv2 model to obtain a visual large model that supports points and boxes; A prompt information acquisition unit is used to extract point cloud data used to generate prompt information from the laser radar point cloud data, and cluster and project the point cloud data to obtain point-form prompt information and box-form prompt information; Segmentation unit: based on the point prompt information, the frame prompt information, and the image data, and using the visual macro model of the support points and the frame, obtains an image segmentation result, removes cluster points that do not conform to the image segmentation result, and completes the segmentation of the hanging object; Wherein, when clustering the point cloud data, it includes: Performing Euclidean distance clustering on the point cloud data to obtain a point cloud clustering result; Obtaining the horizontal field of view center of the laser radar, and traversing the point cloud clustering results to determine whether the distance between each point cloud clustering result and the horizontal field of view center is less than a threshold; The point cloud clustering results with a distance less than the threshold are retained, and the point cloud clustering results with a distance greater than or equal to the threshold are eliminated to obtain the final point cloud clustering results; Wherein, when projecting the point cloud data, it includes: Determine a two-dimensional plane for projection, and project the final point cloud clustering result onto the two-dimensional plane to obtain a plurality of projection points; Using the plurality of projection points as the point-form prompt information; Obtaining a minimum area circumscribed rectangle of the plurality of projection points as the frame-form prompt information; Wherein, the two-dimensional plane is perpendicular to the main optical axis of the image acquisition device, and the projection result is in the form of 0 and 1 without accumulation; Among them, the laser radar and the image acquisition device are installed under the tower crane trolley, and the viewport direction is perpendicular to the ground; the laser radar point cloud data of the suspended object is collected by the laser radar, and the image data of the suspended object is collected by the image acquisition device.

Citation Information

Patent Citations

  • A method and device for determining the rotation angle of engineering mechanical equipment

    CN109903326A

  • Efficient marking method combining laser point cloud and images

    CN109978955A