Crane working environment sensing method, computer equipment and computer readable storage medium

By installing a perception kit on the crane, obtaining and processing perception data, establishing a scenario model and identifying the target object, the problem of not being able to effectively perceive the working environment of the crane in the prior art is solved, and the safe, accurate perception and efficient operation of the crane are achieved.

CN119976640AActive Publication Date: 2025-05-13ZOOMLION HEAVY INDUSTRY SCIENCE AND TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510034896.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-13
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art cannot effectively sense the working environment of the crane, especially when lifting beyond the visual range, and cannot obtain the distance information between the hook, load and environmental obstacles in real time and accurately, resulting in low operating efficiency and high risk.

Method used

Using a perception suite, including a lidar unit, an inertial measurement unit and an imaging unit, is installed at a predetermined location of the crane. By acquiring and processing these perception data, a scene model is established and the target object is identified, and perception enhancement information is output.

Benefits of technology

It realizes the safety and accurate perception of the crane working environment, and can detect the three-dimensional spatial position, size and relative relationship information of the hook, load and surrounding environment in real time, improving the safety and efficiency of working beyond the visual range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119976640A_ABST
    Figure CN119976640A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a crane working environment sensing method, computer equipment and a computer readable storage medium. The method comprises the following steps that sensing data collected by a sensing suite are obtained, the sensing suite comprises at least one of a laser radar unit, an inertia measurement unit and a camera shooting unit, and the sensing suite is installed at the preset position of the crane; a scene model is established according to the sensing data, a target object is recognized, the scene model comprises a static obstacle and / or a dynamic obstacle, and the target object comprises a lifting hook and / or a load; and obtaining perception enhancement information between the scene model and the target object and outputting the perception enhancement information. Therefore, the working environment of the crane can be sensed, and the corresponding distance information is determined and output according to the correlation of the target object in the scene model, so that the three-dimensional space position, size and relative relation information of the lifting hook, the load and the surrounding environment can be completely detected in real time, and over-the-horizon work can be safely realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of intelligent crane technology, and in particular, relates to a crane working environment perception method, computer equipment, and computer-readable storage medium. Background Art

[0002] Cranes are key equipment for transporting goods in the construction industry and are widely used in the field of engineering construction. Their safe operation and maintenance are crucial to the construction industry. Currently, crane safety protection mainly includes monitoring of the lifting point and the surrounding area of ​​the turntable. Most of them use display screens to display the image information of the lifting point or the surrounding area of ​​the turntable.

[0003] Image information can only realize the monitoring function, but cannot provide the actual load, the distance information between the hook and the environmental obstacles. In the existing actual operation process, for example, when performing beyond-visual-range hoisting, it is currently limited by the complexity of the hoisting scene and the lack of algorithm capabilities, and cannot meet the perception needs of the crane in the real dynamic operation scene. It can only be completed through image information and relying on real-time intercom communication with auxiliary personnel. Relying on the cooperation of the driver and the signalman, the operation method is inefficient, and the communication effect depends on the expression and understanding ability of both parties. At the same time, relying on manual judgment of images itself has the technical defects of single information and high risk. It can be seen that the perception of the crane working environment urgently needs to be improved. How to safely and accurately perceive the working environment of the crane is a technical problem that needs to be solved urgently by technical personnel in this field.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the invention

[0005] Based on this, it is necessary to propose a crane working environment perception method, computer equipment and computer-readable storage medium to address the above problems, which can safely and accurately perceive the working environment of the crane.

[0006] The present application solves the technical problem by adopting the following technical solutions:

[0007] The present application provides a method for perceiving a crane working environment, comprising the following steps: acquiring perception data collected by a perception kit, the perception kit comprising at least one of a laser radar unit, an inertial measurement unit, and a camera unit, and the perception kit being installed at a predetermined position of the crane; establishing a scene model based on the perception data, and identifying a target object, the scene model comprising static obstacles and / or dynamic obstacles, and the target object comprising a hook and / or a load; acquiring and outputting perception enhancement information between the scene model and the target object.

[0008] In an optional embodiment of the present application, the perception data collected by the perception kit is obtained, including: according to the predetermined position, using the external parameter calibration method to determine the relative position relationship between the perception kits, the relative position relationship is represented by an external parameter matrix; according to the external parameter matrix, the perception data is unified into the same spatial coordinate system; and the timestamp information of the perception data is calibrated so that the perception data collected by each perception kit is located in a unified time reference system.

[0009] In an optional embodiment of the present application, the perception data collected by the perception kit is obtained, including: obtaining inertial data in the perception data, the inertial data is collected by an inertial measurement unit; calculating a motion change according to the inertial data, the motion change is used to describe the motion change of the perception kit during the collection process; obtaining initial perception data collected by the perception kit, and performing an error compensation operation on the initial perception data according to the motion change to correct errors in the initial perception data; when the initial perception data after the error compensation operation meets the convergence condition, marking the initial perception data as perception data; otherwise, repeatedly performing the error compensation operation on the initial perception data.

[0010] In an optional embodiment of the present application, a scene model is established based on the perception data, including: acquiring image data and point cloud data in the perception data, the image data is collected by a camera unit, and the point cloud data is collected by a lidar unit; performing a pose relationship transformation on the point cloud data within a specified time range to determine that all object surfaces in the scene where the crane is located are position coordinates; performing a filtering operation on the point cloud data after the pose relationship transformation, and concatenating and accumulating the point cloud data after the filtering operation to obtain first point cloud data; determining a scene model based on the first point cloud data and the image data, wherein the target object is not included in the scene model.

[0011] In an optional embodiment of the present application, the filtering operation includes: using a preset straight line detection method to process all image data within a set time period, marking the processed straight line as a wire rope of the crane; identifying a hook area at one end of the wire rope based on the image data; obtaining cargo information, and determining a load area based on the cargo information and the hook area; and filtering out point cloud data within the hook area and the load area.

[0012] In an optional embodiment of the present application, a scene model is established based on the perception data, including: constructing an initial scene model based on the first point cloud data; acquiring all image data within a preset time period, and identifying a static sub-model and a dynamic sub-model from the initial scene model according to the change process of the image data, the static sub-model being the stationary part of the initial scene model, including a background area and a first obstacle area, and the dynamic sub-model being the moving part of the initial scene model, including a hook area, a load area, and a second obstacle area; identifying static obstacles from the static sub-model; deleting the hook area and the load area from the dynamic sub-model, and identifying dynamic obstacles from the deleted second obstacle area; and summarizing the static obstacles and / or dynamic obstacles to obtain a scene model.

[0013] In an optional embodiment of the present application, the target object is identified according to the perception data, including: acquiring image data and point cloud data in the perception data, the image data is acquired by the camera unit, and the point cloud data is acquired by the laser radar unit; entering an initialization phase according to the image data and the point cloud data, including performing a preprocessing operation on the point cloud data to obtain second point cloud data, the preprocessing operation including at least one of coordinate change, multi-frame accumulation, filtering and denoising, ground segmentation, and region of interest filtering; processing the image data by an image semantic segmentation method to determine at least one target object image mask; and using the target object image mask corresponding to the image mask according to the external parameter relationship between the camera unit and the laser radar unit. The method comprises the following steps: extracting a first object point cloud from the second point cloud data of the scene model; determining at least one target object from the first object point cloud after accumulating multiple frames of the first object point cloud; obtaining and outputting perception enhancement information between the scene model and the target object, including: entering a dynamic tracking phase according to the first object point cloud, including accumulating multiple frames of the first object point cloud for a period of time, and determining a pixel area of ​​the target object according to dynamic tracking of the image data; obtaining point cloud data within the pixel area of ​​the target object and marking it as a second object point cloud; registering the first object point cloud with the second object point cloud, replacing the second object point cloud with the registered first object point cloud to obtain a third object point cloud; and obtaining perception enhancement information according to the third point cloud.

[0014] In an optional embodiment of the present application, the perception enhancement information between the scene model and the target object is obtained and output, including: obtaining the relative distance in the perception enhancement information, and the first warning threshold and the second warning threshold corresponding to the relative distance, the relative distance is the distance between the target object and the static obstacle and / or the dynamic obstacle; if the relative distance is less than the first warning threshold and greater than the second warning threshold, then outputting a first warning message, the first warning message is used to prompt the user that there is a risk of collision; if the relative distance is less than the second warning threshold, then outputting a second warning message and / or controlling the crane to stop moving, the second warning message is used to prompt the user that the load will collide. The present application also provides a computer device, including a processor and a memory: the processor is used to execute a computer program stored in the memory to implement the above method.

[0015] The present application also provides a computer-readable storage medium storing a computer program, which implements the aforementioned method when the computer program is executed by a processor.

[0016] The embodiments of the present application have the following beneficial effects:

[0017] The present application can obtain various sensing data through the sensing kit installed at the predetermined position of the crane to construct the scene model of the crane, so as to identify the target objects including the hook, load and / or obstacles. Thus, the working environment of the crane is sensed, and the corresponding distance information is determined and outputted through the correlation of the target objects in the scene model, so that the three-dimensional spatial position, size and relative relationship information of the hook, load and surrounding environment can be detected in real time and completely, so as to facilitate the safe implementation of beyond-visual-range work.

[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented according to the contents of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with the accompanying drawings. It should be understood that the above general description and the detailed description below are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] in:

[0021] Figure 1A schematic diagram of a flow chart of a crane working environment perception method provided by an embodiment;

[0022] Figure 2a A schematic diagram of original point cloud data provided by an embodiment;

[0023] Figure 2b A schematic diagram of point cloud data after filtering cranes provided by an embodiment;

[0024] Figure 2c A schematic diagram of point cloud data after downsampling and filtering provided by an embodiment;

[0025] Figure 2d A schematic diagram of point cloud data after filtering the ground provided by an embodiment;

[0026] Figure 2e A schematic diagram of identifying a hook area provided by an embodiment;

[0027] Figure 2f A schematic diagram of point cloud data of an obstacle provided by an embodiment;

[0028] Figure 2g A schematic diagram of identifying an obstacle provided by an embodiment;

[0029] Figure 3a A first schematic diagram for determining a load area range provided by an embodiment;

[0030] Figure 3b A second schematic diagram for determining a load area range provided in an embodiment;

[0031] Figure 4a A first schematic diagram of determining a load area mask provided by an embodiment;

[0032] Figure 4b A second schematic diagram of determining a load area mask provided by an embodiment;

[0033] Figure 5 A schematic diagram of the effect of obtaining a first object point cloud provided by an embodiment;

[0034] Figure 6 A schematic diagram of a target object and a scene model provided by an embodiment;

[0035] Figure 7 A schematic diagram of extracting load from a third point cloud in real time based on an image mask provided by an embodiment;

[0036] Figure 8 A schematic diagram of real-time matching of a sparse load point cloud to a third point cloud provided by an embodiment;

[0037] Fig. 9A schematic diagram of a first output method of distance information provided by an embodiment;

[0038] Fig.10 A schematic diagram of a second output mode of distance information provided by an embodiment;

[0039] Fig.11 A schematic block diagram of the structure of a computer device provided by an embodiment. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] The existing technology of lifting point monitoring cannot provide actual horizontal and vertical distance information. When lifting beyond visual range, it can only be judged manually through image information, which has single information and high risk; or it can communicate with auxiliary personnel in real time, and the communication effect depends on the expression and understanding ability of both parties. Unmanned lifting technology is the development trend of the industry. As the basic core technology of unmanned lifting technology, intelligent perception technology is currently limited by the complexity of lifting scenes and insufficient algorithm capabilities, and cannot meet the perception needs of cranes in real dynamic operation scenes. At present, the industry lacks relevant technologies, and automatic driving perception technology cannot be applied to lifting operation perception due to differences in scenes, perspectives, working conditions, etc. In recent years, with the development of intelligent technology, a large number of intelligent auxiliary control technologies for the entire process of crane operations have been developed and applied. However, due to the lack of perception (detection) capabilities, there is still a lot of room for improvement in performance. It is urgent to upgrade the crane working environment perception method. In order to overcome the above-mentioned technical defects, this application proposes a crane working environment perception method. In order to clearly describe the method provided in this embodiment, please refer to Figure 1 to Figure 5 , including steps S110 to S130.

[0042] Step S110: Acquire perception data collected by a perception kit, where the perception kit includes at least one of a laser radar unit, an inertial measurement unit, and a camera unit, and the perception kit is installed at a predetermined position of the crane.

[0043] In one embodiment, the perception kit includes at least one of a laser radar unit (Light Detection and Ranging, LiDAR), an inertial measurement unit (Inertial Measurement Unit, IMU), and a camera unit (Camera). The perception kit is installed at a predetermined position of the crane, for example, it can be directly above the hook, at the boom of the crane, or at other positions of the crane, without limitation. Furthermore, the perception kit is not only divided into three categories, but the three categories can be arbitrarily combined; at the same time, the specific number of each category can also be arbitrarily set, for example, the perception kit can include multiple laser radar units, etc. In addition, the laser radar unit, the inertial measurement unit, and the camera unit can be installed at the same predetermined position, or they can be installed at different predetermined positions and installed in a dispersed manner. For example, the laser radar unit is installed directly above the hook, the inertial unit is installed at the farthest end of the boom, and the camera unit is installed on the support frame, etc. The specific installation setting method is not limited and can be arbitrarily set according to actual needs. In the embodiment in which multiple perception kits are installed at the same predetermined position, the perception kits can be connected by a rigid bracket to ensure that they have a certain consistency with each other, so as to facilitate the subsequent coordinate unification, etc.

[0044] The data collected by the perception kit is called perception data, and the perception data can be distinguished according to the data collected by different perception kits. The perception data collected by the lidar unit is called point cloud data; the perception data collected by the inertial measurement unit is called inertial data; and the perception data collected by the camera unit is called image data. In addition, each type of perception data not only includes the data collected by the corresponding perception kit itself, but also has a timestamp of collection to indicate the time of collection. Through the timestamp, different perception data can be aligned to facilitate the subsequent synchronization of the space-time coordinate system.

[0045] In one embodiment, the perception data collected by the perception kit is obtained, including: according to a predetermined position, using an external parameter calibration method to determine the relative position relationship between the perception kits, the relative position relationship is represented by an external parameter matrix; according to the external parameter matrix, the perception data is unified into the same spatial coordinate system; and the timestamp information of the perception data is calibrated so that the perception data collected by each perception kit is located in a unified time reference system.

[0046] In one embodiment, as described above, the perception kit may include multiple subunits, and there must be differences in the installation positions of different components, which will result in differences in the perception data directly collected by different perception kits for the same target. In order to eliminate this difference, the spatiotemporal coordinate systems of different perception kits should be unified when acquiring perception data. First, the relative position relationship between the perception kits can be determined based on the predetermined position. Among them, the predetermined position is a predetermined position, that is, the position of each perception kit is determined, so the relative position relationship between each other can be calculated by a predetermined calculation method. Generally, the relative relationship can be determined by the external parameter calibration method, and the relative relationship is represented by an external parameter matrix.

[0047] Then, the external parameter matrix is ​​used to unify all the perception data into a spatial coordinate system. For example, it can be assumed that one of the perception kits is used as the basis, and the coordinates of the perception data collected by other perception kits are all rotated and uniquely transferred to the corresponding spatial coordinate system. The unification process can be referred to as follows:

[0048]

[0049] In the above formula, x0y0z0 represents the coordinates after unification, and x1y1z1 represents the coordinates before unification; R X , R Y , R Z Represents the rotation of the XYZ axes respectively. Approximately, T XYZ It represents the translation of the XYZ axes.

[0050] In addition, the volume perception data in the previous text also includes timestamp information. While ensuring the uniformity of space, the timestamp can also be calibrated so that various types of perception data can be located in the same time and space. For example, by setting the timestamp, various types of perception data can have the same collection time and collection rate. Specifically, the collection time T of each frame of perception data is set to 100ms, and the collection frequency F is set to 10Hz. It can also be to directly align the timestamps of each perception data. It can also be other methods that enable various perception data to be located in the same space-time coordinate system. This application only makes a simple description of the technology here, and does not impose specific restrictions on the technical solution. The unified space-time reference system lays the foundation for the multi-sensor fusion perception of the subsequent perception kit.

[0051] In one embodiment, acquiring perception data collected by a perception kit includes: acquiring inertial data in the perception data, the inertial data being collected by an inertial measurement unit; calculating a motion change based on the inertial data, the motion change being used to describe a motion change of the perception kit during an acquisition process; acquiring initial perception data collected by the perception kit, and performing an error compensation operation on the initial perception data based on the motion change to correct an error in the initial perception data; when the initial perception data after the error compensation operation meets a convergence condition, marking the initial perception data as perception data; otherwise, repeatedly performing the error compensation operation on the initial perception data.

[0052] In one embodiment, the perception data is continuously acquired. At the same time, it can be understood that the perception kit installed on the crane will be affected by external environments such as vibration and wind caused by the operation of the crane, and it is definitely not static. The influence of the external environment will inevitably cause errors in the perception data collected by the perception kit. In order to ensure the accuracy of the data, it is necessary to eliminate the error. For this, the inertial data collected by the inertial measurement unit in the perception data can be obtained, and the motion change amount is calculated based on the inertial data. The motion change amount is used to describe the motion change of the perception kit during the acquisition process, which may specifically include position change and posture change. The motion change amount can be estimated by integrating the inertial data to determine the motion change amount during the acquisition of each frame of perception data. Among them, the motion amount can be more accurately obtained and calculated in combination with the position where the perception kit is installed. For example, the perception kit is installed on the boom head, and the perception kit is rigidly connected to the boom head. The position of the boom head can be obtained by obtaining the position of the perception kit. The motion change amount of the perception kit obtained in the above text is also the motion change amount of the boom head.

[0053] Get the initial perception data collected by the perception kit, and use the motion change to perform error compensation on the initial perception data to correct the error in the initial perception data. Specifically, it can include first performing motion compensation on each corresponding perception data of the same frame according to the motion change. Taking the point cloud data in the perception data as an example, the purpose of motion compensation is to compensate all the point cloud data to a certain moment, so that the point cloud data collected in the frame point cloud data within the T time period can be unified to a time point. The unification process can refer to the following example:

[0054] P start =T start-current *P currentent (2)

[0055] Match the compensated point cloud data to obtain a more accurate point cloud motion to determine the motion estimation error. Point cloud matching can use algorithms such as ICP (Iterative Closest Point) and NDT (Normal Distributions Transform). This application is only an example and not a limitation. The motion estimation error is redistributed to the point cloud data for error compensation.

[0056] Repeat the error compensation operation and determine whether the initial perception data meets the preset convergence conditions. If so, mark the initial perception data as perception data; otherwise, continue to repeat the operation until it is satisfied. It is worth noting that the compensation here is to compensate for the original perception data, so as to ensure that the errors in the perception data can be eliminated and the accuracy of the data can be guaranteed. On the one hand, motion estimation improves the accuracy of perception, and on the other hand, it helps to integrate the LiDAR point cloud data at different times; thereby improving the point cloud density and the detection accuracy of the downstream LiDAR monitoring module, which is applicable to the working conditions (altitude) and meets the long-distance 80m / 100m level detection.

[0057] Step S120: establishing a scene model according to the perception data, and identifying the target object, the scene model includes static obstacles and / or dynamic obstacles, and the target object includes the hook and / or the load.

[0058] In one embodiment, a scene model is established based on the perception data, including: acquiring image data and point cloud data in the perception data, the image data is collected by a camera unit, and the point cloud data is collected by a lidar unit; performing a posture relationship transformation on the point cloud data within a specified time range to determine that all object surfaces in the scene where the crane is located are position coordinates; performing a filtering operation on the point cloud data after the posture relationship transformation, and concatenating and accumulating the point cloud data after the filtering operation to obtain first point cloud data; determining a scene model based on the first point cloud data and the image data, wherein the scene model does not include the target object.

[0059] In one embodiment, for the convenience of analysis, a scene model of the crane can be constructed. The scene model can be used to more accurately understand the environment in which the crane is located, which can be represented by environmental objects. Environmental objects include static obstacles and dynamic obstacles. Static obstacles specifically include but are not limited to the ground, the crane itself, fixed facilities, etc.; dynamic obstacles include but are not limited to pedestrians, moving vehicles, other equipment at work, etc. The scene model is different from the target object. Here, only the scene model is constructed, and the target object is not identified. The processing of the target object will be described in detail later. The target object includes a hook and / or a load. In addition to the target object, the scene model does not include booms, wire ropes, etc.

[0060] Specifically, in the process of crane boom movement, the single-frame point cloud data within the specified time range T0 is transformed into the pose relationship by combining the crane boom information. start The point cloud data is spliced ​​and accumulated based on the time, thereby forming dense point cloud data, which is called the first point cloud data. The boom movement may include but is not limited to boom extension, boom amplitude change, boom rotation, and boom movement caused by deflection deformation during the process. Boom information may include boom rotation angle, boom length, boom angle (root, head), etc. Boom information is data determined by the crane and can be directly obtained.

[0061] The point cloud data within the specified time range is transformed into a pose relationship. The pose relationship transformation can be specifically carried out using Point-LIO, Fast-LIO, Fast-LIO2 and other real-time mapping and positioning algorithms, so as to determine the position coordinates of all object surfaces in the scene where the crane is located by integrating inertial data, image data and / or boom data based on the point cloud data.

[0062] Thirdly, the changed point cloud data is spliced ​​and accumulated within the specified time range to form dense scene point cloud data, and the filtering operation is performed to splice and accumulate the point cloud data after the filtering operation to obtain the first point cloud data. The filtering operation is to filter out objects such as hooks, loads, booms, wire ropes, etc. from the point cloud data.

[0063] In one embodiment, the filtering operation includes: using a preset straight line detection method to process all image data within a set time period, marking the processed straight line as a wire rope of the crane; identifying a hook area at one end of the wire rope based on the image data; obtaining cargo information, and determining a load area based on the cargo information and the hook area; and filtering out point cloud data within the hook area and the load area.

[0064] Image detection can be performed on the image data, such as using a straight line detection method to obtain the straight line part in the image data. The straight line detection method can be a corresponding algorithm such as Hough transform, and this application does not make specific restrictions. In combination with the installation position of the photographic unit, the wire rope of the crane is determined from the identified straight line. It can be understood that one end of the wire rope is connected to the crane boom and the other end is connected to the hook. According to this rule, the hook area can be identified. In addition, since the appearance of the hook is similar and there are fewer types, the position and outline of the hook can be obtained through a hook detection model generated by training a certain number of samples, thereby determining the hook outline. At the same time, the hook outline should have multiple wire ropes intersecting with it; if not, it is not considered to belong to the hook, and the hook area can also be identified based on this.

[0065] As for the load, due to the variety of loads and different appearances, it is impossible to directly detect the load using a deep learning model. However, the hook area can be used for direct identification. For example, the hook load is usually a load, so the load can be directly determined. For loads that have not yet been loaded, cargo information can also be obtained. The cargo information is used to indicate the cargo that the current crane needs to handle and the loading area of ​​the cargo, that is, the load area can be determined. It can also be obtained in combination with manual interaction. For example, the position or area range of the load can be manually identified in the image. The position is at least one pixel point within the load range in the image; the area range is the pixel point outline or the outer polygon outline of the load that can contain at least 3 points in the image. The image data is accurately segmented and extracted, and combined with the manual identification information, the position and outline information of the final output load precision wheel are determined, thereby determining the load area.

[0066] Finally, the point cloud data in the hook area and the load area are filtered out, so as to ensure that the target object and the corresponding boom, wire rope and other objects are not included in the constructed scene model. The point cloud data after the filtering operation is spliced ​​and accumulated to obtain the first point cloud data.

[0067] In one embodiment, a scene model is established based on perception data, including: constructing an initial scene model based on first point cloud data; acquiring all image data within a preset time period, and identifying a static sub-model and a dynamic sub-model from the initial scene model according to the change process of the image data, wherein the static sub-model is the stationary part of the initial scene model, including a background area and a first obstacle area, and the dynamic sub-model is the moving part of the initial scene model, including a hook area, a load area, and a second obstacle area; identifying static obstacles from the static sub-model; deleting the hook area and the load area from the dynamic sub-model, and identifying dynamic obstacles from the deleted second obstacle area; and summarizing the static obstacles and / or dynamic obstacles to obtain a scene model.

[0068] In one embodiment, the scene model is used as a basis to more quickly determine the target object, which includes one or more of the hook, load and / or obstacle. It is understandable that for the working environment of the crane, the important thing is the area of ​​the hook and the load part, and the relationship between the obstacles. The rest can be regarded as the background and will not affect the operation of the crane. At the same time, it is understandable that during the operation of the crane, usually only the hook, load and boom are moving, while the rest are relatively static. Therefore, the scene model can be used to distinguish the moving part and the static part, and divided into two types of sub-models. Then, the two types of sub-models are analyzed to locate the target object faster and more accurately.

[0069] Based on the above ideas, image data at any time can be obtained, and the static sub-model and dynamic sub-model can be identified from the scene model in real time according to the image data. The static sub-model is the stationary part of the scene model, including the background area and the first obstacle area; the dynamic sub-model is the moving part of the scene model, including the hook area, the load area and the second obstacle area. The background area is the environment and background of the crane, which can include the crane itself, surrounding buildings and facilities, etc. The first obstacle area is a relatively static obstacle, such as other temporarily placed goods, parked vehicles, etc. The hook area and the load area correspond to the hook and the load. The second obstacle area is a relatively moving obstacle, which specifically includes but is not limited to pedestrians, moving vehicles, and other facilities that are working.

[0070] Acquire all image data within a preset time period, which can specifically be the time period during which the crane is working. Objects in the image data may or may not move relative to each other within the time range. Based on the above differences, the static and changing parts of the scene model can be determined from the change process of the image data, and the static part of the scene model can be identified as a static sub-model, and the changing part of the scene model can be identified as a dynamic sub-model. Methods for distinguishing dynamic and static obstacles include: 1. Point cloud-based: identifying changes in point cloud space and finding dynamic points; 2. Image-based: identifying changes in pixels, finding dynamic pixels, and corresponding to the point cloud based on external parameter relationships; 3. Based on the fusion of the two.

[0071] The acquisition of the static sub-model can be determined by acquiring the installation environment information of the crane, which is used to indicate the environment in which the crane is set and the equipment information of the crane. The installation environment information can be the environment in which the crane is located, including but not limited to the installation location of the crane, the building information around the crane tower, etc.; the equipment information indicates the equipment data of the crane itself, which can include the boom information mentioned above and the height of the crane tower, etc. It can be understood that for the installation environment information of the crane, the related objects are relatively fixed, so for the scene model, it can be directly determined as the static sub-model part. Further, it can be understood that although they are both static parts, there are still some differences. For example, the buildings around the crane are fixed parts and are absolutely static; but the goods around the crane change every time they are transported and every day, and their static is only relative. Based on the above differences, the background area and the first obstacle area can be distinguished from the static sub-model according to the installation environment information. Among them, the background area is the absolutely static part in the static sub-model, which will not change accordingly for a long time and can be directly regarded as the background. The first obstacle area is a relatively static part, such as cargo, which may not change during one operation of the crane, but will produce significant differences after multiple accumulations. Buildings under renovation may also change over time. The relatively static part can be regarded as an obstacle, and the area where it is located is called the first obstacle area. This is for the convenience of subsequent identification.

[0072] For the dynamic sub-model, the crane hook area or load area can be detected in real time, and the remaining part can be regarded as the second obstacle area. All parts of the dynamic sub-model except the hook area and the load area can be marked as the second obstacle area. The difference between the obstacles included in the second obstacle area and the first obstacle area is that the obstacles in the former are moving, which can specifically include but are not limited to: staff, moving cars, other working machinery, etc. For the image data at the subsequent moment, combined with the position and contour information of the hook / load at the previous moment, the position and contour information of the hook / load at the current moment are tracked and predicted. The end-to-end visual object single target tracker can be mainly used to extract and merge features through the hybrid attention mechanism of the load target of the previous frame and the current search area, and obtain the correlation matching degree between the current area and the load target. When the matching area to be identified, it can be combined with the point cloud data for more specific identification, so as to determine the spatial information of each target formation, such as position and spatial size. The specific identification process will be described in detail later. The image detection results are integrated to achieve real-time detection and tracking of hooks / loads. The stability and robustness of the detection are ensured. According to the previous description, the dynamic sub-model will be deleted from it to avoid the impact on the construction of the scene model, that is, the operation corresponding to the filtering operation, which will not be repeated here.

[0073] The scene model is determined according to the first point cloud data and the image data. For the scene model determined by the first point cloud data, semantic division can be further performed: including extracting / removing dynamic target points, extracting ground plane equations, extracting the crane's own point cloud data, etc. Thus, the scene model is preliminarily distinguished to facilitate the subsequent division of regions. The representation of the scene model can include one or more of the following: ① A scene model based on point cloud, including the three-dimensional spatial coordinate points on the surface of all objects in the scene, and saving the appropriate point cloud density as needed; ② A scene model based on voxels, including the probability of whether each spatial position in the scene is occupied, and the distribution of the point cloud in the voxel (such as represented by a three-dimensional space Gaussian model), and selecting the appropriate voxel size as needed; ③ A scene model based on Mesh, using a series of polygons (usually triangles) of similar size and shape to approximate the model of a three-dimensional object. Scene modeling improves the modeling accuracy of static scenes and helps to improve the perception of static obstacles.

[0074] In one embodiment, identifying a target object based on perception data includes: acquiring image data and point cloud data in the perception data, the image data being acquired by a camera unit, and the point cloud data being acquired by a laser radar unit; entering an initialization phase based on the image data and the point cloud data, including performing a preprocessing operation on the point cloud data to obtain second point cloud data, the preprocessing operation including at least one of coordinate change, multi-frame accumulation, filtering and denoising, ground segmentation, and region of interest filtering; processing the image data by an image semantic segmentation method to determine at least one target object image mask; extracting a first object point cloud using the second point cloud data corresponding to the target object image mask based on an external parameter relationship between the camera unit and the laser radar unit; and determining at least one target object from the first object point cloud after accumulating multiple frames of the first object point cloud.

[0075] In one implementation, identifying a target object based on perception data may specifically include two stages, an initialization stage and a dynamic tracking stage, which will be described in stages later.

[0076] First, the initialization phase is entered, which mainly includes point cloud extraction and target object area recognition. A preprocessing operation is performed on the point cloud data to obtain the second point cloud data. The preprocessing operation includes at least one of coordinate change, multi-frame accumulation, filtering and denoising, ground segmentation, and interest area filtering. The preprocessing operations are explained one by one in the following.

[0077] Coordinate change: Using the result T of laser positioning, the point cloud single frame is transformed to obtain the point cloud in the vehicle coordinate system.

[0078] Multi-frame accumulation: Since single-frame point cloud data is sparse and unevenly distributed, it may not be possible to detect the complete shape of some obstacles. Multi-frame data is superimposed for processing to generate a dense point cloud, improve point cloud coverage, and support subsequent processing.

[0079] Filter denoising: remove point cloud noise and achieve point cloud data compression. Several filters can be used, such as SOR filter (Statistical Outlier Removal): a statistical outlier filtering method can be used to determine whether the point is a noise point by calculating the standard deviation and average of the distance between each point and the points in its neighborhood. If the distance between the point and the points in the neighborhood exceeds a certain multiple of the standard deviation, the point is considered to be a noise point; radius filter, to determine whether the points around the point cloud are noise; downsampling: for the filtered point cloud, there is an uneven density distribution, and the density in some areas is too high. The point cloud can be downsampled in the entire space to reduce the amount of data and improve processing efficiency while ensuring accuracy.

[0080] Ground segmentation: You can use a segmented plane model or a grid-based method. The premise of the segmented plane model is to assume that the ground is flat within a certain range, and then use the RANSAC or PCA principal component analysis method to obtain the normal vector, and combine the height threshold to identify the ground. You can also use algorithms such as cloth simulation according to the actual scene to improve the adaptability to slope and elevation changes. Combine the positioning information of the boom head and the ground height information to estimate the boom head height.

[0081] Filtering of interest area: In the vehicle body coordinate system, the vehicle body range is fixed (length, width and height), which belongs to the background point cloud. It can be filtered in advance according to the spatial coordinate range to reduce the amount of calculation.

[0082] The point cloud data after the preprocessing operation is called the second point cloud data, and the second point cloud data is a dense point cloud with a long multi-frame accumulation time.

[0083] The image data is then processed by the image semantic segmentation method to determine at least one target object image mask. Based on the external reference relationship between the camera unit and the lidar unit, the first object point cloud is extracted using the second point cloud data corresponding to the target object image mask. The specific recognition method can refer to the method of determining the area to be recognized by the image data: determine the hook, load and obstacle one by one. For the point cloud data recognition process of the target object, please refer to Figure 2a to Figure 2g , thereby determining Figure 2e and Figure 2g The hook area and obstacles shown are different. The specific identification process of each type of target object can be referred to as follows.

[0084] Rope and hook identification: The side view obtained from the point cloud data is projected onto a two-dimensional image to construct a binary image of the foreground and background. The Hough transform algorithm is used to detect straight lines on the binary image. The straight line closest to the current boom head is the boom rope, and the other end of the rope is the endpoint of the hook top. When the rope is detected, the point cloud below the rope is appropriately segmented using the height threshold (prior attribute information such as the length, width, and height of the hook), and the point cloud identification of the hook can be completed.

[0085] Load identification: Utilize the external parameter relationship between the laser radar and the camera to project the point cloud after removing the ground, hook / wire rope, etc. onto the image, build a corresponding relationship between the point cloud and the image, and combine the position and contour information of the load in the image to obtain the point cloud data within the contour range, i.e., the load point cloud.

[0086] The pixel area of ​​the target object is the area where the hook and the load are located in the image. Due to the variety of loads and their different shapes, it is impossible to directly detect the load using a deep learning model. Therefore, it can be obtained by combining manual interaction. The position or area range of the load needs to be manually identified in the image. The identified position is at least one pixel point within the load range in the image; the area range is the pixel point outline or the outer polygon outline of the load that can contain at least 3 points in the image. Among them, taking the selection of load as an example, the determination method can be as follows Figure 3a and Figure 3b As shown. Figure 3a The green dots in the figure indicate that the load is determined by pixel points. Figure 3b The green box in the figure is determined by the box selection method. In actual situations, there may be other options and determination methods. This is just a simple description of the solution, not a limitation. After the above processing flow, the pixel area of ​​the target object is finally determined. After determining the target object area, the mask of accurate segmentation can be obtained, and the effect of obtaining continuation can be obtained. Figure 3a and Figure 3b , you can get Figure 4a and Figure 4b The red rectangular area in the figure is the part that needs to be retained relative to the target object recognition, and the rest can be selected to be masked or disabled.

[0087] The identification of the area where the hook load is located has been described in detail in the previous part describing the dynamic sub-model. Please refer to the previous text for details. Here is a brief description. That is, the wire rope is identified by the preset straight line recognition method in combination with the installation position of the camera unit, and then the hook is identified based on the area at the end of the wire rope. Since the appearance of the hook is similar and there are fewer types, the position and outline of the hook can be obtained by training a hook detection model generated by a certain number of samples. The outline of the hook should have multiple wire ropes intersecting with it; if not, it is not considered to belong to the hook.

[0088] Obstacle identification: Clustering algorithms such as Euclidean clustering or density clustering (Density-Based Spatial Clustering of Applications with Noise, DBSCAN) can be used to cluster point cloud clusters. However, this process often has over-segmentation and under-segmentation, and it is necessary to add post-processing algorithms according to the actual situation, such as additional splitting and merging processing based on a certain pre-designed two-dimensional projection area size, or combined with intensity information. At this time, except for the hook, rope and load, the remaining obstacles do not have category information. By combining real-time load detection with the prior model, the accuracy of spatial position relationship calculation is improved; and a variety of methods for obtaining prior models and their information are proposed. According to the external parameter relationship between the camera unit and the lidar unit, the second point cloud data corresponding to the target object image mask is used to extract the first object point cloud. The external parameter relationship has been described in detail in the previous processing process of the same space-time system, so it will not be repeated here. For the first object point cloud, which is a dense point cloud, the acquisition effect can be referred to. Figure 5 shown.

[0089] Furthermore, the position of subsequent target objects can also be predicted. An end-to-end visual object single target tracker can be used to extract and merge features through a hybrid attention mechanism between the load target of the previous frame and the current search area, obtain the correlation matching degree between the current area and the load target, and select the highest matching area as the predicted bounding box for output when the matching degree reaches the benchmark. The processing flow of the tracker belongs to the processing method of the dynamic tracking stage, which will be used to more accurately identify the target object or determine the perception enhancement information. For the convenience of explanation, it will be explained in the process of obtaining the perception enhancement information later.

[0090] Step S130: Acquire and output the perception enhancement information between the scene model and the target object.

[0091] In one embodiment, it is understandable that the target object is in continuous motion and change, and the perception enhancement information between the target objects is also acquired through dynamic changes. The perception enhancement information includes but is not limited to 1. hook / load posture, 3D contour; 2. distance relative to obstacles; 3. arm head posture; 4. arm head height above the ground; 5. load height; 6. arm deflection and amplitude; 7. arm head-load wire rope length, etc. The perception enhancement information can be used for including but not limited to: 1. output-display distance information; 2. multi-level early warning based on distance information; 3. can support the operation of other subsequent control systems. For example, providing load height information for load translation; providing load position information for load sway reduction control; providing collision information for active safety, etc.

[0092] In order to obtain perception enhancement information, the image data and point cloud data can be combined for dynamic tracking and updating. That is, the dynamic tracking stage is entered according to the first object point cloud, including a multi-frame accumulation time period of the first object point cloud, and the pixel area of ​​the target object is determined according to the dynamic tracking of the image data; the point cloud data in the pixel area of ​​the target object is obtained and marked as the second object point cloud. The first object point cloud is aligned with the second object point cloud. The alignment method includes ICP or NDT, etc., and there is no specific limitation on this. The second object point cloud is replaced by the aligned first object point cloud to obtain the third object point cloud. The point cloud density is higher, the point cloud information is richer, and the subsequent information extraction accuracy is higher through alignment. Perception enhancement information is obtained through the third point cloud. For the acquisition process of the first object point cloud, that is, the method of determining the pixel area of ​​the target object, please refer to the previous text and will not be repeated here.

[0093] Perception enhancement information is mainly obtained through scene models and target objects, where the target object is specifically expressed in the form of a third point cloud. The relationship between the two can be referred to Figure 6 As shown. Taking the load obtained as an example in the above embodiments, Figure 6 The red part is the load represented by the third point cloud; the remaining green part represents the scene model, specifically the static sub-model part of the scene model. Through the correlation between the two, the various types of perception enhancement information mentioned above can be determined. And further, it should be noted that the perception enhancement information can also be obtained using the scene model or the target object alone, or even directly without the two, such as boom height, crane height, etc.

[0094] In order to continuously track and identify the target object, each target object can be tracked with a tracking model, and all tracking models form a tracking queue. When a new target object is identified, each object in the tracking queue first uses the tracking model to predict the current position and speed, and then the KM algorithm (Kuhn-Munkres Algorithm, Hungarian algorithm) can be used to associate the tracking observation with the tracking model, and finally the tracking observation is used to correct the prediction value of the corresponding tracking target and update the model.

[0095] The tracking model can be a Kalman filter (KF), and each target object is tracked using a KF model. When a new target object is identified, a KF model is set to track it, and the tracking queue is updated, and the tracking data in the tracking queue is continuously obtained to generate distance information. When the tracking is lost for a period of time / a specified number of frames, the KF model corresponding to the lost target object is cleared, thereby dynamically updating the tracking list.

[0096] Furthermore, the present application can identify more than one target object, so it can also achieve multi-target tracking and matching. The current target object (one or more of the hook, load, and obstacle) forms a source group, called the src group (Source Group); the targets in the tracking queue form a target group, called the tar group (Target Group). Traverse the target tars in the tar group i , calculate tar i The predicted value at the current moment In src Target src in the neighborhood j , remember dist i,j is the Euclidean distance between two targets, and calculates:

[0097] weight i,j =-dist i,j -1 (3)

[0098] For src not in Target src in the neighborhood k , recorded as:

[0099] weight i,k =MIN_EXPECT (4)

[0100] Among them, MIN_EXPECT and MAX_EXPECT are the expected boundaries, which should be set reasonably according to the task to avoid addition and subtraction overflow. i,j} matrix, and use the minimum weight KM algorithm to obtain the tracking data including the association relationship between the target in src and the target in tar.

[0101] For the extracted hook / load, since the point cloud data collected in real time is relatively sparse, it is impossible to obtain complete object surface data. A more complete (accurate) spatial occupancy range of the hook / load can be obtained by the following method: aligning the hook / load model acquired in advance with the hook / load point cloud extracted in real time, and the alignment method includes ICP (Iterative Closest Point) or NDT (Normal Distributions Transform). Taking the load in the previous article as an example, the schematic diagram of extracting the load from the third point cloud in real time based on the image mask can be referred to Figure 7 As shown; real-time matching of sparse load point cloud to the third point cloud can refer to Figure 8 As shown, the colored portion is the upper surface of the identified load, which can represent the load to be tracked and identified.

[0102] The method of obtaining the size information between the hook and the load also includes: importing a known 3D model, such as a design model (drawing), etc., using a measuring tool, such as a handheld automatic 3D scanner, etc., to scan the site and automatically generate a model, using a simple distance measurement tool to measure on site, and collecting and inputting the basic dimensions of the hook load (length, width, height, etc.). It also includes automatic detection, extracting point cloud data in the background area of ​​the static sub-model to indicate the ground; extracting the upper surface information of the load in the target object, and the position range of the load in the map space. Calculate the distance from the upper surface height of the load to the ground as the load height, and the load can meet the corresponding conditions, such as lifting on the ground, rigid objects, and keeping level in the air. And after lifting the load, obtain the upper surface height of the load at the moment of leaving the ground, which also needs to meet the corresponding conditions: such as the stable trend after the weight of the force limiter increases, and cooperate with the roll-up-amplitude change. And during the lifting process, calculate the distance from the upper surface height of the load to the ground as the load height, and update the load height. After all loads leave their original positions, calculate the upper surface height of the original position range, calculate the difference between the upper surface height when the load leaves the ground and the upper surface height of the original position range as the load height, update the load height, and apply to loads under other conditions, such as: lifting from non-level ground, non-level in the air, and non-rigid objects.

[0103] In one embodiment, perception enhancement information between a scene model and a target object is obtained and output, including: obtaining a relative distance in the perception enhancement information, and a first warning threshold and a second warning threshold corresponding to the relative distance, wherein the relative distance is the distance between the target object and a static obstacle and / or a dynamic obstacle; if the relative distance is less than the first warning threshold and greater than the second warning threshold, a first warning message is output, and the first warning message is used to prompt a user that there is a risk of collision; if the relative distance is less than the second warning threshold, a second warning message is output and / or the crane is controlled to stop moving, and the second warning message is used to prompt the user that the load will collide.

[0104] In one embodiment, for multi-level warning using distance information in the perception enhancement information, warning or control can be performed according to different warning thresholds. The distance information may include the location, outline, size, etc. of the hook, load, and obstacle. For the output diagram of the distance information, please refer to Fig. 9 shown.

[0105] The early warning can not only output the distance information of the target object itself, but also output the relationship between the target objects and the surrounding environment for alarm. The main one is the relative distance between the load and the obstacle, which is called relative distance. The relative distance is used to indicate the distance and distance relationship between the load and various directions / surrounding areas in the environment. The risk level can also be divided according to the distance, represented by different colors. The risk level can be determined by the relationship between the preset first warning threshold (called C1, for example, it can be 1m) and the second warning threshold (called C2, for example, it can be 0.5m). When the relative distance is greater than C1, it can be displayed in green, indicating that there is no collision risk in this direction; when it is less than C1 and greater than C2, it can be displayed in yellow, and the first warning information is output. The first warning information is used to prompt the user that there is a collision risk; when it is less than or equal to C2, it can be displayed in red, and the second warning information is output. The first warning information and the second warning information are used to prompt the user that the load will collide with the obstacle. The crane can also be controlled to stop moving, or make an appropriate amount of reverse movement to avoid a collision. For the output display effect under this implementation, please refer to Fig.10 .like Fig.10 As shown, the output mode may include the shortest distance in different directions, such as left and right in the figure for the slewing direction, where left is left slewing and right is right slewing; as shown in the figure, up and down are the luffing directions, where up is the falling luffing and down is the rising luffing. The distance information may also include other data, such as the height of the boom head to the ground (crane supporting the ground), the height of the hook / load to the ground (crane supporting the ground), the length of the wire rope from the boom head to the hook, and other information.

[0106] Therefore, the present application can obtain various perception data through the perception kit installed at the predetermined position of the crane to construct a scene model where the crane is located, so as to facilitate the identification of target objects including hooks, loads and / or obstacles. In this way, the working environment of the crane is perceived, and the corresponding distance information is determined for output through the correlation between the target objects in the scene model, so that real-time images of the hook / load surroundings and perception enhancement information of the surrounding environment can be provided, including distance, risk warning prompts, etc. Active safety braking control upgrade requirements are supported. It can meet the needs of crane operators to achieve safe lifting beyond visual range, greatly improving the convenience and safety of lifting. In addition, the method provided in the present application can be applied to any type of hooks and loads, improving the adaptability and detection accuracy of the method.

[0107] Fig.11 FIG. 1 shows an internal structure diagram of a computer device in an embodiment. The computer device may be a terminal or a server. Fig.11As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement the crane working environment perception method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement the crane working environment perception method. Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0108] In one embodiment, the present application further proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method described in any of the aforementioned embodiments.

[0109] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0110] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for sensing a crane working environment, characterized in that: The steps include: Acquiring perception data collected by a perception kit, wherein the perception kit includes at least one of a laser radar unit, an inertial measurement unit, and a camera unit, and the perception kit is installed at a predetermined position of the crane; Establishing a scene model according to the perception data, and identifying a target object, wherein the scene model includes static obstacles and / or dynamic obstacles, and the target object includes a hook and / or a load; Acquire and output perception enhancement information between the scene model and the target object.

2. The crane working environment perception method according to claim 1, characterized in that: The acquiring of the perception data collected by the perception kit includes: According to the predetermined positions, the relative position relationship between the sensing kits is determined by using an external parameter calibration method, wherein the relative position relationship is represented by an external parameter matrix; Unifying the perception data into the same spatial coordinate system according to the extrinsic parameter matrix; The timestamp information of the perception data is calibrated so that the perception data collected by each perception kit is located in a unified time reference system.

3. The crane working environment perception method according to claim 1, characterized in that: The acquiring of the perception data collected by the perception kit includes: Acquire inertial data from the perception data, where the inertial data is collected and acquired by the inertial measurement unit; Calculating a motion change amount according to the inertial data, wherein the motion change amount is used to describe the motion change of the perception kit during the acquisition process; Acquire initial perception data collected by the perception kit, and perform an error compensation operation on the initial perception data according to the motion change amount to correct errors in the initial perception data; When the initial perception data after performing the error compensation operation meets the convergence condition, the initial perception data is marked as the perception data; otherwise, the error compensation operation is repeatedly performed on the initial perception data.

4. The crane working environment perception method according to claim 1, characterized in that: The step of establishing a scene model according to the perception data comprises: Acquire image data and point cloud data in the perception data, wherein the image data is acquired by the camera unit, and the point cloud data is acquired by the laser radar unit; Performing a posture relationship transformation on the point cloud data within a specified time range to determine the position coordinates of all object surfaces in the scene where the crane is located; Performing a filtering operation on the point cloud data after the posture relationship transformation, and concatenating and accumulating the point cloud data after the filtering operation to obtain first point cloud data; The scene model is determined according to the first point cloud data and the image data, wherein the scene model does not include the target object.

5. The crane working environment perception method according to claim 4, characterized in that: The filtering operation comprises: Using a preset straight line detection method to process all the image data within a set time period, marking the processed straight line as the steel wire rope of the crane; identifying the hook area at one end of the steel wire rope according to the image data; Acquiring cargo information, and determining a load area according to the cargo information and the hook area; The point cloud data in the hook area and the load area are filtered out.

6. The crane working environment perception method according to claim 4, characterized in that: The step of establishing a scene model according to the perception data comprises: Constructing an initial scene model according to the first point cloud data; Acquire all the image data within a preset time period, and identify a static sub-model and a dynamic sub-model from the initial scene model according to the change process of the image data, wherein the static sub-model is a stationary part of the initial scene model, including a background area and a first obstacle area, and the dynamic sub-model is a moving part of the initial scene model, including a hook area, a load area, and a second obstacle area; Identify the static obstacle from the static sub-model; delete the hook area and the load area from the dynamic sub-model, and identify the dynamic obstacle from the deleted second obstacle area; The static obstacles and / or the dynamic obstacles are aggregated to obtain the scene model.

7. The crane working environment perception method according to claim 1, characterized in that: The identifying the target object according to the perception data comprises: Acquire image data and point cloud data in the perception data, wherein the image data is acquired by the camera unit, and the point cloud data is acquired by the laser radar unit; Entering an initialization phase according to the image data and the point cloud data, including performing a preprocessing operation on the point cloud data to obtain second point cloud data, the preprocessing operation including at least one of coordinate change, multi-frame accumulation, filtering and denoising, ground segmentation, and region of interest filtering; Processing the image data by an image semantic segmentation method to determine at least one target object image mask; extracting a first object point cloud using the second point cloud data corresponding to the target object image mask according to an external parameter relationship between the camera unit and the laser radar unit; and determining at least one target object from the first object point cloud after accumulating multiple frames of the first object point cloud; The acquiring and outputting the perception enhancement information between the scene model and the target object includes: Entering a dynamic tracking phase according to the first object point cloud, including accumulating a multi-frame time period of the first object point cloud, and determining a pixel area of ​​a target object according to the dynamic tracking of the image data; Acquire point cloud data within a pixel area of ​​the target object and mark it as a second object point cloud; align the first object point cloud with the second object point cloud, and replace the second object point cloud with the aligned first object point cloud to obtain a third object point cloud; The perception enhancement information is acquired according to the third point cloud.

8. The crane working environment perception method according to claim 1, characterized in that: The acquiring and outputting the perception enhancement information between the scene model and the target object includes: Acquire a relative distance in the perception enhancement information, and a first warning threshold and a second warning threshold corresponding to the relative distance, wherein the relative distance is a distance between the target object and the static obstacle and / or the dynamic obstacle; If the relative distance is less than the first warning threshold and greater than the second warning threshold, a first warning message is output, where the first warning message is used to remind the user that there is a collision risk; If the relative distance is less than the second warning threshold, a second warning message is output and / or the crane is controlled to stop moving, and the second warning message is used to remind the user that the load will collide.

9. A computer device, characterized in that: including a processor and a memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Bridge crane hoisting safety anti-collision system and method based on dynamic binocular vision

    CN112418103A

  • Tower crane hoisting object identification and collision information measurement system and method

    CN113860178A

  • Engineering machinery and dynamic Anti-collision method, device, and system for operation space of the engineering machinery

    US20210171324A1

Cited By

  • Mobile crane operation state real-time monitoring method based on digital twinning

    CN122403281A