Crane working environment perception methods, computer equipment and computer-readable storage media
By installing a perception kit consisting of lidar, inertial measurement, and camera units on a crane, unified perception data and construction of scene models have solved the problem of accuracy in crane working environment perception, enabling safe and real-time detection of hooks and loads, and improving the safety and efficiency of crane operation.
Patent Information
- Application Number
- CN202510034896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing crane working environment perception methods cannot provide accurate information on the distance between the load, hook, and environmental obstacles, resulting in low operational efficiency during beyond-line-of-sight lifting and risks associated with relying on manual judgment.
A perception kit consisting of a lidar unit, an inertial measurement unit, and a camera unit is used. The spatiotemporal coordinate system of the perception data is unified through the external parameter calibration method. Combined with inertial data error compensation and image data processing, a scene model of the crane is constructed and the target object is identified, and perception enhancement information is output.
It enables real-time and complete monitoring of the crane's working environment, providing three-dimensional spatial position and relative relationship information of the hook, load, and surrounding environment, thereby improving safety and operational efficiency.
Smart Images

Figure CN119976640B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent crane technology, and in particular relates to a crane working environment perception method, computer equipment, and computer-readable storage medium. Background Technology
[0002] Cranes are key equipment for transporting goods in the construction industry and are widely used in engineering construction. Their safe operation and maintenance are crucial to the construction industry. Currently, crane safety protection mainly includes monitoring of the lifting points and the area around the turntable. Most of these monitoring systems use displays to show the transmitted image information of the lifting points or the area around the turntable.
[0003] Image information can only achieve monitoring functions and cannot provide actual information about the load, the distance between the hook and environmental obstacles. In current practical operations, such as performing beyond-line-of-sight lifting, the complexity of the lifting scenario and the inadequacy of algorithms currently limit the crane's perception needs in real dynamic operating scenarios. It can only be accomplished through real-time communication with support personnel based on image information. Relying on the operator and signalman to work together is inefficient, and the effectiveness of communication depends on the expression and understanding abilities of both parties. Furthermore, relying on manual image interpretation inherently has the technical drawbacks of limited information and high risk. Therefore, the perception of the crane's working environment urgently needs improvement. How to safely and accurately perceive the crane's working environment is a technical problem that urgently needs to be solved by those skilled in the art.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] Based on this, it is necessary to propose a crane working environment perception method, computer equipment, and computer-readable storage medium to address the above problems, which can safely and accurately perceive the crane's working environment.
[0006] The technical problem solved by this application is achieved by the following technical solution:
[0007] This application provides a method for sensing the working environment of a crane, comprising the following steps: acquiring sensing data collected by a sensing kit, the sensing kit including at least one of a lidar unit, an inertial measurement unit, and a camera unit, the sensing kit being installed at a predetermined position on the crane; establishing a scene model based on the sensing data and identifying target objects, the scene model including static obstacles and / or dynamic obstacles, the target objects including hooks and / or loads; acquiring and outputting enhanced sensing information between the scene model and the target objects.
[0008] In an optional embodiment of this application, acquiring the sensing data collected by the sensing kits includes: determining the relative positional relationship between the sensing kits according to a predetermined position using an extrinsic parameter calibration method, wherein the relative positional relationship is represented by an extrinsic parameter matrix; unifying the sensing data to the same spatial coordinate system according to the extrinsic parameter matrix; and calibrating the timestamp information of the sensing data so that the sensing data collected by each sensing kit is located in a unified time reference system.
[0009] In an optional embodiment of this application, acquiring sensing data collected by the sensing kit includes: acquiring inertial data from the sensing data, wherein the inertial data is acquired by an inertial measurement unit; calculating motion change based on the inertial data, wherein the motion change is used to describe the motion change of the sensing kit during the acquisition process; acquiring initial sensing data collected by the sensing kit; performing an error compensation operation on the initial sensing data based on the motion change to correct errors in the initial sensing data; when the initial sensing data after the error compensation operation meets the convergence condition, marking the initial sensing data as sensing data; otherwise, repeating the error compensation operation on the initial sensing data.
[0010] In an optional embodiment of this application, establishing a scene model based on perception data includes: acquiring image data and point cloud data from the perception data, wherein the image data is acquired by a camera unit and the point cloud data is acquired by a lidar unit; performing pose relationship transformation on the point cloud data within a specified time range to determine the position coordinates of all object surfaces in the scene where the crane is located; performing a filtering operation on the point cloud data after pose relationship transformation; and stitching and accumulating the point cloud data after the filtering operation to obtain first point cloud data; and determining a scene model based on the first point cloud data and image data, wherein the scene model does not include the target object.
[0011] In an optional embodiment of this application, the filtering operation includes: processing all image data within a set time period using a preset straight line detection method, marking the processed straight line as the crane's wire rope; identifying the hook area at one end of the wire rope based on the image data; obtaining cargo information, determining the load area based on the cargo information and the hook area; and filtering out the point cloud data within the hook area and the load area.
[0012] In an optional embodiment of this application, establishing a scene model based on perception data includes: constructing an initial scene model based on first point cloud data; acquiring all image data within a preset time period; identifying a static sub-model and a dynamic sub-model from the initial scene model based on the change process of the image data; the static sub-model being the stationary part of the initial scene model, including a background area and a first obstacle area; the dynamic sub-model being the moving part of the initial scene model, including a hook area, a load area, and a second obstacle area; identifying static obstacles from the static sub-model; deleting the hook area and load area from the dynamic sub-model; identifying dynamic obstacles from the deleted second obstacle area; and summarizing the static obstacles and / or dynamic obstacles to obtain a scene model.
[0013] In an optional embodiment of this application, identifying a target object based on perception data includes: acquiring image data and point cloud data from the perception data, wherein the image data is acquired by a camera unit and the point cloud data is acquired by a lidar unit; entering an initialization phase based on the image data and point cloud data, including performing preprocessing operations on the point cloud data to obtain second point cloud data, wherein the preprocessing operations include at least one of coordinate transformation, multi-frame accumulation, filtering and noise reduction, ground segmentation, and region of interest filtering; processing the image data using an image semantic segmentation method to determine at least one target object image mask; and using the target object image mask corresponding to the external parameter relationship between the camera unit and the lidar unit. The second point cloud data is used to extract the first object point cloud; after accumulating the first object point cloud over multiple frames, at least one target object is identified from the first object point cloud; perceptual enhancement information between the scene model and the target object is acquired and output, including: entering the dynamic tracking stage based on the first object point cloud, including accumulating the first object point cloud over multiple frames over a time period, and dynamically tracking and determining the pixel region of the target object based on the image data; acquiring point cloud data within the pixel region of the target object and marking it as the second object point cloud; registering the first object point cloud with the second object point cloud, and replacing the second object point cloud with the registered first object point cloud to obtain the third object point cloud; and acquiring perceptual enhancement information based on the third point cloud.
[0014] In an optional embodiment of this application, acquiring and outputting enhanced perception information between a scene model and a target object includes: acquiring the relative distance in the enhanced perception information, and a first warning threshold and a second warning threshold corresponding to the relative distance, wherein the relative distance is the distance between the target object and static obstacles and / or dynamic obstacles; if the relative distance is less than the first warning threshold and greater than the second warning threshold, then outputting a first alarm message, which is used to alert the user of a collision risk; if the relative distance is less than the second warning threshold, then outputting a second warning message and / or controlling the crane to stop moving, which is used to alert the user that a collision will occur with the load. This application also provides a computer device, including a processor and a memory: the processor is used to execute a computer program stored in the memory to implement the method described above.
[0015] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as described above.
[0016] The embodiments of this application have the following beneficial effects:
[0017] This application enables the acquisition of diverse sensing data through a sensing kit installed at a predetermined location on a crane to construct a scene model of the crane's location, facilitating the identification of target objects, including hooks, loads, and / or obstacles. This allows for the perception of the crane's working environment, and the determination of corresponding distance information based on the relationships between target objects in the scene model. This enables real-time and complete detection of the three-dimensional spatial position, size, and relative relationships of the hook, load, and surrounding environment, facilitating safe beyond-line-of-sight work.
[0018] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it according to the contents of the specification, and to make the above and other objects, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] in:
[0021] Figure 1A flowchart illustrating a crane working environment perception method provided in one embodiment;
[0022] Figure 2a A schematic diagram of raw point cloud data provided in one embodiment;
[0023] Figure 2b A schematic diagram of point cloud data after filtering a crane, provided in one embodiment;
[0024] Figure 2c A schematic diagram of downsampled and filtered point cloud data provided in one embodiment;
[0025] Figure 2d A schematic diagram of point cloud data after filtering the ground, provided in one embodiment;
[0026] Figure 2e This is a schematic diagram illustrating the identification of the hook area in one embodiment;
[0027] Figure 2f A schematic diagram of point cloud data of an obstacle provided in one embodiment;
[0028] Figure 2g A schematic diagram illustrating the identification of obstacles provided in one embodiment;
[0029] Figure 3a A first schematic diagram defining the load region range provided in one embodiment;
[0030] Figure 3b A second schematic diagram defining the load region range provided in one embodiment;
[0031] Figure 4a A first schematic diagram of load region mask determination provided in one embodiment;
[0032] Figure 4b A second schematic diagram showing the determination of a load region mask provided in one embodiment;
[0033] Figure 5 This is a schematic diagram illustrating the point cloud acquisition effect of a first object provided in one embodiment;
[0034] Figure 6 A schematic diagram of the target object and scene model provided in one embodiment;
[0035] Figure 7 This is a schematic diagram illustrating real-time extraction of payload from a third point cloud based on an image mask, as provided in one embodiment.
[0036] Figure 8 A schematic diagram illustrating real-time matching of sparse load point cloud to a third point cloud as provided in one embodiment;
[0037] Figure 9This is a schematic diagram of a first method for outputting distance information according to an embodiment;
[0038] Figure 10 This is a schematic diagram illustrating the second output method of distance information provided in one embodiment;
[0039] Figure 11 This is a schematic block diagram of the structure of a computer device provided in one embodiment. Detailed Implementation
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0041] Existing technologies for monitoring lifting points cannot provide actual horizontal and vertical distance information. In lifting operations beyond visual range, judgment must be made manually based on image information, resulting in limited information and significant risks. Alternatively, real-time communication with support personnel is required, but the effectiveness depends on the expression and comprehension abilities of both parties. Unmanned lifting technology is the trend in the industry. Intelligent sensing technology, as the foundational core technology of unmanned lifting, is currently limited by the complexity of lifting scenarios and insufficient algorithm capabilities, failing to meet the perception needs of cranes in real dynamic operating scenarios. Currently, the industry lacks relevant technologies, and autonomous driving sensing technologies, due to differences in scenarios, perspectives, and operating conditions, are also unsuitable for lifting operation perception. In recent years, with the development of intelligent technologies, a large number of intelligent auxiliary control technologies for the entire crane operation process have been developed and implemented. However, due to insufficient perception (detection) capabilities, there is still considerable room for improvement in performance. There is an urgent need to upgrade crane working environment perception methods. To overcome the above-mentioned technical deficiencies, this application proposes a crane working environment perception method. For a clear description of the method provided in this embodiment, please refer to... Figures 1-5 This includes steps S110 to S130.
[0042] Step S110: Acquire sensing data collected by the sensing kit. The sensing kit includes at least one of a lidar unit, an inertial measurement unit, and a camera unit. The sensing kit is installed at a predetermined position on the crane.
[0043] In one embodiment, the sensing kit includes at least one of a LiDAR (Light Detection and Ranging) unit, an Inertial Measurement Unit (IMU), and a camera unit. The sensing kit is installed at a predetermined location on the crane, such as directly above the hook, on the crane boom, or at other locations on the crane; this is not limited. Furthermore, the sensing kit is not only divided into three categories, which can be arbitrarily combined, but the specific quantity of each category can also be arbitrarily set; for example, the sensing kit may include multiple LiDAR units. In addition, the LiDAR unit, IMU, and camera unit can be installed at the same predetermined location or at different predetermined locations, distributed across the installation. For example, the LiDAR unit can be installed directly above the hook, the IMU at the farthest end of the boom, and the camera unit on a support frame; the specific installation method is not limited and can be arbitrarily set according to actual needs. In embodiments where multiple sensing kits are installed at the same predetermined location, the sensing kits can be connected by rigid supports to ensure a certain degree of consistency, facilitating subsequent coordinate unification, etc.
[0044] The data collected by the sensing kits is called sensing data, and sensing data can be further categorized according to the sensing kits used. Sensing data collected by the LiDAR unit is called point cloud data; sensing data collected by the inertial measurement unit is called inertial data; and sensing data collected by the camera unit is called image data. Furthermore, each type of sensing data includes not only the data itself collected by the corresponding sensing kit but also a timestamp indicating the time of collection. The timestamp allows for the alignment of different sensing data sets, facilitating subsequent synchronization of the spatiotemporal coordinate system.
[0045] In one embodiment, acquiring sensing data collected by a sensing kit includes: determining the relative positional relationship between sensing kits based on a predetermined location using an extrinsic calibration method, wherein the relative positional relationship is represented by an extrinsic matrix; unifying the sensing data to the same spatial coordinate system based on the extrinsic matrix; and calibrating the timestamp information of the sensing data so that the sensing data collected by each sensing kit is located under a unified time reference system.
[0046] In one embodiment, as described above, a sensing kit may include multiple seed units. The installation positions of these different components will inevitably differ, leading to discrepancies in the sensing data directly collected by different sensing kits for the same target. To eliminate this discrepancy, the spatiotemporal coordinate system of different sensing kits should be unified when acquiring sensing data. First, the relative positional relationship between the sensing kits can be determined based on predetermined positions. These predetermined positions are pre-determined, meaning the positions of each sensing kit are fixed. Therefore, the relative positional relationship can be calculated using a predetermined calculation method. Generally, extrinsic parameter calibration can be used to determine the relative relationship, and this relationship is represented by an extrinsic parameter matrix.
[0047] Then, using the extrinsic parameter matrix, all the sensing data is unified into a single spatial coordinate system. For example, we can assume that one sensing kit is used as a base, and rotate and unify the coordinates of all sensing data collected by other sensing kits into the corresponding spatial coordinate system. The unification process can be referenced as follows:
[0048]
[0049] In the above formula, x0y0z0 represents the unified coordinates, and x1y1z1 represents the original coordinates; R X R Y R Z These represent the rotation amounts along the X, Y, and Z axes, respectively. Approximately, T... XYZ This represents the translation amount along the XYZ axes.
[0050] Furthermore, the volumetric sensing data mentioned above also includes timestamp information. While ensuring spatial uniformity, timestamps can be assigned to ensure that various types of sensing data reside in the same spatiotemporal context. For example, by setting timestamps, various types of sensing data can have the same acquisition duration and acquisition frequency. Specifically, this could be achieved by setting the acquisition duration T for each sensing data point in a single frame to 100ms and the acquisition frequency F to 10Hz. Alternatively, the timestamps of each sensing data point can be directly aligned. Other methods that enable various sensing data points to reside in the same spatiotemporal coordinate system are also possible. This application only provides a brief technical description and does not impose specific limitations on the technical solutions. A unified spatiotemporal reference system lays the foundation for multi-sensor fusion sensing in subsequent sensing suites.
[0051] In one embodiment, acquiring sensing data collected by a sensing kit includes: acquiring inertial data from the sensing data, the inertial data being acquired by an inertial measurement unit; calculating motion change based on the inertial data, the motion change being used to describe the motion change of the sensing kit during the acquisition process; acquiring initial sensing data collected by the sensing kit; performing an error compensation operation on the initial sensing data based on the motion change to correct errors in the initial sensing data; when the initial sensing data after the error compensation operation meets the convergence condition, marking the initial sensing data as sensing data; otherwise, repeating the error compensation operation on the initial sensing data.
[0052] In one implementation, the sensing data is continuously acquired. It is understood that the sensing kit installed on the crane is subject to external environmental influences such as vibrations from crane operation and wind, and is therefore not stationary. These external environmental influences inevitably lead to errors in the sensing data collected by the sensing kit. To ensure data accuracy, this error must be eliminated. This can be achieved by acquiring inertial data collected by the inertial measurement unit (IMU) within the sensing data, and calculating the motion change based on this data. The motion change describes the motion changes of the sensing kit during the acquisition process, specifically including position and attitude changes. The motion change can be estimated by integrating the inertial data, thus determining the motion change during each frame of sensing data acquisition. Furthermore, the calculated motion change can be more accurately obtained by considering the installation location of the sensing kit. For example, if the sensing kit is installed at the boom head and rigidly connected to it, obtaining the position of the sensing kit will yield the position of the boom head, and the previously calculated motion change of the sensing kit will also be the obtained motion change of the boom head.
[0053] The initial sensing data collected by the sensing kit is acquired, and error compensation is performed on the initial sensing data using motion changes to correct errors in the initial sensing data. Specifically, this may include performing motion compensation on each corresponding sensing data in the same frame based on motion changes. Taking point cloud data in the sensing data as an example, the purpose of motion compensation is to compensate all point cloud data to a certain time point. This allows the point cloud data collected within a time period T of that frame to be unified to a single point in time. The unification process can be seen in the following example:
[0054] P start =T start-current *P currentent (2)
[0055] The compensated point cloud data is matched to obtain more accurate point cloud motion data, thereby determining the motion estimation error. Point cloud matching can employ algorithms such as ICP (Iterative Closest Point) and NDT (Normal Distributions Transform); this application is merely an example and not a limitation. The motion estimation error is then redistributed to the point cloud data for error compensation.
[0056] Error compensation is repeatedly performed, and the initial sensing data is checked to see if it meets the preset convergence criteria. If it does, the initial sensing data is marked as sensing data; otherwise, the process continues until the criteria are met. It's important to note that this compensation applies to the original sensing data, ensuring the elimination of errors within the sensing data and guaranteeing its accuracy. Motion estimation improves sensing accuracy and facilitates the fusion of LiDAR point cloud data from different times, thereby increasing point cloud density and improving the detection accuracy of downstream LiDAR monitoring modules. This applies to various working conditions (altitude) and meets the requirements for long-range detection at 80m / 100m levels.
[0057] Step S120: Establish a scene model based on the perception data and identify the target object. The scene model includes static obstacles and / or dynamic obstacles, and the target object includes hooks and / or loads.
[0058] In one embodiment, establishing a scene model based on perception data includes: acquiring image data and point cloud data from the perception data, wherein the image data is acquired by a camera unit and the point cloud data is acquired by a lidar unit; performing pose transformation on the point cloud data within a specified time range to determine the position coordinates of all object surfaces in the scene where the crane is located; performing a filtering operation on the point cloud data after pose transformation; and stitching and accumulating the point cloud data after the filtering operation to obtain first point cloud data; and determining a scene model based on the first point cloud data and the image data, wherein the scene model does not include the target object.
[0059] In one implementation, to facilitate analysis, a scene model of the crane's location can be constructed. This scene model provides a more accurate understanding of the crane's environment and can be represented by environmental objects. Environmental objects include static obstacles and dynamic obstacles. Static obstacles specifically include, but are not limited to, the ground, the crane itself, and fixed facilities; dynamic obstacles include, but are not limited to, pedestrians, moving vehicles, and other operating equipment. The scene model differs from the target object; here, only the scene model is constructed, and the target object is not identified. The processing of the target object will be described in detail later. The target object includes the hook and / or load. Furthermore, the scene model does not include the boom, wire rope, etc., in addition to excluding the target object.
[0060] Specifically, the establishment of the scene model involves transforming the pose relationships of single-frame point cloud data within a specified time range T0 during the movement of the crane boom, based on the crane boom information. start Point cloud data is stitched and accumulated based on a specific time frame to form a dense point cloud, referred to as the first point cloud data. Boom movement can include, but is not limited to, boom extension, boom luffing, boom slewing, and the boom movement caused by deflection deformation during these processes. Boom information can include boom slewing angle, boom length, boom angle (root and head), etc. Boom information is data determined by the crane and can be directly obtained.
[0061] The pose relationship transformation is performed on point cloud data within a specified time range. Specifically, the pose relationship transformation can employ real-time mapping and localization algorithms such as Point-LIO, Fast-LIO, and Fast-LIO2 to determine the position coordinates of all object surfaces in the scene where the crane is located, based on the point cloud data and by fusing inertial data, image data, and / or boom data.
[0062] Next, the modified point cloud data is stitched and accumulated within a specified time range to form dense scene point cloud data. A filtering operation is then performed, and the point cloud data after the filtering operation is stitched and accumulated to obtain the first point cloud data. The filtering operation removes objects such as hooks, loads, booms, and wire ropes from the point cloud data.
[0063] In one embodiment, the filtering operation includes: processing all image data within a preset time period using a preset straight line detection method, marking the processed straight lines as the crane's wire rope; identifying the hook area at one end of the wire rope based on the image data; obtaining cargo information, determining the load area based on the cargo information and the hook area; and filtering out the point cloud data within the hook area and the load area.
[0064] Image detection can be performed on the image data, such as using line detection methods to obtain the straight lines in the image data. Line detection methods can include algorithms such as the Hough transform, and this application does not impose specific limitations. Combined with the installation position of the camera unit, the crane's wire rope is determined from the identified straight lines. It is understood that one end of the wire rope is connected to the crane boom, and the other part is connected to the hook. Based on this rule, the hook area can be identified. Furthermore, since hooks have similar shapes and few types, the position and outline of the hook can be obtained through a hook detection model trained with a certain number of samples, thereby determining the hook outline. Simultaneously, the hook outline should intersect with multiple wire ropes; if not, it is not considered a hook, and based on this, the hook area can also be identified.
[0065] For loads, due to their diverse types and shapes, deep learning models cannot be used to directly detect them. However, they can be directly identified using the hook area. For example, hook loads are usually loads, so the load can be directly determined. For loads that are not yet loaded, cargo information can be obtained. This cargo information indicates the cargo that the crane needs to handle and the loading area of the cargo, thus determining the load area. Manual interaction can also be used to obtain this information. For example, the location or area of the load can be manually marked in the image. The location is at least one pixel within the load area in the image; the area is the pixel outline or bounding polygon outline of the load that can contain at least three points in the image. Accurate segmentation and extraction of the image data, combined with the manually marked information, determines the final output of the precise position and outline information of the load wheel, thereby determining the load area.
[0066] Finally, the point cloud data in the hook area and load area are filtered out to ensure that the target object, as well as the corresponding boom, wire rope, and other objects, are not included in the constructed scene model. The point cloud data after the filtering operation is performed are then stitched together and accumulated to obtain the first point cloud data.
[0067] In one embodiment, establishing a scene model based on perception data includes: constructing an initial scene model based on first point cloud data; acquiring all image data within a preset time period; identifying a static sub-model and a dynamic sub-model from the initial scene model based on the change process of the image data; the static sub-model being the stationary part of the initial scene model, including a background area and a first obstacle area; the dynamic sub-model being the moving part of the initial scene model, including a hook area, a load area, and a second obstacle area; identifying static obstacles from the static sub-model; deleting the hook area and load area from the dynamic sub-model; identifying dynamic obstacles from the deleted second obstacle area; and summarizing the static obstacles and / or dynamic obstacles to obtain a scene model.
[0068] In one implementation, using a scene model as a basis allows for faster identification of target objects, including one or more of hooks, loads, and / or obstacles. It is understood that for the crane's working environment, the area containing the hook and load, and its relationship to obstacles, are crucial; the remaining areas can be considered background and will not affect the crane's operation. Furthermore, it is understood that during crane operation, typically only the hook, load, and boom are in motion, while the rest are relatively stationary. Therefore, the scene model can be used to distinguish between moving and stationary parts, dividing them into two sub-models. Further analysis of these two sub-models allows for faster and more accurate location of the target object.
[0069] Based on the above approach, image data at any given moment can be acquired. Static and dynamic sub-models can then be identified in real-time from the scene model based on this image data. The static sub-model refers to the stationary parts of the scene model, including the background area and the first obstacle area. The dynamic sub-model refers to the moving parts of the scene model, including the hook area, the load area, and the second obstacle area. The background area represents the environment in which the crane is located, which may include the crane itself and surrounding buildings and facilities. The first obstacle area consists of relatively static obstacles, such as temporarily placed goods or parked vehicles. The hook and load areas correspond to the hook and load, respectively. The second obstacle area consists of relatively moving obstacles, including but not limited to pedestrians, moving vehicles, and other operating facilities.
[0070] Acquire all image data within a preset time period, specifically the period during which the crane operates. Objects within the image data may or may not move relative to each other within this timeframe. Based on these differences, the static and dynamic parts of the scene model can be identified from the changes in the image data. The static parts are identified as static sub-models, and the dynamic parts as dynamic sub-models. Methods for distinguishing between dynamic and static obstacles include: 1. Point cloud-based: identifying changes in the point cloud space to find dynamic points; 2. Image-based: identifying pixel changes to find dynamic pixels and mapping them to the point cloud based on extrinsic parameters; 3. A fusion of both methods.
[0071] The static sub-model can be obtained by acquiring the crane's installation environment information. This information indicates the environment in which the crane is set up and its equipment information. The installation environment information includes, but is not limited to, the crane's installation location and surrounding building information. The equipment information indicates the crane's own equipment data, which may include the boom information mentioned earlier, as well as the tower height. It's understandable that the objects in the crane's installation environment information are relatively fixed, therefore, for the scene model, it can be directly identified as the static sub-model. Furthermore, it's understandable that although both are static, there are still some differences. For example, the buildings around the crane are fixed and absolutely static; however, the goods around the crane change with each transport and each day, so their static state is only relative. Based on these differences, the background area and the first obstacle area can be distinguished from the static sub-model based on the installation environment information. The background area is the absolutely static part of the static sub-model, and it will not change over a long period of time; it can be directly considered as the background. The first obstacle area consists of relatively static parts, such as cargo. This cargo might not change during a single crane operation, but multiple operations will show significant differences. Similarly, a building under repair may change over time. These relatively static parts can be considered obstacles, and the area containing them is called the first obstacle area. This is for subsequent identification purposes.
[0072] For the dynamic sub-model, the crane hook area or load area can be detected in real time, and the remaining part can be regarded as the second obstacle area. All parts of the dynamic sub-model other than the hook and load areas can be marked as the second obstacle area. The difference between the obstacles in the second obstacle area and those in the first obstacle area is that the obstacles in the former are moving, and can include, but are not limited to, workers, moving vehicles, and other machinery. For subsequent image data, the position and contour information of the hook / load at the previous moment is combined to track and predict the position and contour information of the hook / load at the current moment. This can be mainly achieved using an end-to-end visual object single-target tracker. Feature extraction and merging are performed using a hybrid attention mechanism between the load target in the previous frame and the current search area to obtain the correlation matching degree between the current area and the load target. When the matching degree is within the target area to be identified, point cloud data can be combined for more specific identification, thereby determining the spatial information of each target formation, such as position and spatial dimensions. The specific identification process will be described in detail later. By integrating image detection results, real-time detection and tracking of the hook / load can be achieved. This ensures the stability and robustness of the detection. Furthermore, as described above, the dynamic sub-model will be removed to avoid impacting the scene model construction; this is the operation corresponding to the filtering operation, which will not be elaborated upon here.
[0073] Based on the initial point cloud data and image data, a scene model is determined. For the scene model determined from the initial point cloud data, further semantic segmentation can be performed, including extracting / removing dynamic target points, extracting ground plane equations, and extracting the crane's own point cloud data. This initially distinguishes the scene models, facilitating subsequent region division. The representation of the scene model can include one or more of the following: ① Point cloud-based scene model, including the 3D spatial coordinates of all object surfaces in the scene, saving appropriate point cloud density as needed; ② Voxel-based scene model, including the probability of each spatial location being occupied and the distribution of point cloud within that voxel (e.g., represented by a 3D Gaussian model), selecting appropriate voxel size as needed; ③ Mesh-based scene model, using a series of polygons (usually triangles) of similar size and shape to approximate the 3D object model. Scene modeling improves the modeling accuracy of static scenes and helps enhance the perception of static obstacles.
[0074] In one embodiment, identifying a target object based on perception data includes: acquiring image data and point cloud data from the perception data, wherein the image data is acquired by a camera unit and the point cloud data is acquired by a lidar unit; entering an initialization phase based on the image data and point cloud data, including performing preprocessing operations on the point cloud data to obtain second point cloud data, wherein the preprocessing operations include at least one of coordinate transformation, multi-frame accumulation, filtering and noise reduction, ground segmentation, and region of interest filtering; processing the image data using an image semantic segmentation method to determine at least one target object image mask; extracting a first object point cloud using the second point cloud data corresponding to the target object image mask based on the extrinsic parameter relationship between the camera unit and the lidar unit; and determining at least one target object from the first object point cloud after multi-frame accumulation of the first object point cloud.
[0075] In one embodiment, identifying the target object based on the sensing data may specifically include two stages: an initialization stage and a dynamic tracking stage, which will be described in detail below.
[0076] The initialization phase begins, primarily involving point cloud extraction and target object region identification. Preprocessing operations are then performed on the point cloud data to obtain second point cloud data. These preprocessing operations include at least one of the following: coordinate transformation, multi-frame accumulation, filtering and denoising, ground segmentation, and region of interest filtering. Each preprocessing operation will be explained in detail below.
[0077] Coordinate transformation: Using the laser positioning result T, the coordinates of a single frame of the point cloud are transformed to obtain the point cloud in the vehicle coordinate system.
[0078] Multi-frame accumulation: Since single-frame point cloud data is relatively sparse and unevenly distributed, it may be impossible to detect the complete shape of some obstacles. By superimposing multiple frames of data for processing, a dense point cloud is generated, improving the point cloud coverage and supporting subsequent processing.
[0079] Filtering and denoising: Removing noise from point clouds and compressing the point cloud data. Several filters can be used, such as the SOR filter (Statistical Outlier Removal): This uses a statistical outlier filtering method that calculates the standard deviation and average distance between each point and its neighbors to determine if a point is noise. If the distance between a point and its neighbors exceeds a certain multiple of the standard deviation, the point is considered noise. Radius filters are also used to determine whether points around the point cloud are noise. Downsampling: For filtered point clouds with uneven density distribution and excessively high density in some areas, point cloud downsampling can be performed across the entire space to reduce data volume and improve processing efficiency while ensuring accuracy.
[0080] Ground segmentation: This can be achieved using either a piecewise planar model or a grid-based method. The piecewise planar model approach assumes the ground is planar within a certain range, then uses RANSAC or PCA principal component analysis to obtain normal vectors, which are then combined with height thresholds for ground identification. Algorithms such as cloth simulation can also be used, depending on the specific scenario, to improve adaptability to changes in slope and elevation. Finally, the boom height is estimated by combining the boom's positioning information and ground height information.
[0081] Region of interest filtering: In the vehicle coordinate system, the vehicle body range is fixed (length, width, and height), which belongs to the background point cloud. It can be filtered in advance based on the spatial coordinate range to reduce the amount of calculation.
[0082] The point cloud data after preprocessing is called the second point cloud data, which is a dense point cloud with a long cumulative time of multiple frames.
[0083] The image data is then processed using image semantic segmentation to determine at least one target object image mask. Based on the extrinsic parameter relationship between the camera unit and the lidar unit, the first object point cloud is extracted using the second point cloud data corresponding to the target object image mask. The specific recognition method can refer to the method for determining the region to be recognized from image data: identifying the hook, load, and obstacles one by one. The point cloud data recognition process for the target object can refer to... Figures 2a to 2g Thus, it can be determined that... Figure 2e and Figure 2g The hook area and obstacles are shown in the diagram. The specific identification process for each type of target object can be found in the following description.
[0084] Rope and hook identification: The side view obtained from the point cloud data is projected onto a 2D image to construct a binary map of the foreground and background. A Hough transform algorithm is used to detect lines on the binary map. The line closest to the current boom head is identified as the boom rope, and the other end of the rope is the endpoint of the hook top. After the rope has been detected, the point cloud below the rope is appropriately segmented using a height threshold (using prior attributes such as hook length, width, and height), thus completing the point cloud identification of the hook.
[0085] Load identification: Utilizing the extrinsic relationship between the lidar and the camera, the point cloud after removing the ground, hooks / wire ropes, etc., is projected onto the image to construct the correspondence between the point cloud and the image. Combining the position and contour information of the load in the image, the point cloud data falling within the contour range is obtained, i.e., the load point cloud.
[0086] The target pixel region is the area in the image where the hook and load are located. Because loads are diverse in type and shape, they cannot be directly detected using deep learning models. Therefore, manual interaction is necessary. The location or region of the load needs to be manually identified in the image. The identified location is at least one pixel within the load's range in the image; the region is the pixel outline or enclosing polygon outline of the load that can contain at least three points in the image. Taking load selection as an example, the determination method can be as follows: Figure 3a and Figure 3b As shown. Among them. Figure 3a The green dots in the image represent the load determined by pixels; approximately, for Figure 3b The green box in the image is defined by selecting the area. In practice, other selection and definition methods are possible; this is merely a simplified explanation and not a limitation. Through the above processing steps, the pixel region of the target object is finally determined. After determining the target object region, a precisely segmented mask is obtained, allowing for the continuation of the desired effect. Figure 3a and Figure 3b You can get Figure 4a and Figure 4b The mask shown in the image represents the portion of the target object that needs to be preserved for identification, while the remaining portion can be masked or disabled.
[0087] The identification of the area where the hook load is located has been described in detail in the previous section on the dynamic sub-model; please refer to that section for details. Here, we will only briefly describe it. Specifically, the wire rope is identified using a preset straight-line recognition method based on the installation position of the camera unit, and then the hook is identified based on the area at the end of the wire rope. Since hooks are similar in shape and relatively few in type, the position and outline of the hook can be obtained from a hook detection model generated through training with a certain number of samples. The outline of the hook should have multiple wire ropes intersecting with it; if not, it is not considered a hook.
[0088] Obstacle recognition: Clustering algorithms such as Euclidean clustering or density-based spatial clustering of applications with noise (DBSCAN) can be used to cluster point clouds into clusters. However, this process often results in over-segmentation and under-segmentation, requiring additional post-processing algorithms based on the actual situation, such as further splitting and merging based on a pre-designed two-dimensional projected area or by incorporating intensity information. At this point, apart from hooks, ropes, and loads, other obstacles do not yet have category information. By combining real-time load detection with a prior model, the accuracy of spatial position relationship calculation is improved; and various methods for obtaining the prior model and its information are proposed. Based on the extrinsic parameter relationship between the camera unit and the lidar unit, the first object point cloud is extracted using the second point cloud data corresponding to the target object image mask. The extrinsic parameter relationship has been explained in detail in the previous section on processing within the same spatiotemporal system, and will not be repeated here. For the first object point cloud, which is a dense point cloud, the acquisition effect can be referenced. Figure 5 As shown.
[0089] Furthermore, the position of subsequent target objects can be predicted. An end-to-end visual object single-target tracker can be used, employing a hybrid attention mechanism between the payload target in the previous frame and the current search region for feature extraction and merging. This obtains the correlation matching degree between the current region and the payload target. When the matching degree reaches a benchmark, the region with the highest matching degree is selected as the predicted bounding box for output. The tracker's processing flow belongs to the dynamic tracking stage, which will be used to more accurately identify target objects or determine perceptual enhancement information. For ease of explanation, this will be elaborated in the subsequent section on obtaining perceptual enhancement information.
[0090] Step S130: Obtain and output the perception enhancement information between the scene model and the target object.
[0091] In one embodiment, it is understood that the target object is in continuous motion and change, and the enhanced perception information between target objects is also dynamically acquired. This enhanced perception information includes, but is not limited to: 1. hook / load pose and 3D contour; 2. distance relative to obstacles; 3. boom pose; 4. boom height above ground; 5. load height; 6. boom deflection and amplitude; 7. boom-load wire rope length, etc. Enhanced perception information can be used for, but is not limited to: 1. outputting and displaying distance information; 2. providing multi-level early warnings based on distance information; 3. supporting the operation of other subsequent control systems. For example, providing load height information for load translation; providing load position information for load sway control; providing collision information for active safety, etc.
[0092] To acquire enhanced perception information, image data and point cloud data can be dynamically tracked and updated. Specifically, the dynamic tracking phase begins with the first object point cloud, involving accumulating the first object point cloud over multiple frames and dynamically tracking and determining the target object's pixel region based on image data. Point cloud data within the target object's pixel region is then acquired and labeled as the second object point cloud. The first and second object point clouds are then registered. Registration methods include ICP or NDT, among others, with no specific limitation. The registered first object point cloud replaces the second object point cloud to obtain the third object point cloud. Registration results in higher point cloud density, richer point cloud information, and higher accuracy in subsequent information extraction. Enhanced perception information is then acquired from the third point cloud. The process for acquiring the first object point cloud, i.e., determining the target object's pixel region, can be found in the previous text and will not be repeated here.
[0093] Perceptual enhancement information is primarily acquired through scene models and target objects, with the target objects specifically represented as third-point clouds. The relationship between the two can be found in [reference needed]. Figure 6 As shown in the previous embodiments, taking the acquisition of load as an example, wherein... Figure 6 The red portion represents the payload, specifically the third point cloud; the remaining green portion represents the scene model, specifically the static sub-model portion within the scene model. By understanding the relationship between these two, the various types of perception enhancement information mentioned earlier can be determined. Furthermore, it should be noted that perception enhancement information can also be obtained independently using the scene model or the target object, or even directly without these two, such as boom height or crane height.
[0094] To continuously track and identify target objects, a tracking model can be used for each target object, and all tracking models form a tracking queue. When a new target object is identified, each object in the tracking queue first uses the tracking model to predict its current position and velocity. Then, the KM algorithm (Kuhn-Munkres Algorithm) can be used to associate the tracking observations with the tracking model. Finally, the tracking observations are used to correct the predicted values of the corresponding tracked target and update the model.
[0095] The tracking model can specifically be a Kalman Filter (KF), with each target object tracked using a different KF model. When a new target object is detected, a KF model is set to track it, and the tracking queue is updated. Tracking data in the tracking queue is continuously acquired to generate distance information. When tracking is lost for a period of time / a specified number of frames, the KF model corresponding to the lost target object is cleared, thereby dynamically updating the tracking list.
[0096] Furthermore, this application can identify more than one target object, thus enabling multi-target tracking and matching. The current target object (one or more of the following: hook, load, obstacle) forms a source group, referred to as the src group; the targets in the tracking queue form a target group, referred to as the tar group. The tracking process iterates through the targets within the tar group. i Calculate tar i Predicted value at the current moment In src Target src within the neighborhood j , remember dist i,j Let Euclidean distance be the distance between the two targets, and calculate it:
[0097] weight i,j =-dist i,j -1 (3)
[0098] Not in src Target src within the neighborhood k , denoted as:
[0099] weight i,k =MIN_EXPECT (4)
[0100] Among them, MIN_EXPECT and MAX_EXPECT are the set expected boundaries, which should be set reasonably according to the task to avoid overflow. Based on the obtained {weight i,j The matrix is used to obtain tracking data that includes the correlation between targets in src and targets in tar using the minimum weight KM algorithm.
[0101] For the extracted hook / load, since the real-time acquired point cloud data is relatively sparse and cannot obtain complete object surface data, a more complete (accurate) spatial occupancy range of the hook / load can be obtained by registering the pre-acquired hook / load model with the real-time extracted hook / load point cloud. The registration method includes ICP (Iterative Closest Point) or NDT (Normal Distributions Transform), etc. For example, using the load mentioned earlier, a diagram illustrating the real-time extraction of the load from the third point cloud based on image masking can be found here. Figure 7 As shown; real-time matching of sparse payload point clouds to third point clouds can be referenced. Figure 8 As shown, the colored part is the upper surface of the identified load, which can represent the load being tracked and identified.
[0102] Methods for obtaining dimensional information between the hook and load include: importing known 3D models, such as design models (drawings), using measuring tools, such as handheld automatic 3D scanners, to perform on-site scanning and automatically generate a model; using simple distance measuring tools for on-site measurement; and collecting and inputting the basic dimensions (length, width, height, etc.) of the hook load. This also includes automatic detection to extract point cloud data from the background area of the static sub-model to indicate the ground; extracting the upper surface information of the load in the target object, as well as the load's position range in map space. The distance from the upper surface height of the load to the ground is calculated as the load height, and the load can meet certain conditions, such as being lifted from level ground, being a rigid object, or remaining horizontal in the air. After lifting the load, the upper surface height of the load at the moment it leaves the ground is obtained, which also needs to meet certain conditions: for example, the stabilization trend after the force limiter's weight increases, in conjunction with the winch-lamp adjustment. During the lifting process, the distance from the upper surface height of the load to the ground is calculated as the load height, and the load height is updated. After the load has completely left its original position, calculate the height of the upper surface of the original position range. Calculate the difference between the height of the upper surface of the load at the moment it leaves the ground and the height of the upper surface of the original position range as the load height. Update the load height. This method is applicable to loads under other conditions, such as: lifting on non-flat ground, non-horizontal objects in the air, and non-rigid objects.
[0103] In one embodiment, acquiring and outputting enhanced perception information between the scene model and the target object includes: acquiring the relative distance in the enhanced perception information, and a first warning threshold and a second warning threshold corresponding to the relative distance, wherein the relative distance is the distance between the target object and static obstacles and / or dynamic obstacles; if the relative distance is less than the first warning threshold and greater than the second warning threshold, then outputting a first alarm message, which is used to alert the user that there is a collision risk; if the relative distance is less than the second warning threshold, then outputting a second warning message and / or controlling the crane to stop moving, wherein the second alarm message is used to alert the user that the load will collide.
[0104] In one embodiment, for multi-level early warning using distance information from enhanced perception information, warnings or controls can be implemented based on different alert thresholds. The distance information may specifically include information such as the position, outline, and size of the hook, load, and obstacles. A schematic diagram of the distance information output can be found in [reference needed]. Figure 9 As shown.
[0105] The warning system can output not only the distance information of the target object itself, but also the relationship between target objects and each other, as well as with the surrounding environment. The primary information is the relative distance between the load and obstacles, referred to as the relative distance. The relative distance indicates the distance and distance relationship between the load and various directions / surrounding areas in the environment. Risk levels can also be categorized based on distance, represented by different colors. Risk levels are determined by the relationship between a preset first warning threshold (C1, for example, 1m) and a second warning threshold (C2, for example, 0.5m). When the relative distance is greater than C1, it is displayed in green, indicating no collision risk in that direction; when it is less than C1 but greater than C2, it is displayed in yellow, outputting the first warning message to alert the user of a potential collision risk; when it is less than or equal to C2, it is displayed in red, outputting the second warning message. These warnings alert the user that a collision between the load and an obstacle is imminent, and can also control the crane to stop moving or perform a slight reverse movement to avoid a collision. For the output display effect of this implementation, please refer to [reference needed]. Figure 10 .like Figure 10 As shown, the output can include the nearest distance in different directions. For example, left and right represent the turning direction, with left for left turn and right for right turn; up and down represent the luffing direction, with up for lowering and down for raising. Distance information can also include other data, such as the boom height above the ground (where the crane supports the ground), the hook / load height above the ground (where the crane supports the ground), and the length of the wire rope from the boom to the hook.
[0106] Therefore, this application can acquire diverse sensing data through a sensing kit installed at a predetermined location on the crane to construct a scene model of the crane's location, facilitating the identification of target objects, including hooks, loads, and / or obstacles. This allows for the perception of the crane's working environment, determining corresponding distance information based on the correlation between target objects and the scene model, and outputting this information. This enables the provision of real-time images of the hook / load's surroundings and enhanced perception information about the surrounding environment, including distance and risk warnings. It supports upgrades to active safety braking control. It allows crane operators to achieve safe lifting beyond visual range, significantly improving the convenience and safety of lifting operations. Furthermore, the method provided in this application is applicable to any type of hook and load, enhancing the method's adaptability and detection accuracy.
[0107] Figure 11 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 11As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a crane working environment perception method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement the crane working environment perception method. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0108] In one embodiment, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in any of the foregoing embodiments.
[0109] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0111] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for sensing the working environment of a crane, characterized in that, Includes the following steps: The process involves acquiring sensing data collected by a sensing kit, which includes at least one of a lidar unit, an inertial measurement unit, and a camera unit, and is installed at a predetermined position on the crane. The acquisition of the sensing data includes: acquiring inertial data from the sensing data, the inertial data being acquired by the inertial measurement unit; calculating a motion change based on the inertial data, the motion change describing the motion change of the sensing kit during the acquisition process; acquiring initial sensing data collected by the sensing kit; performing an error compensation operation on the initial sensing data based on the motion change to correct errors in the initial sensing data; when the initial sensing data after the error compensation operation meets the convergence condition, marking the initial sensing data as the sensing data; otherwise, repeating the error compensation operation on the initial sensing data. A scene model is established based on the perceived data, and target objects are identified. The scene model includes static obstacles and / or dynamic obstacles, and the target objects include hooks and / or loads. Obtain and output the perception enhancement information between the scene model and the target object.
2. The crane working environment perception method as described in claim 1, characterized in that, The acquisition of sensory data collected by the sensory kit includes: Based on the predetermined position, the relative positional relationship between the sensing kits is determined using an extrinsic parameter calibration method, and the relative positional relationship is represented by an extrinsic parameter matrix. The perceived data are unified to the same spatial coordinate system based on the extrinsic parameter matrix. The timestamp information of the sensing data is calibrated so that the sensing data collected by each sensing kit is located under a unified time reference system.
3. The crane working environment perception method as described in claim 1, characterized in that, The step of establishing a scene model based on the perceived data includes: The image data and point cloud data in the perception data are acquired, wherein the image data is acquired by the camera unit and the point cloud data is acquired by the lidar unit; The point cloud data within a specified time range is transformed to determine the position coordinates of all object surfaces in the scene where the crane is located. The point cloud data after the pose relationship transformation is filtered out, and the point cloud data after the filtering operation is performed is spliced and accumulated to obtain the first point cloud data. The scene model is determined based on the first point cloud data and the image data, wherein the scene model does not include the target object.
4. The crane working environment perception method as described in claim 3, characterized in that, The filtering operation includes: All image data within a given time period are processed using a preset straight line detection method, and the processed straight lines are marked as the wire rope of the crane; the hook area at one end of the wire rope is identified based on the image data; Obtain cargo information, and determine the load area based on the cargo information and the hook area; The point cloud data in the hook area and the load area are filtered out.
5. The crane working environment perception method as described in claim 3, characterized in that, The step of establishing a scene model based on the perceived data includes: An initial scene model is constructed based on the first point of cloud data; Acquire all the image data within a preset time period, and identify static sub-models and dynamic sub-models from the initial scene model based on the change process of the image data. The static sub-model is the stationary part of the initial scene model, including a background area and a first obstacle area. The dynamic sub-model is the moving part of the initial scene model, including a hook area, a load area, and a second obstacle area. The static obstacle is identified from the static sub-model; the hook area and the load area are deleted from the dynamic sub-model, and the dynamic obstacle is identified from the deleted second obstacle area; The static obstacles and / or the dynamic obstacles are combined to obtain the scene model.
6. The crane working environment perception method as described in claim 1, characterized in that, The step of identifying the target object based on the perceived data includes: The image data and point cloud data in the perception data are acquired, wherein the image data is acquired by the camera unit and the point cloud data is acquired by the lidar unit; The initialization phase is initiated based on the image data and the point cloud data, including performing preprocessing operations on the point cloud data to obtain second point cloud data. The preprocessing operations include at least one of coordinate transformation, multi-frame accumulation, filtering and noise reduction, ground segmentation, and region of interest filtering. The image data is processed by image semantic segmentation to determine at least one target object image mask; based on the external parameter relationship between the camera unit and the lidar unit, the second point cloud data corresponding to the target object image mask is used to extract the first object point cloud; after accumulating the first object point cloud over multiple frames, at least one target object is determined from the first object point cloud; The step of acquiring and outputting the perception enhancement information between the scene model and the target object includes: The dynamic tracking stage is initiated based on the first object point cloud, which includes accumulating the first object point cloud over a multi-frame time period and dynamically tracking and determining the target object pixel region based on the image data. Obtain point cloud data within the pixel region of the target object and mark it as the second object point cloud; register the first object point cloud with the second object point cloud, and replace the second object point cloud with the registered first object point cloud to obtain the third object point cloud; The perception enhancement information is obtained based on the point cloud of the third object.
7. The crane working environment perception method as described in claim 1, characterized in that, The step of acquiring and outputting the perception enhancement information between the scene model and the target object includes: The relative distance in the perception enhancement information is obtained, as well as the first warning threshold and the second warning threshold corresponding to the relative distance, wherein the relative distance is the distance between the target object and the static obstacle and / or the dynamic obstacle; If the relative distance is less than the first warning threshold and greater than the second warning threshold, then a first alarm message is output, which is used to alert the user that there is a risk of collision. If the relative distance is less than the second warning threshold, a second alarm message is output and / or the crane is controlled to stop moving. The second alarm message is used to alert the user that the load will collide.
8. A computer device, characterized in that, Including processor and memory; The processor is used to execute a computer program stored in the memory to implement the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Bridge crane hoisting safety anti-collision system and method based on dynamic binocular vision
CN112418103A
Tower crane hoisting object identification and collision information measurement system and method
CN113860178A