Method, device and equipment for automatic feeding and discharging of profile based on visual recognition
By integrating 3D point cloud data and 2D image data in multiple dimensions, and combining the pose data of the loading and unloading actuator, the system achieves accurate identification and stable gripping of profiles, solving the efficiency and robustness issues of automatic loading and unloading after profile cutting, and is suitable for complex industrial environments.
Patent Information
- Application Number
- CN202511787241.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-12-01
AI Technical Summary
Existing technologies struggle to achieve efficient and stable automatic loading and unloading after profile cutting, especially in complex scenarios where recognition robustness and grasping success rate are insufficient, affecting the overall operational efficiency of the production line.
An automatic profile loading and unloading method based on vision recognition is adopted. By multi-dimensional fusion of 3D point cloud data and 2D image data, the tilt angle and size information of the profile are identified, and combined with the pose data of the loading and unloading actuator, the precise loading and unloading actions are realized.
It improves the robustness of profile recognition and the success rate of grasping, and can effectively cope with complex scenarios such as profile stacking, partial occlusion and posture changes, thereby improving the flexibility and efficiency of automated operations.
Smart Images

Figure CN121213332B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual recognition technology, and in particular to a method, apparatus and equipment for automatic loading and unloading of profiles based on visual recognition. Background Technology
[0002] Driven by the continuous development of intelligent manufacturing, industrial automation and intelligence have become key directions for improving enterprise production efficiency and processing precision. For example, in industries such as shipbuilding and new energy equipment, current industrial profile cutting equipment has achieved a high level of intelligence, capable of quickly responding to automatic cutting tasks of multi-segment profiles, significantly improving production efficiency and processing flexibility. These advanced devices, through CNC and automated programming, have achieved high-precision, high-speed cutting of profiles of different sizes and shapes, providing a solid foundation for automated loading and unloading and assembly operations. As production lines increasingly demand flexible manufacturing and unmanned operation, how to achieve efficient and stable automated loading and unloading processes after cutting has become one of the key links in improving the overall intelligence level of the production line. Moreover, the level of intelligence in the loading and unloading process directly affects the operating efficiency and cycle time of the entire production line. Especially with the increasing prevalence of precision machining processes such as laser cutting and high-speed milling, how to efficiently and accurately automate the auxiliary profile cutting machine and achieve automatic profile identification and loading / unloading tasks has become a core technical challenge in the process of industrial automation upgrading. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, apparatus and equipment for automatic loading and unloading of profiles based on visual recognition, which can not only realize automatic loading and unloading of profiles, but also greatly improve the robustness of profile recognition and the success rate of grasping.
[0004] In a first aspect, the present invention provides a method for automatic loading and unloading of profiles based on visual recognition, comprising:
[0005] The visual image capture points are determined based on the attribute information of the profiles to be sorted.
[0006] According to at least one target visual photography point, collect the profile visual data corresponding to the profile to be sorted, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result corresponding to the profile to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information corresponding to the profile to be sorted.
[0007] Based on the visual recognition results, the target pose data corresponding to the loading and unloading actuator is determined, so as to control the loading and unloading actuator to perform loading and unloading actions on the profiles to be sorted according to the target pose data.
[0008] In one implementation, visual data of the profiles to be sorted is collected according to at least one target visual imaging point, and profile recognition processing based on multi-dimensional visual fusion is performed on the data to obtain the visual recognition result of the profiles to be sorted, including:
[0009] Collect visual data of the profiles to be sorted from a target visual photography point.
[0010] The integrity of the profiles to be sorted displayed in the profile visual data is checked. If the integrity check fails, the profile visual data corresponding to the profiles to be sorted is collected according to multiple target visual photography points.
[0011] For the visual data of profiles under at least one target visual capture point, perform profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted.
[0012] In one implementation, the visual data of profiles at at least one target visual capture point is subjected to profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted, including:
[0013] The following processing is performed on the profile visual data under each target visual capture point: based on 3D point cloud data, or based on 3D point cloud data and 2D image data, the visual recognition result corresponding to the profile to be sorted is identified.
[0014] In one implementation, the visual recognition result corresponding to the profile to be sorted is identified based on 3D point cloud data, or based on 3D point cloud data and 2D image data, including:
[0015] Based on 3D point cloud data, the profile to be sorted and its sorting area are segmented to obtain the first segmentation result corresponding to the profile to be sorted.
[0016] Determine whether the first segmentation result corresponding to the profile to be sorted meets the preset conditions;
[0017] If so, then based on the first segmentation result, identify the visual recognition result corresponding to the profile to be sorted;
[0018] If not, the 2D image data is used to segment the profile to be sorted and the sorting area in which it is located to obtain the second segmentation result corresponding to the profile to be sorted. The second segmentation result and the 3D point cloud data are combined to identify the visual recognition result corresponding to the profile to be sorted.
[0019] In one implementation, identifying the visual recognition result corresponding to the profile to be sorted includes:
[0020] Based on the edge of the profile to be sorted, characterized by the first segmentation result or the second segmentation result, determine the unilateral information of the profile to be sorted.
[0021] Based on the unilateral information of the profile to be sorted and the depth information described by 3D point cloud data, the surface unevenness information of the profile to be sorted is determined.
[0022] Based on the unilateral information and surface unevenness information of the profile to be sorted, the visual recognition result corresponding to the profile to be sorted is determined.
[0023] In one implementation, determining the target pose data corresponding to the loading / unloading actuator based on visual recognition results includes:
[0024] The tilt angle and size information contained in the visual recognition results are transformed from the camera coordinate system to the actuator coordinate system;
[0025] Based on the hand-eye calibration results, the tilt angle and size information in the actuator coordinate system, determine the positional deviation of the loading and unloading actuator;
[0026] Based on the pose deviation, the current pose data of the loading / unloading actuator is adjusted to obtain the target pose data of the loading / unloading actuator.
[0027] Secondly, the present invention also provides a visual recognition-based automatic profile loading and unloading device, comprising:
[0028] The location determination module is used to determine the visual image capture location based on the attribute information of the profile to be sorted;
[0029] The visual recognition module is used to collect the visual data of the profile to be sorted at at least one target visual photography point, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result of the profile to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information of the profile to be sorted.
[0030] The loading and unloading module is used to determine the target pose data of the loading and unloading actuator based on the visual recognition results, so as to control the loading and unloading actuator to perform loading and unloading actions on the profiles to be sorted according to the target pose data.
[0031] In one implementation, the visual recognition module is specifically used for:
[0032] Collect visual data of the profiles to be sorted from a target visual photography point.
[0033] The integrity of the profiles to be sorted displayed in the profile visual data is checked. If the integrity check fails, the profile visual data corresponding to the profiles to be sorted is collected according to multiple target visual photography points.
[0034] For the visual data of profiles under at least one target visual capture point, perform profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted.
[0035] Thirdly, the present invention also provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.
[0036] Fourthly, the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.
[0037] This invention provides a method, apparatus, and device for automatic loading and unloading of profiles based on visual recognition. First, visual image capture points are determined based on the attribute information of the profiles to be sorted. Then, visual data of the profiles to be sorted is collected according to at least one target visual image capture point, and multi-dimensional visual fusion-based profile recognition processing is performed to obtain the visual recognition result corresponding to the profiles to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information of the profiles to be sorted. Finally, the target pose data corresponding to the loading and unloading actuator is determined based on the visual recognition result, so as to control the loading and unloading actuator to perform loading and unloading actions on the profiles to be sorted according to the target pose data. The above method collects 2D image data and 3D point cloud data of the profile to be sorted by determining at least one target visual imaging point based on attribute information, and obtains the corresponding visual recognition result through profile recognition processing based on multi-dimensional visual fusion. On this basis, the target pose data corresponding to the loading and unloading execution mechanism is determined, so as to control the loading and unloading execution mechanism to perform loading and unloading actions for the profile to be sorted. This invention can not only accurately restore the spatial position and geometric contour of the workpiece, but also has stronger adaptability. It can effectively deal with complex scenarios such as dense profile stacking, partial occlusion, and posture changes, and greatly improve the recognition robustness and grasping success rate, providing a solid guarantee for flexible production and automated operation in intelligent manufacturing.
[0038] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0039] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0040] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating a method for automatic loading and unloading of profiles based on visual recognition, provided in an embodiment of the present invention;
[0042] Figure 2 A technical framework diagram of a method for automatic loading and unloading of profiles based on visual recognition provided in an embodiment of the present invention;
[0043] Figure 3 This is a schematic diagram of profile area analysis provided in an embodiment of the present invention;
[0044] Figure 4 This is a schematic diagram illustrating the calculation of profile concavity / convexity information based on the X-direction of an image, provided by an embodiment of the present invention.
[0045] Figure 5 This is a schematic diagram illustrating the fusion of 2D and 3D visual recognition for profile grasping information, provided in an embodiment of the present invention.
[0046] Figure 6 This is a schematic diagram illustrating a method for identifying long profiles at different locations, provided by an embodiment of the present invention.
[0047] Figure 7 A schematic diagram of the overall process of a method for automatic loading and unloading of profiles based on visual recognition provided in an embodiment of the present invention;
[0048] Figure 8 A schematic diagram of the overall process of another method for automatic loading and unloading of profiles based on vision recognition provided in an embodiment of the present invention;
[0049] Figure 9 A schematic diagram of a device for automatic loading and unloading of profiles based on visual recognition, provided in an embodiment of the present invention;
[0050] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Currently, in industries such as shipbuilding and new energy equipment, industrial profile cutting machines have achieved a high degree of automation, enabling them to quickly and efficiently complete the automatic cutting of multiple profile segments. To meet the demands for high-efficiency, unmanned automatic profile cutting and loading / unloading operations, this invention provides a method, apparatus, and equipment for automatic profile loading / unloading based on vision recognition. This not only enables automatic profile loading / unloading but also significantly improves the robustness of profile recognition and the success rate of grasping.
[0053] To facilitate understanding of this embodiment, a method for automatic profile loading and unloading based on vision recognition, disclosed in this embodiment of the invention, will first be described in detail. (See [link to relevant documentation]). Figure 1 The diagram shows a flow chart of a method for automatic loading and unloading of profiles based on vision recognition. The method mainly includes the following steps S102 to S106:
[0054] Step S102: Determine the visual image capture points based on the attribute information of the profiles to be sorted.
[0055] The profiles to be sorted can be ship profiles (1-12 meters), and the attribute information can include the type of profile to be sorted and its initial length, width, and height information. The visual capture point refers to the location of the visual device (such as a 3D large-view structure camera) when collecting data. In one embodiment, multiple visual capture points can be determined based on the field of view of the visual device and the initial length, width, and height information.
[0056] Step S104: Collect visual data of the profiles to be sorted according to at least one target visual photography point, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result of the profiles to be sorted.
[0057] The profile visual data includes 2D image data and 3D point cloud data. The visual recognition result includes the tilt angle and size information of the profile to be sorted, which is the length, width, and height of the profile to be sorted obtained through profile recognition processing. In one embodiment, visual images are taken at preset rules. First, a target visual image point is determined. It is then determined whether the profile visual data collected at the target visual image point can cover the entire profile to be sorted. If it can, the visual recognition result is directly determined using the profile visual data. If not, all visual image points are determined as target visual image points, and the corresponding profile visual data is collected sequentially. Finally, the visual recognition result is determined by combining the profile visual data from all target visual image points.
[0058] Step S106: Determine the target pose data corresponding to the loading / unloading actuator based on the visual recognition results, so as to control the loading / unloading actuator to perform loading / unloading actions on the profiles to be sorted according to the target pose data.
[0059] The loading / unloading mechanism can be a gripper / robotic arm that supports telescopic rotation. The target pose data includes at least the rotation angle, surface angle, gripping height, gripping point, and magnetic activation point of the loading / unloading mechanism. In one example, the pose offset can be determined based on hand-eye calibration results, the tilt angle and dimensions of the profile to be sorted, and this offset can be used to adjust the current pose data of the loading / unloading mechanism, enabling it to perform loading / unloading actions according to the adjusted target pose data.
[0060] The method for automatic loading and unloading of profiles based on visual recognition provided in this invention utilizes 3D+2D multimodal vision fusion technology. It can simultaneously acquire the three-dimensional structural information and two-dimensional texture features of the workpiece, providing more robust and accurate technical support for workpiece recognition, posture estimation, and automatic grasping in complex industrial environments. Compared with traditional recognition methods that rely solely on 2D images, 3D vision can not only accurately reconstruct the spatial position and geometric contour of the workpiece but also possesses stronger adaptability. It can effectively cope with complex scenarios such as densely stacked profiles, partial occlusion, and posture changes, significantly improving recognition robustness and grasping success rate, providing a solid guarantee for flexible production and automated operations in intelligent manufacturing.
[0061] In one implementation, see Figure 2 The diagram illustrates a technical framework for a vision-based automatic profile loading and unloading method, including processes such as 3D depth information analysis, combined with 2D image information analysis, and vision-guided scaffolding / robotic arm gripper positioning. For ease of understanding, this invention provides a specific implementation of the vision-based automatic profile loading and unloading method:
[0062] Regarding the aforementioned step S104, this embodiment of the invention provides an implementation method for collecting visual data of the profiles to be sorted at at least one target visual imaging point, performing profile recognition processing based on multi-dimensional visual fusion, and obtaining the visual recognition result corresponding to the profiles to be sorted, including:
[0063] (1A) Collect visual data of the profiles to be sorted according to a target visual imaging point. In one example, first control only the 3D large field of view structure camera to collect 2D image data and 3D point cloud data according to a point.
[0064] (1B) Perform an integrity check on the profiles to be sorted displayed in the profile visual data. If the integrity check fails, collect the profile visual data corresponding to the profiles to be sorted at multiple target visual photography points. The integrity check is to check whether the profile visual data collected in (1A) can cover the entire profile to be sorted. If it covers the entire profile to be sorted, the integrity check is passed. If it does not cover the entire profile to be sorted, the integrity check fails. In this case, the corresponding profile visual data will be collected at all visual photography points.
[0065] (1C) Perform profile recognition processing based on multi-dimensional visual fusion on the profile visual data under at least one target visual photography point to obtain the visual recognition result corresponding to the profile to be sorted.
[0066] Specifically, the visual data of the profiles at each target visual capture point is processed as follows: Based on 3D point cloud data, or based on 3D point cloud data and 2D image data, the visual recognition result corresponding to the profile to be sorted is identified. In specific implementation, if the visual data of the profile captured in a single shot passes the integrity check (i.e., a single shot can cover the complete profile to be sorted), then only the visual data of the profile captured in this shot needs to be processed to obtain the corresponding visual recognition result; if the visual data of the profile captured in a single shot fails the integrity check (i.e., a single shot does not cover the complete profile to be sorted), for example, due to the limitation of the field of view of the 3D camera (the profiles are long, for example, shipyard profiles are mostly in the range of 6~16m), in order to accurately identify the offset information of long profiles, it is necessary to process the visual data of the profiles at multiple capture points to obtain the corresponding visual recognition result. It should be noted that, regardless of whether the visual data of the profile captured in a single shot is processed or the visual data of the profiles at multiple capture points is processed, the processing can include the following steps:
[0067] (1) Based on 3D point cloud data, the profile to be sorted and its sorting area are segmented to obtain the first segmentation result corresponding to the profile to be sorted.
[0068] For example, a recognition method based on 3D point cloud data sets an ROI region according to the sorting area, and within the set area, uses pass-through filtering and region growing algorithms based on the depth information of 3D point cloud data to achieve segmentation between the profile to be sorted and the sorting area, thus obtaining the first segmentation result.
[0069] (2) Determine whether the first segmentation result corresponding to the profile to be sorted meets the preset conditions. The preset conditions may be that the first segmentation result is not obscured or missing, that is, if the first segmentation result is obscured by other objects or a part of the area is missing, it does not meet the preset conditions.
[0070] (3) If so, then based on the first segmentation result, identify the visual recognition result corresponding to the profile to be sorted.
[0071] In one example, the process of identifying the visual recognition result corresponding to the profile to be sorted based on the first segmentation result is as follows: Based on the edge of the profile to be sorted represented by the first segmentation result, the one-sided information of the profile to be sorted is determined; based on the one-sided information of the profile to be sorted and the depth information described by the 3D point cloud data, the surface concavity and convexity information of the profile to be sorted is determined; based on the one-sided information of the profile to be sorted and the surface concavity and convexity information, the visual recognition result corresponding to the profile to be sorted is determined. Here, the one-sided information is also the length, width, and height information.
[0072] For example, taking bulb flat steel profiles as an example, see [link to relevant documentation]. Figure 3 The diagram illustrates a profile area analysis: the left image shows the segmentation and identification of the feeding bulb flat steel profile. By combining depth information from 3D point cloud data, the single-sided information (length, width, and height) of the bulb flat steel profile is obtained. The single-sided information proposed in this embodiment is obtained by setting a Region of Interest (ROI) within the feeding area and acquiring the lateral information of each identified profile (…). Figure 3 The material identification process (based on the x-direction of the image) calculates the long side, short side, and height information of the bulb flat steel profile in the x-direction.
[0073] See Figure 4 The diagram illustrates a method for calculating profile concavity / convexity information based on the X-axis of an image. Based on the single-sided information of the profile (image X-axis), and combined with the depth information of 3D point cloud data, this embodiment of the invention proposes a method for obtaining concavity / convexity information at different heights along the X / Y axes of an image / point cloud, tailored to specific profile types, by combining depth information. For example, this concavity / convexity information could be the concavity / convexity information of the plane formed by the single-sided edge on the left side and the single-sided edge at the center of a bulb flat steel profile.
[0074] The tilt angle of the spherical flat steel profile can be analyzed by combining the unilateral and concave / convex information of the photo-taking point (in conjunction with the basic coordinate system of the robotic arm).
[0075] (4) If not, then based on the 2D image data, the profile to be sorted and its sorting area are segmented to obtain the second segmentation result corresponding to the profile to be sorted. Then, combining the second segmentation result with the 3D point cloud data, the visual recognition result corresponding to the profile to be sorted is identified, such as... Figure 5 The diagram shows a method for integrating 2D and 3D visual recognition to capture information from profiles.
[0076] In one example, the process of segmenting the profile to be sorted and its sorting area based on 2D image data to obtain the second segmentation result corresponding to the profile to be sorted is shown below. The YOLOv8-seg model can output the corresponding detection box, instance segmentation mask and category, etc., based on the input 2D image data and 3D point cloud data.
[0077] In one example, the process of identifying the visual recognition result corresponding to the profile to be sorted by combining the second segmentation result and 3D point cloud data is as follows: Based on the edge of the profile to be sorted represented by the second segmentation result, the one-sided information of the profile to be sorted is determined; based on the one-sided information of the profile to be sorted and the depth information described by the 3D point cloud data, the surface concavity and convexity information of the profile to be sorted is determined; based on the one-sided information of the profile to be sorted and the surface concavity and convexity information, the visual recognition result corresponding to the profile to be sorted is determined. For details, please refer to the foregoing embodiments; the embodiments of this invention will not elaborate further.
[0078] The method provided in this invention fully combines the advantages of 3D point cloud data in spatial structure perception with the ability of 2D image data in texture feature extraction. It can effectively identify profiles in different postures and stacking states, achieving stable and accurate disordered grasping and intelligent sorting, significantly improving the flexibility and operational efficiency of automated loading and unloading systems. Furthermore, by fusing visual information from multiple shooting points, the system can accurately estimate the spatial posture of the workpiece, thereby achieving precise identification of the profile's position and dynamic correction of the scaffolding grasping angle. This provides a more reasonable grasping posture plan for the scaffolding mechanism, effectively improving the stability and success rate of grasping.
[0079] Regarding the aforementioned step S106, this embodiment of the invention also provides an implementation method for determining the target pose data corresponding to the loading / unloading actuator based on the visual recognition result, including: converting the tilt angle and size information contained in the visual recognition result from the camera coordinate system to the actuator coordinate system; determining the pose deviation amount corresponding to the loading / unloading actuator based on the hand-eye calibration result and the tilt angle and size information in the actuator coordinate system; and adjusting the current pose data corresponding to the loading / unloading actuator based on the pose deviation amount to obtain the target pose data corresponding to the loading / unloading actuator.
[0080] In practical implementation, the target pose data is determined by combining the visual recognition results obtained from a single or multiple camera points with the hand-eye calibration results of the loading / unloading actuator. For example, to determine the target pose data based on visual recognition results from multiple camera points and the hand-eye calibration results of the loading / unloading actuator, see [link to relevant documentation]. Figure 6 The diagram shows a method for identifying long profiles at different locations. Figure 6 The long side, short side, highest point, and overall offset angle of the profile are converted into world coordinates for the scaffold / robotic arm gripper, which then uses these coordinates to grasp the profile. Considering the length of the profile, if the entire profile can be captured in a single image view, the overall deflection angle to be grasped can be calculated using this single image view (to provide the T-axis rotation for the scaffold / robotic arm); and the rotational offset for each point of contact is determined based on the information of each ROI.
[0081] This invention can adaptively calculate the surface height (H), tilt angle, gripper gripping height, and profile type of different profiles. This method integrates multimodal perception data such as depth information, geometric features, and topological structure to perform refined analysis of point clouds, achieving highly robust identification and parameter extraction for profiles of various postures and types. This provides reliable data support for subsequent gripping planning and path generation, significantly improving the system's automated adaptability under complex working conditions.
[0082] This invention further provides an application example of a vision-based method for automatic loading and unloading of profiles, taking bulb flat steel profiles as an example. (See attached image.) Figure 7The diagram illustrates the overall process of an automatic profile loading and unloading method based on visual recognition, including: Step 1: Calculate visual image capture points once or twice based on the profile length and visual range in the attribute information of the bulb flat steel profile (the incoming material error needs to be limited). Taking ship profiles as an example, the system can obtain attribute information issued by the shipyard and determine the visual image capture points by combining it with the field of view of a 3D large-view structure camera. Step 2: Determine whether all images can be captured at once. If yes, proceed to Step 3; otherwise, proceed to Step 4. Step 3: Use profile recognition processing based on multi-dimensional visual fusion (also known as multimodal analysis / depth information ROI block recognition method) to obtain visual recognition results. Step 4: Use profile recognition processing based on multi-dimensional visual fusion (also known as multimodal analysis / depth information ROI block recognition method) to obtain visual recognition results for image capture points 1 and 2 respectively. Here, image capture points 1 and 2 can be understood as determining image capture points at least twice; the first time, only one image capture point is determined, and the second time, multiple image capture points are determined. Step 5: Analyze and identify the tilt angle, face type, and angle of the bulb flat steel profile. Step 6: Based on the tilt angle, face type, and angle of the bulb flat steel profile, control the loading / unloading mechanism (i.e., the scaffold / robotic arm gripper) to perform rotation adjustment, face angle adjustment, and gripping height adjustment.
[0083] Based on this, the embodiments of the present invention further provide a specific implementation method when the segmentation of profiles and sorting areas fails solely relying on 3D point cloud data, see [link to implementation details]. Figure 8 The diagram illustrates the overall process of another vision-based automatic profile loading and unloading method. 2D image data and 3D point cloud data are input into the YOLOv8-seg model, which outputs detection boxes, instance segmentation masks, and categories. The profile contour region is extracted using the masks. The following operations are performed using the 3D point cloud data: the profile centerline, profile normal vector, and profile tilt angle are calculated in the camera coordinate system; these are then converted to the mechanical coordinate system of the scaffold / robotic arm gripper; and the gripping posture of the scaffold / robotic arm gripper is controlled using hand-eye calibration.
[0084] In summary, the embodiments of the present invention have at least the following characteristics:
[0085] 1. Based on the selected camera field of view, it can realize automatic loading and unloading of profiles of different lengths from multiple points;
[0086] 2. Integrating 2D and 3D multimodal visual recognition methods to achieve high-precision positioning of profiles under stacked / tilted / complex backgrounds;
[0087] 3. Supports multi-point image fusion recognition, suitable for spatial attitude estimation and multi-point correction grasping of ultra-long profiles;
[0088] 4. A method for estimating profile type, clamping height and tilt angle by combining 3D point cloud feature analysis is proposed, which can realize the identification of multiple types of steel / I-beams, angle steel and other types of spherical flat steel, and can adapt different profile grippers according to the identified information;
[0089] 5. Significantly improves the flexibility and robustness of automated loading and unloading, making it particularly suitable for complex profile processing scenarios such as shipbuilding and new energy.
[0090] Based on the foregoing embodiments, this invention provides a device for automatic loading and unloading of profiles based on visual recognition. See [link to related documentation]. Figure 9 The diagram shows a structural schematic of a vision recognition-based automatic profile loading and unloading device, which mainly includes the following parts:
[0091] The location determination module 902 is used to determine the visual image capture location based on the attribute information of the profile to be sorted;
[0092] The visual recognition module 904 is used to collect the visual data of the profile to be sorted according to at least one target visual photography point, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result of the profile to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information of the profile to be sorted.
[0093] The loading / unloading module 906 is used to determine the target pose data corresponding to the loading / unloading actuator based on the visual recognition results, so as to control the loading / unloading actuator to perform loading / unloading actions on the profiles to be sorted according to the target pose data.
[0094] The automatic profile loading and unloading device based on vision recognition provided in this invention utilizes 3D+2D multimodal vision fusion technology to simultaneously acquire the three-dimensional structural information and two-dimensional texture features of the workpiece. This provides more robust and accurate technical support for workpiece recognition, posture estimation, and automatic grasping in complex industrial environments. Compared with traditional recognition methods that rely solely on 2D images, 3D vision can not only accurately reconstruct the spatial position and geometric contour of the workpiece but also possesses stronger adaptability. It can effectively cope with complex scenarios such as densely stacked profiles, partial occlusion, and posture changes, significantly improving recognition robustness and grasping success rate, and providing a solid guarantee for flexible production and automated operations in intelligent manufacturing.
[0095] In one embodiment, the visual recognition module 904 is specifically used for:
[0096] Collect visual data of the profiles to be sorted from a target visual photography point.
[0097] The integrity of the profiles to be sorted displayed in the profile visual data is checked. If the integrity check fails, the profile visual data corresponding to the profiles to be sorted is collected according to multiple target visual photography points.
[0098] For the visual data of profiles under at least one target visual capture point, perform profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted.
[0099] In one embodiment, the visual recognition module 904 is specifically used for:
[0100] The following processing is performed on the profile visual data under each target visual capture point: based on 3D point cloud data, or based on 3D point cloud data and 2D image data, the visual recognition result corresponding to the profile to be sorted is identified.
[0101] In one embodiment, the visual recognition module 904 is specifically used for:
[0102] Based on 3D point cloud data, the profile to be sorted and its sorting area are segmented to obtain the first segmentation result corresponding to the profile to be sorted.
[0103] Determine whether the first segmentation result corresponding to the profile to be sorted meets the preset conditions;
[0104] If so, then based on the first segmentation result, identify the visual recognition result corresponding to the profile to be sorted;
[0105] If not, the 2D image data is used to segment the profile to be sorted and the sorting area in which it is located to obtain the second segmentation result corresponding to the profile to be sorted. The second segmentation result and the 3D point cloud data are combined to identify the visual recognition result corresponding to the profile to be sorted.
[0106] In one embodiment, the visual recognition module 904 is specifically used for:
[0107] Based on the edge of the profile to be sorted, characterized by the first segmentation result or the second segmentation result, determine the unilateral information of the profile to be sorted.
[0108] Based on the unilateral information of the profile to be sorted and the depth information described by 3D point cloud data, the surface unevenness information of the profile to be sorted is determined.
[0109] Based on the unilateral information and surface unevenness information of the profile to be sorted, the visual recognition result corresponding to the profile to be sorted is determined.
[0110] In one embodiment, the loading / unloading module 906 is specifically used for:
[0111] The tilt angle and size information contained in the visual recognition results are transformed from the camera coordinate system to the actuator coordinate system;
[0112] Based on the hand-eye calibration results, the tilt angle and size information in the actuator coordinate system, determine the positional deviation of the loading and unloading actuator;
[0113] Based on the pose deviation, the current pose data of the loading / unloading actuator is adjusted to obtain the target pose data of the loading / unloading actuator.
[0114] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0115] This invention provides an electronic device, specifically, the electronic device includes a processor and a memory; the memory stores a computer program, which, when run by the processor, executes the method described in any of the above embodiments.
[0116] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 10, a memory 11, a bus 12 and a communication interface 13. The processor 10, the communication interface 13 and the memory 11 are connected through the bus 12. The processor 10 is used to execute executable modules, such as computer programs, stored in the memory 11.
[0117] The memory 11 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 13 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0118] Bus 12 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0119] The memory 11 is used to store programs. After receiving an execution instruction, the processor 10 executes the programs. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 10 or implemented by the processor 10.
[0120] Processor 10 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 10 or by instructions in software form. Processor 10 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 11. The processor 10 reads the information in memory 11 and, in conjunction with its hardware, completes the steps of the above method.
[0121] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0122] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0123] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automatic loading and unloading of profiles based on visual recognition, characterized in that, include: The visual image capture points are determined based on the attribute information of the profiles to be sorted. According to at least one target visual photography point, collect the profile visual data corresponding to the profile to be sorted, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result corresponding to the profile to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information corresponding to the profile to be sorted. Based on the visual recognition results, the target pose data corresponding to the loading and unloading actuator is determined, so as to control the loading and unloading actuator to perform loading and unloading actions on the profile to be sorted according to the target pose data; According to at least one target visual photography point, the visual data of the profile to be sorted is collected, and profile recognition processing based on multi-dimensional visual fusion is performed on it to obtain the visual recognition result of the profile to be sorted. The process includes: collecting the visual data of the profile to be sorted according to one target visual photography point; performing an integrity check on the profile to be sorted displayed in the profile visual data, and if the integrity check fails, collecting the visual data of the profile to be sorted according to multiple target visual photography points respectively; performing profile recognition processing based on multi-dimensional visual fusion on the profile visual data under at least one target visual photography point to obtain the visual recognition result of the profile to be sorted. For the profile visual data under at least one of the target visual photography points, perform profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted, including: performing the following processing on the profile visual data under each of the target visual photography points: based on the 3D point cloud data, or based on the 3D point cloud data and the 2D image data, identify the visual recognition result corresponding to the profile to be sorted; Based on the 3D point cloud data, or based on the 3D point cloud data and the 2D image data, the visual recognition result corresponding to the profile to be sorted is identified, including: segmenting the profile to be sorted and the sorting area it is located in based on the 3D point cloud data to obtain a first segmentation result corresponding to the profile to be sorted; determining whether the first segmentation result corresponding to the profile to be sorted meets a preset condition; if yes, then based on the first segmentation result, the visual recognition result corresponding to the profile to be sorted is identified; if no, then based on the 2D image data, the profile to be sorted and the sorting area it is located are segmented to obtain a second segmentation result corresponding to the profile to be sorted, and combining the second segmentation result and the 3D point cloud data, the visual recognition result corresponding to the profile to be sorted is identified.
2. The method for automatic loading and unloading of profiles based on visual recognition according to claim 1, characterized in that, Identifying the visual recognition result corresponding to the profile to be sorted includes: Based on the edge of the profile to be sorted as characterized by the first segmentation result or the second segmentation result, determine the one-sided information of the profile to be sorted. Based on the one-sided information of the profile to be sorted and the depth information described by the 3D point cloud data, the surface unevenness information of the profile to be sorted is determined; Based on the single-sided information and surface unevenness information of the profile to be sorted, the visual recognition result corresponding to the profile to be sorted is determined.
3. The method for automatic loading and unloading of profiles based on visual recognition according to claim 1, characterized in that, Based on the visual recognition results, the target pose data corresponding to the loading and unloading actuator is determined, including: The tilt angle and size information contained in the visual recognition result are transformed from the camera coordinate system to the actuator coordinate system; Based on the hand-eye calibration results, the tilt angle and size information in the coordinate system of the actuator, determine the positional deviation of the loading and unloading actuator; Based on the pose deviation, the current pose data corresponding to the loading / unloading actuator is adjusted to obtain the target pose data corresponding to the loading / unloading actuator.
4. A device for automatic loading and unloading of profiles based on vision recognition, characterized in that, include: The location determination module is used to determine the visual image capture location based on the attribute information of the profile to be sorted; The visual recognition module is used to collect the profile visual data corresponding to the profile to be sorted according to at least one target visual photography point, perform profile recognition processing based on multi-dimensional visual fusion, and obtain the visual recognition result corresponding to the profile to be sorted. The profile visual data includes 2D image data and 3D point cloud data, and the visual recognition result includes the tilt angle and size information corresponding to the profile to be sorted. The loading and unloading module is used to determine the target pose data corresponding to the loading and unloading actuator based on the visual recognition result, so as to control the loading and unloading actuator to perform loading and unloading actions on the profile to be sorted according to the target pose data; The visual recognition module is specifically used for: collecting profile visual data corresponding to the profile to be sorted according to one of the target visual photography points; performing an integrity check on the profile to be sorted displayed in the profile visual data; and if the integrity check fails, collecting profile visual data corresponding to the profile to be sorted according to multiple target visual photography points respectively. For the profile visual data under at least one of the target visual capture points, perform profile recognition processing based on multi-dimensional visual fusion to obtain the visual recognition result corresponding to the profile to be sorted. The visual recognition module is specifically used to: perform the following processing on the profile visual data under each target visual photography point: based on the 3D point cloud data, or based on the 3D point cloud data and the 2D image data, identify the visual recognition result corresponding to the profile to be sorted; The visual recognition module is specifically used to: segment the profile to be sorted and its sorting area based on the 3D point cloud data to obtain a first segmentation result corresponding to the profile to be sorted; and determine whether the first segmentation result corresponding to the profile to be sorted meets a preset condition. If so, then based on the first segmentation result, identify the visual recognition result corresponding to the profile to be sorted; If not, the profile to be sorted and the sorting area in which it is located are segmented based on the 2D image data to obtain a second segmentation result corresponding to the profile to be sorted. The second segmentation result and the 3D point cloud data are then combined to identify the visual recognition result corresponding to the profile to be sorted.
5. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Robot disordered grabbing method and system based on machine vision and storage medium
CN112070818A
Three-dimensional machine vision practical training method and system, computer equipment and storage medium
CN116721582A