Visual perception method and device, equipment and medium

Through the cooperation of multi-spectral cameras and lidar, the problem of industrial robots identifying the position of packaging box stacks in narrow environments is solved, and high-precision packaging box grabbing is achieved, meeting the needs of automated operations in the container.

CN120339405APending Publication Date: 2025-07-18XYZ ROBOTICS CHINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410070596.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when industrial robots perform rapid palletization or unpalletization in narrow environments such as containers, the camera is large in size and insufficient in accuracy, and the imaging field of view cannot meet the positional requirements for identifying the packing box pallets.

Method used

Using a multi-spectral camera and lidar, the 2D image and point cloud image of the packaging box stack are collected by calibrating the associated field of view, the deep learning model is used to detect the packaging box area, and its position is projected into the point cloud image, and sent to the robot end to perform the grabbing action.

Benefits of technology

The imaging field of view has been expanded, the imaging clarity has been improved, and the application needs of automatic loading and unloading in containers has been met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339405A_ABST
    Figure CN120339405A_ABST
Patent Text Reader

Abstract

The invention provides a visual perception method and device, equipment and a medium, and the method comprises the steps: obtaining a 2D image and a point cloud image of a packaging box stack body, the 2D image is collected through a multispectral camera, the point cloud image is collected through a laser radar, and the multispectral camera and the laser radar carry out the visual field correlation through calibration; conveying the 2D image of the packaging box stack body to a deep learning model, and detecting an area of each packaging box on the 2D image through the deep learning model; projecting the detected 2D image of the position of each packaging box into the point cloud image, and determining the pose of the packaging box according to the point cloud corresponding to each packaging box; and the pose of each packaging box is sent to a robot end, so that the robot executes the grabbing action on the packaging boxes. Through cooperation of the fisheye camera, the multispectral camera and the laser radar, the imaging FOV is expanded, the imaging definition is improved, and application of automatic feeding and discharging in the container is met.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] An industrial robot is an intelligent device equipped with sensors, objective lenses, and electro-optical systems, capable of quickly sorting and handling goods.

[0003] More and more visual sensors and force sensors are being used on industrial robots, making industrial robots increasingly intelligent. With the progress of technologies such as sensing and recognition systems and artificial intelligence, robots are evolving from being unidirectionally controlled to storing and applying data by themselves, gradually becoming more information-based.

[0004] To expand the application scenarios and scope of industrial robots, in the prior art, mobile robots are manufactured by installing industrial robots on mobile bases, thereby enabling the movement of industrial robots to achieve functions such as mobile depalletizing and mobile picking. However, when performing rapid-paced palletizing or depalletizing in narrow environments such as inside containers, it is necessary to identify the packaging box stack, determine the pose of each box. The cameras in the prior art are large in size, insufficient in accuracy, and the imaging field of view cannot meet the application scenarios in such cases. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a visual perception method, device, equipment, and medium.

[0006] According to the visual perception method provided by the present invention, it includes:

[0007] Step S1: Obtain a 2D image and a point cloud image of the packaging box stack. The 2D image is collected by a multispectral camera, and the point cloud image is collected by a lidar. The multispectral camera and the lidar are associated with each other's fields of view through calibration;

[0008] Step S2: Transmit the 2D image of the packaging box stack to a deep learning model, and detect the area of each packaging box on the 2D image through the deep learning model;

[0009] Step S3: Project the 2D image where the position of each detected packaging box is located into the point cloud image, and determine the pose of the packaging box according to the point cloud corresponding to each packaging box;

[0010] Step S4: Send the pose of each packaging box to the robot end, so that the robot performs a grasping action on the packaging box.

[0011] Preferably, the step S1 includes the following steps:

[0012] Step S101: Calibrate the multispectral camera and the lidar to associate the fields of view of the multispectral camera and the lidar;

[0013] Step S102: Control the multispectral camera to collect the 2D image of the packing case stack;

[0014] Step S103: Control the lidar to collect the point cloud image of the packing case stack to obtain the 2D image and the point cloud image of the packing case stack.

[0015] Preferably, the step S101 includes the following steps:

[0016] Step S1011: Control the multispectral camera to collect the calibrated 2D image of the target calibration board, and determine the first calibration board area in the calibrated 2D image;

[0017] Step S1012: Control the lidar to collect the calibrated point cloud image of the target calibration board, and determine the second calibration board area in the calibrated point cloud image;

[0018] Step S1013: Register the first calibration board area and the second calibration board area to determine the mapping relationship between the first calibration board area and the second calibration board area, that is, realize the association of the fields of view of the multispectral camera and the lidar.

[0019] Preferably, the step S3 includes the following steps:

[0020] Step S301: Obtain the detection frame of each packing case detected on the 2D image;

[0021] Step S302: Project the detection frame of each packing case into the point cloud image to determine the point cloud area corresponding to the packing case;

[0022] Step S303: Determine the pose of each packing case according to the point cloud area corresponding to the detection frame of each packing case.

[0023] Preferably, the step S3 includes the following steps:

[0024] Step S301: Obtain the detection frame of each packing case detected on the 2D image, and further determine the box mask within the detection frame;

[0025] Step S302: Project the box mask of each packing case into the point cloud image to determine the point cloud area corresponding to the packing case;

[0026] Step S303: Determine the pose of each packing case according to the point cloud area corresponding to the box mask of each packing case.

[0027] Preferably, the step S3 includes the following steps:

[0028] Step S301: Obtain the detection frame of each packing box detected on the 2D image, and further determine the box mask within the detection frame;

[0029] Step S302: Project the box mask of each packing box into the point cloud image to determine the point cloud region corresponding to the box mask;

[0030] Step S302: Determine the minimum outer contour of the point cloud formed by the point cloud region corresponding to the box mask, and determine the pose of the packing box according to the point cloud corresponding to the minimum outer contour.

[0031] Preferably, the multispectral camera is provided with a wide-angle lens, the FOV of the multispectral camera is greater than 120°, and a stereoscopic projection model is adopted.

[0032] According to the visual perception system provided by the present invention, it includes the following modules:

[0033] An image acquisition module, configured to acquire a 2D image and a point cloud image of a stack of packing boxes, the 2D image is acquired by a multispectral camera, the point cloud image is acquired by a lidar, and the multispectral camera and the lidar are associated with each other's fields of view through calibration;

[0034] A detection module, configured to send the 2D image of the stack of packing boxes to a deep learning model, and detect the region of each packing box on the 2D image through the deep learning model;

[0035] A pose calculation module, configured to project the 2D image in which the position of each detected packing box is located into the point cloud image, and determine the pose of the packing box according to the point cloud corresponding to each packing box;

[0036] A grasping control module, configured to send the pose of each packing box to the robot side, so that the robot performs a grasping action on the packing box.

[0037] According to the visual perception device provided by the present invention, it includes:

[0038] A processor;

[0039] A memory module, which stores executable instructions of the processor;

[0040] Wherein, the processor is configured to execute the steps of the visual perception method by executing the executable instructions.

[0041] According to the computer-readable storage medium provided by the present invention, it is used to store a program, and when the program is executed, the steps of the visual perception method are implemented.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] In the present invention, a 2D image of a stack of packaging boxes is collected by a multispectral camera, and a point cloud image is collected by a lidar. The 2D image is transmitted to a deep learning model, and the area of each packaging box is detected on the 2D image through the deep learning model. The 2D image at the position of each detected packaging box is projected into the point cloud image, and the pose of the packaging box is determined according to the point cloud corresponding to each packaging box. The pose of each packaging box is sent to the robot end so that the robot can perform a grasping action on the packaging box. Through the cooperation of the multispectral camera and the lidar, the FOV of imaging is expanded and the clarity of imaging is improved, meeting the application of automatic loading and unloading inside the container. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts. By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:

[0045] Figure 1 It is a flowchart of the steps of the visual perception method in the embodiment of the present invention;

[0046] Figure 2 It is a flowchart of the steps of obtaining the 2D image and the point cloud image of the stack of packaging boxes in the embodiment of the present invention;

[0047] Figure 3 It is a flowchart of the steps of realizing the association of the fields of view of the multispectral camera and the lidar in the embodiment of the present invention;

[0048] Figure 4 It is a flowchart of the steps of determining the pose of the packaging box in the first embodiment of the present invention;

[0049] Figure 5 It is a flowchart of the steps of determining the pose of the packaging box in the second embodiment of the present invention;

[0050] Figure 6 It is a flowchart of the steps of determining the pose of the packaging box in the third embodiment of the present invention;

[0051] Figure 7 It is a schematic diagram of the stack of packaging boxes after detection in the embodiment of the present invention;

[0052] Figure 8 It is a schematic diagram of the modules of the visual perception system in the embodiment of the present invention;

[0053] Figure 9 The structural schematic diagram of the visual perception device in the embodiment of the present invention; and

[0054] Figure 10 The structural schematic diagram of the computer-readable storage medium in the embodiment of the present invention. Detailed implementation manners

[0055] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.

[0056] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0057] The technical solution of the present invention will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0058] The technical solution of the present invention and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the drawings.

[0059] Figure 1 The step flowchart of the visual perception method in the embodiment of the present invention, as Figure 1 shown, the visual perception method provided by the present invention includes:

[0060] Step S1: Obtain a 2D image and a point cloud image of the packing case stack. The 2D image is collected by a multispectral camera, the point cloud image is collected by a lidar, and the multispectral camera and the lidar are associated with the field of view through calibration;

[0061] In an embodiment of the present invention, the lidar includes a first lidar and a second lidar; the multi-spectral camera is disposed between the first lidar and the second lidar.

[0062] In an embodiment of the present invention, the FOV of the multi-spectral camera is greater than 120° and a stereoscopic projection model is adopted. The multi-spectral camera is provided with a wide-angle lens; specifically, the wide-angle lens is a fish-eye lens; the multi-spectral camera adopts any one of an RGB camera, a grayscale camera, and an infrared camera, and preferably an RGB camera.

[0063] In an embodiment of the present invention, at this time, the number of lidars can be set to 2 or 4. The overall field of view of one multi-spectral camera is aligned with that of two lidars, or the overall field of view of two multi-spectral cameras is aligned with that of four lidars.

[0064] Figure 2 It is a flowchart of steps for obtaining a 2D image and a point cloud image of a packing case stack in an embodiment of the present invention. As Figure 2 shown, the step S1 includes the following steps:

[0065] Step S101: Calibrate the multi-spectral camera and the lidar to associate the fields of view of the multi-spectral camera and the lidar;

[0066] Step S102: Control the multi-spectral camera to collect a 2D image of the packing case stack;

[0067] Step S103: Control the lidar to collect a point cloud image of the packing case stack to obtain a 2D image and a point cloud image of the packing case stack.

[0068] In an embodiment of the present invention, the multi-spectral camera can be controlled to collect a 2D image of the packing case stack first, and then the lidar can be controlled to collect a point cloud image of the packing case stack. Or the lidar can be controlled to collect a point cloud image of the packing case stack first, and then the multi-spectral camera can be controlled to collect a 2D image of the packing case stack. It is also possible to simultaneously control the multi-spectral camera to collect a 2D image of the packing case stack and the lidar to collect a point cloud image of the packing case stack.

[0069] Figure 3 It is a flowchart of steps for realizing the association of the fields of view of the multi-spectral camera and the lidar in an embodiment of the present invention. As Figure 3 shown, the step S101 includes the following steps:

[0070] Step S1011: Control the multi-spectral camera to collect a calibrated 2D image of a target calibration board and determine a first calibration board area in the calibrated 2D image;

[0071] Step S1012: Control the lidar to collect the point cloud image of the target calibration board, and determine the second calibration board area in the point cloud image of the calibration board;

[0072] Step S1013: Register the first calibration board area and the second calibration board area to determine the mapping relationship between the first calibration board area and the second calibration board area, that is, realize the association of the fields of view of the multispectral camera and the lidar.

[0073] In the embodiment of the present invention, when registering the first calibration board area and the second calibration board area, specifically, a coordinate system is established, the first plane equation of the first calibration board area is generated, the second plane equation of the second calibration board area is generated, and then the transformation matrix between the first plane equation and the second plane equation is generated, that is, the registration of the first calibration board area and the second calibration board area is realized.

[0074] Step S2: Transmit the 2D image of the packing box stack to the deep learning model, and detect the area of each packing box on the 2D image through the deep learning model;

[0075] In the embodiment of the present invention, the deep learning model is trained and generated by using a convolutional neural network model.

[0076] Step S3: Project the 2D image detecting the position of each packing box into the point cloud image, and determine the pose of the packing box according to the point cloud corresponding to each packing box;

[0077] Step S4: Send the pose of each packing box to the robot end, so that the robot performs a grasping action on the packing box.

[0078] Figure 4 This is the flowchart of the steps for determining the pose of the packing box in the first embodiment of the present invention. As Figure 4 shown, the step S3 includes the following steps:

[0079] Step S301: Obtain the detection frame of each packing box detected on the 2D image;

[0080] Step S302: Project the detection frame of each packing box into the point cloud image to determine the point cloud area corresponding to the packing box;

[0081] Step S303: Determine the pose of the packing box according to the point cloud area corresponding to the detection frame of each packing box.

[0082] Figure 5 This is the flowchart of the steps for determining the pose of the packing box in the second embodiment of the present invention. As Figure 5As shown, step S3 includes the following steps:

[0083] Step S301: Obtain the detection frame of each packing box detected on the 2D image, and further determine the box mask within the detection frame;

[0084] Step S302: Project the box mask of each packing box into the point cloud image to determine the point cloud area corresponding to the packing box;

[0085] Step S303: Determine the pose of the packing box according to the point cloud area corresponding to the box mask of each packing box.

[0086] In the embodiment of the present invention, since the area of some detection frames is larger than the actual area of the packing box area, there are often errors in calculating the pose of the packing box directly through the point cloud within the detection frame, and dropping of parts will occur during the grasping of the packing box. Calculating the pose of the packing box through the point cloud corresponding to the box mask can significantly improve the calculation accuracy of the pose.

[0087] In the embodiment of the present invention, the segmentation algorithm is used to calculate the box mask.

[0088] Figure 6 It is the flow chart of the steps to determine the pose of the packing box in the third embodiment of the present invention. As Figure 6 shown, step S3 includes the following steps:

[0089] Step S301: Obtain the detection frame of each packing box detected on the 2D image, and further determine the box mask within the detection frame;

[0090] Step S302: Project the box mask of each packing box into the point cloud image to determine the point cloud area corresponding to the box mask;

[0091] Step S302: Determine the minimum outer contour of the point cloud formed by the point cloud area corresponding to the box mask, and determine the pose of the packing box according to the point cloud within the minimum outer contour.

[0092] In the embodiment of the present invention, since some box masks will include adjacent areas, there are often errors in calculating the pose of the packing box directly through the point cloud of the box mask within the detection frame, and dropping of parts will occur during the grasping of the packing box. Calculating the pose of the packing box through the minimum outer contour of the point cloud formed by the point cloud area corresponding to the box mask can significantly improve the calculation accuracy of the pose.

[0093] In the embodiment of the present invention, the minimum outer contour of the point cloud is a rectangular frame corresponding to the outer side of the packing box. When generating the minimum outer contour of the point cloud, it is generated by clustering the point cloud and at the same time excluding the points far from the majority of the point cloud planes.

[0094] Figure 8 This is a schematic diagram of the modules of the visual perception system in the embodiments of the present invention. As Figure 8 shown, the visual perception system provided by the present invention includes the following modules:

[0095] An image acquisition module, configured to acquire a 2D image and a point cloud image of a packaging box stack. The 2D image is collected by a multispectral camera, and the point cloud image is collected by a lidar. The multispectral camera and the lidar are associated with each other's fields of view through calibration;

[0096] A detection module, configured to send the 2D image of the packaging box stack to a deep learning model, and detect the area of each packaging box on the 2D image through the deep learning model;

[0097] A pose calculation module, configured to project the 2D image in which the position of each detected packaging box is located into the point cloud image, and determine the pose of the packaging box according to the point cloud corresponding to each packaging box;

[0098] A grasping control module, configured to send the pose of each packaging box to the robot side, so that the robot performs a grasping action on the packaging box.

[0099] In the embodiments of the present invention, a visual perception device is further provided, including a processor and a memory. The memory stores executable instructions of the processor. Wherein, the processor is configured to execute the steps of the visual perception method by executing the executable instructions.

[0100] As described above, in this embodiment, a 2D image of the packaging box stack is collected by a multispectral camera, and the point cloud image is collected by a lidar. The 2D image is sent to a deep learning model, and the area of each packaging box is detected on the 2D image through the deep learning model. The 2D image in which the position of each detected packaging box is located is projected into the point cloud image, and the pose of the packaging box is determined according to the point cloud corresponding to each packaging box. The pose of each packaging box is sent to the robot side, so that the robot performs a grasping action on the packaging box. Through the cooperation of the multispectral camera and the lidar, the FOV of imaging is expanded and the clarity of imaging is improved, meeting the application of automatic loading and unloading inside the container.

[0101] Those skilled in the art of the present technical field can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0102] Figure 9 This is a schematic structural diagram of a visual perception device in an embodiment of the present invention. The following will refer to Figure 9 to describe the electronic device 600 according to this embodiment of the present invention. Figure 9 The electronic device 600 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0103] As Figure 9 shown, the electronic device 600 is presented in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0104] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the above-mentioned visual perception method part of this specification. For example, the processing unit 610 can execute the steps as Figure 1 shown.

[0105] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0106] The storage unit 620 may further include a program / utility 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0107] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0108] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, a camera, a depth camera, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Also, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 9 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0109] In an embodiment of the present invention, there is also provided a computer-readable storage medium for storing a program, and the steps of the visual perception method are implemented when the program is executed. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above-mentioned visual perception method part of this specification.

[0110] As shown above, when the program of the computer-readable storage medium in this embodiment is executed, a 2D image of the packing case stack is collected by a multispectral camera, the point cloud image is collected by using a lidar, the 2D image is sent to a deep learning model, the area of each packing case is detected on the 2D image by the deep learning model, the 2D image detecting the position of each packing case is projected into the point cloud image, the pose of the packing case is determined according to the point cloud corresponding to each packing case, and the pose of each packing case is sent to the robot end so that the robot performs a grasping action on the packing case. Through the cooperation of the multispectral camera and the lidar, the FOV of imaging is expanded and the clarity of imaging is improved, meeting the application of automatic loading and unloading inside the container.

[0111] Figure 10 is a schematic structural diagram of the computer-readable storage medium in an embodiment of the present invention. Refer to Figure 10As shown, a program product 800 for implementing the above method according to an embodiment of the present invention is described. It may employ a portable compact disc read-only memory (CD-ROM), include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0112] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0113] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0114] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0115] In the embodiments of the present invention, a 2D image of a stack of packaging boxes is collected by a multispectral camera, and a point cloud image is collected by using a lidar. The 2D image is transmitted to a deep learning model. The area of each packaging box is detected on the 2D image through the deep learning model. The 2D image with the detected position of each packaging box is projected into the point cloud image. The pose of the packaging box is determined according to the point cloud corresponding to each packaging box. The pose of each packaging box is sent to the robot end so that the robot performs a grasping action on the packaging box. Through the cooperation of the multispectral camera and the lidar, the FOV of imaging is expanded and the clarity of imaging is improved, meeting the application of automatic loading and unloading inside the container.

[0116] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0117] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A visual perception method, characterized in that, Including: Step S1: Obtain the 2D image and the point cloud image of the packing case stack. The 2D image is collected by a multispectral camera, the point cloud image is collected by a lidar, and the multispectral camera and the lidar are associated with each other in terms of the field of view through calibration. Step S2: Transmit the 2D image of the packing case stack to a deep learning model, and detect the area of each packing case on the 2D image through the deep learning model. Step S3: Project the 2D image where the position of each detected packing case is located onto the point cloud image, and determine the pose of the packing case according to the point cloud corresponding to each packing case. Step S4: Send the pose of each packing case to the robot side, so that the robot performs a grasping action on the packing case.

2. The visual perception method according to claim 1, wherein The step S1 includes the following steps: Step S101: Calibrate the multispectral camera and the lidar to associate the fields of view of the multispectral camera and the lidar. Step S102: Control the multispectral camera to collect the 2D image of the packing case stack. Step S103: Control the lidar to collect the point cloud image of the packing case stack to obtain the 2D image and the point cloud image of the packing case stack.

3. The visual perception method according to claim 1, wherein The step S101 includes the following steps: Step S1011: Control the multispectral camera to collect the calibrated 2D image of the target calibration board, and determine the first calibration board area in the calibrated 2D image. Step S1012: Control the lidar to collect the calibrated point cloud image of the target calibration board, and determine the second calibration board area in the calibrated point cloud image. Step S1013: Register the first calibration board area and the second calibration board area to determine the mapping relationship between the first calibration board area and the second calibration board area, that is, realize the association of the fields of view of the multispectral camera and the lidar.

4. The visual perception method according to claim 1, characterized in that, The step S3 includes the following steps: Step S301: Obtain the detection frame of each packing case detected on the 2D image. Step S302: Project the detection frame of each packing case onto the point cloud image to determine the point cloud area corresponding to the packing case. Step S303: Determine the pose of the packing case according to the point cloud area corresponding to the detection frame of each packing case.

5. The visual perception method according to claim 1, wherein The step S3 includes the following steps: Step S301: Obtain the detection frame of each packing case detected on the 2D image, and further determine the box mask within the detection frame. Step S302: Project the box mask of each packing case onto the point cloud image to determine the point cloud area corresponding to the packing case. Step S303: Determine the pose of the packing case according to the point cloud area corresponding to the box mask of each packing case.

6. The visual perception method according to claim 1, characterized in that, The step S3 includes the following steps: Step S301: Obtain the detection frame of each packing case detected on the 2D image, and further determine the box mask within the detection frame. Step S302: Project the box mask of each packing case onto the point cloud image to determine the point cloud area corresponding to the box mask. Step S302: Determine the minimum outer contour of the point cloud formed by the point cloud region corresponding to the box mask, and determine the pose of the packaging box according to the point cloud corresponding to the minimum outer contour.

7. The visual perception method according to claim 1, characterized in that The multi-spectral camera is provided with a wide-angle lens, and the FOV of the multi-spectral camera is greater than 120° and a stereoscopic projection model is adopted.

8. A visual perception system, characterized in that, It includes the following modules: An image acquisition module, configured to acquire a 2D image and a point cloud image of a stack of packaging boxes. The 2D image is acquired by a multi-spectral camera, the point cloud image is acquired by a lidar, and the multi-spectral camera and the lidar are associated with each other in terms of field of view through calibration; A detection module, configured to send the 2D image of the stack of packaging boxes to a deep learning model, and detect the region of each packaging box on the 2D image through the deep learning model; A pose calculation module, configured to project the 2D image detecting the position of each packaging box into the point cloud image, and determine the pose of the packaging box according to the point cloud corresponding to each packaging box; A grasping control module, configured to send the pose of each packaging box to the robot side, so that the robot performs a grasping action on the packaging box.

9. A visual perception device, characterized in that, It includes: A processor; A memory module, in which executable instructions of the processor are stored; Wherein, the processor is configured to execute the steps of the visual perception method according to any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that, When the program is executed, the steps of the visual perception method according to any one of claims 1 to 7 are implemented.