Camera pose adjusting method and device and computing device cluster
By using single-view camera shooting and image processing model accuracy evaluation to adjust pose, the problem of low efficiency in existing technologies is solved, and efficient and accurate object acquisition is achieved.
Patent Information
- Application Number
- CN202411154992.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2024-08-21
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, methods that use multiple poses to photograph an object to determine the grasping or suction point are inefficient and take a long time to capture.
A single-view camera is used for shooting, and the camera pose is adjusted by evaluating the accuracy of the image processing model. The output of the image processing model is optimized to guide the object acquisition device to successfully acquire the object.
By using single-view shooting and pose adjustment, shooting efficiency is improved, the accuracy of the output results of the image processing model is ensured, thereby increasing the success rate of the object acquisition device.
Smart Images

Figure CN121010640A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision, and in particular to a camera pose adjustment method, apparatus, and computing device cluster. Background Technology
[0002] Object acquisition is one of the applications that combines robotics and machine vision technologies. When using a robot (such as a multi-degree-of-freedom robotic arm connected to a gripper or suction cup) to grasp or pick up an object, it is necessary to determine the grasping point or picking point of the object.
[0003] Currently, the object's grasping or suction point is determined by taking multiple photos of the object from different poses using the same camera. However, this method involves long shooting times and low efficiency. Summary of the Invention
[0004] This invention provides a camera pose adjustment method, apparatus, and computing device cluster. The camera performs single-view shooting of an object to save shooting time. At the same time, based on the accuracy evaluation of the camera pose by the image processing model, the camera pose is adjusted so that when the image processing model processes the single-view image taken by the camera, the output of the image processing model can guide the object acquisition device to successfully acquire the object.
[0005] In a first aspect, the present invention provides a camera pose adjustment method, the method comprising: determining a first pose of the camera; the camera taking a picture of an object in the transport area of an object transport device in the first pose; inputting the image taken by the camera in the first pose into an image processing model, and determining the output result of the image processing model; the output result being used to indicate the three-dimensional reconstruction result of the object and / or the acquisition position point of the object; and the output result being used to guide the object acquisition device to acquire the object in the transport area.
[0006] Based on the output of the image processing model, evaluation information is determined. The evaluation information is used to indicate the accuracy evaluation of the image processing model in the first pose. Based on the evaluation information, the first pose is adjusted to obtain the second pose of the camera.
[0007] In the solution provided by this invention, the camera takes a single-view shot of the object to save shooting time; at the same time, the camera pose is adjusted according to the accuracy evaluation of the camera pose by the image processing model, so that when the image processing model processes the image taken by the camera from a single view, the output of the image processing model can guide the object acquisition device to successfully acquire the object.
[0008] In some possible implementations, the first pose is adjusted based on the evaluation information to obtain the second pose of the camera, including: based on the evaluation information, if it is determined that the first pose needs to be adjusted, the first pose is adjusted to obtain the second pose of the camera.
[0009] In this scheme, the pose is adjusted when the evaluation information indicates that the accuracy of the image processing model is low, thereby improving the accuracy of the image processing model in the adjusted pose.
[0010] In some possible implementations, the first pose is adjusted to obtain the camera's second pose, including:
[0011] Starting from the first pose, perform pose search according to the first scale step and / or the first angle step to determine multiple first candidate poses of the camera; determine the second pose of the camera from the multiple first candidate poses.
[0012] In one embodiment of this implementation, the first scale step size and / or the first angle step size are determined by evaluation information.
[0013] In this scheme, the step size of pose adjustment is guided by the accuracy evaluation of the image processing model, so as to find poses with better participation value within a suitable range and improve the performance of the image processing model in the adjusted pose.
[0014] In one embodiment of this implementation, determining the second pose of the camera from a plurality of first candidate poses includes:
[0015] For each of the multiple first candidate poses, determine the score of the image captured by the camera when it is in each first candidate pose; and take the first candidate pose with the highest score as the second pose of the camera.
[0016] In this scheme, the camera poses are selected by evaluating the images captured by the camera and obtaining the camera poses with higher participation value.
[0017] In one embodiment of this implementation, determining the score of an image captured when the camera is in a first candidate pose includes:
[0018] Determine the text, which describes the shooting features and / or object features of the object captured by the camera; determine the similarity between the image captured by the camera in the first candidate pose and the text; and determine the score of the image captured by the camera in the first candidate pose based on the similarity.
[0019] In this approach, images captured by the camera are evaluated using text, and camera poses are filtered to obtain camera poses that are more consistent with the text descriptions and have higher reference value.
[0020] In some possible implementations, the first pose of the camera is determined, including:
[0021] Starting from the initial pose of the camera, pose search is performed according to the second scale step and / or the second angle step to obtain multiple second candidate poses of the camera; the second angle step is greater than the first angle step; the second scale step is greater than the first scale step; from the multiple second candidate poses, the first pose of the camera is determined.
[0022] In this scheme, a coarse-grained step size is selected for pose search during the initial pose selection. Subsequently, when further pose adjustment is needed, a fine-grained step size is selected for pose search to achieve more precise pose adjustment.
[0023] In some possible implementations, the first pose or initial pose is determined based on job information, which is used to indicate the transport area of the object transport device, the acquisition area of the object acquisition device, and the camera parameters of the camera; the acquisition area includes at least a portion of the transport area.
[0024] In some possible implementations, the evaluation information includes at least one of the following: acquisition success rate, accuracy of the acquisition location points predicted by the image processing model, and accuracy of the 3D reconstruction results predicted by the image processing model; the acquisition success rate is used to indicate the proportion of times the object acquisition device successfully acquires an object out of the total number of times it acquires an object based on the output results.
[0025] In some possible implementations, the output may also include a score for the location point.
[0026] Secondly, embodiments of the present invention provide a camera pose adjustment device, which includes several modules. Each module is used to execute various steps in the camera pose method provided in the first aspect of the present invention. The division of modules is not limited here. For the specific functions performed by each module of this camera pose adjustment device and the beneficial effects achieved, please refer to the functions of each step in the camera pose method provided in the first aspect of the present invention; they will not be repeated here. The second aspect or any implementation thereof is a device implementation corresponding to the first aspect or any implementation thereof, and the descriptions in the first aspect or any implementation thereof are applicable to the second aspect or any implementation thereof.
[0027] For example, the camera pose adjustment device includes:
[0028] The pose determination module is used to determine the first pose of the camera; the camera takes pictures of objects on the conveying area of the object conveying device in the first pose.
[0029] The prediction module is used to input the image captured by the camera in the first pose into the image processing model and determine the output result of the image processing model; the output result is used to indicate the 3D reconstruction result of the object and / or the acquisition position point of the object; the output result is used to guide the object acquisition device to acquire the object on the transport area.
[0030] The evaluation module is used to determine evaluation information based on the output of the image processing model. The evaluation information is used to indicate the accuracy evaluation of the image processing model in the first pose.
[0031] The pose adjustment module is used to adjust the first pose based on the evaluation information to obtain the second pose of the camera.
[0032] In some possible implementations, a pose adjustment module is used to adjust the first pose based on evaluation information, if it is determined that the first pose needs adjustment, to obtain the second pose of the camera.
[0033] In some possible implementations, the pose adjustment module is used to perform pose search starting from the first pose, according to the first scale step and / or the first angle step, to determine multiple first candidate poses of the camera; and to determine the second pose of the camera from the multiple first candidate poses.
[0034] In one embodiment of this implementation, the first scale step size and / or the first angle step size are determined by evaluation information.
[0035] In one embodiment of this implementation, the pose adjustment module is used to determine the score of the image captured by the camera when it is in each of the multiple first candidate poses; and to take the first candidate pose with the highest score as the second pose of the camera.
[0036] In one embodiment of this implementation, the pose adjustment module is used to determine text, which describes the shooting features and / or object features of the object being photographed by the camera; determine the similarity between the image and the text when the camera is in a first candidate pose; and determine a score for the image when the camera is in the first candidate pose based on the similarity.
[0037] In some possible implementations, the pose determination module is used to perform pose search starting from the initial pose of the camera, according to a second scale step and / or a second angle step, to obtain multiple second candidate poses of the camera; the second angle step is greater than the first angle step; the second scale step is greater than the first scale step; and the first pose of the camera is determined from the multiple second candidate poses.
[0038] In some possible implementations, the first pose or initial pose is determined based on job information, which indicates the transport area of the object transport device, the acquisition area of the object acquisition device, and the camera parameters; the acquisition area includes at least a portion of the transport area. In some possible implementations, the evaluation information includes at least one of the following: acquisition success rate, the accuracy of the acquisition position points predicted by the image processing model, and the accuracy of the 3D reconstruction results predicted by the image processing model; the acquisition success rate indicates the proportion of successful object acquisitions out of the total number of acquisitions performed by the object acquisition device based on the output results.
[0039] In some possible implementations, the output may also include a score for the location point.
[0040] Thirdly, embodiments of the present invention provide a camera pose adjustment device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is used to execute the method provided in the first aspect.
[0041] Fourthly, embodiments of the present invention provide a camera pose adjustment device, which executes computer program instructions to perform the method provided in the first aspect. Exemplarily, the device may be a chip or a processor.
[0042] In one example, the device may include a processor that can be coupled to memory, read instructions from the memory, and execute the methods provided in the first aspect according to those instructions. The memory may be integrated into the chip or processor, or it may be independent of the chip or processor.
[0043] Fifthly, the present invention provides a computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the method described in the first aspect and any implementation thereof.
[0044] In a sixth aspect, the present invention provides a computer program product containing instructions that, when executed by a cluster of computer devices, cause the cluster of computer devices to perform the method described in the first aspect and any implementation thereof.
[0045] In a seventh aspect, the present invention provides a computer-readable storage medium including computer program instructions, which, when executed by a cluster of computing devices, enable the cluster of computing devices to perform the method described in the first aspect and any implementation thereof. Attached Figure Description
[0046] Figure 1a This is a schematic diagram of the camera pose adjustment system provided in an embodiment of the present invention;
[0047] Figure 1b This is a schematic diagram of the camera pose adjustment system provided in an embodiment of the present invention. Figure 2 ;
[0048] Figure 2 This is a schematic flowchart of the camera pose adjustment method provided in an embodiment of the present invention;
[0049] Figure 3 This is a flowchart illustrating the camera pose adjustment method provided in an embodiment of the present invention. Figure 2 ;
[0050] Figure 4 This is a schematic diagram illustrating an application scenario of the camera pose adjustment method provided in an embodiment of the present invention;
[0051] Figure 5a This is a schematic diagram of the camera pose adjustment framework provided in an embodiment of the present invention;
[0052] Figure 5b yes Figure 5a A schematic diagram illustrating the implementation principle of the camera module in the diagram;
[0053] Figure 6 This is a schematic flowchart of the camera pose adjustment device provided in an embodiment of the present invention;
[0054] Figure 7 This is a schematic diagram of the structure of the computing device provided in an embodiment of the present invention;
[0055] Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present invention;
[0056] Figure 9 This is a schematic diagram of computing devices in a computer cluster connected via a network, as provided in an embodiment of the present invention. Detailed Implementation
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0059] References to "one embodiment" or "some embodiments" as used in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the invention. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including, but not limited to," unless otherwise specifically emphasized.
[0060] The following explanations cover some of the terms used in this embodiment. It should be noted that these explanations are for the convenience of those skilled in the art and are not intended to limit the scope of protection claimed by this invention.
[0061] Pose: The position and orientation of an object (rigid body) within its own coordinate system, graphically represented by a set of coordinate axes. Pose is a collective term for position and orientation, representing the precise and unique positional state of any rigid body in a spatial coordinate system. Specifically, pose comprises two main parts: position and orientation. Position: This part describes the specific location of the rigid body in three-dimensional space, usually represented by x, y, and z coordinates. Orientation: This part describes the orientation or rotational state of the rigid body in three-dimensional space. A single rotation of a rigid body in three-dimensional space can be characterized by the angles of rotation around the x, y, and z axes. These three angles are roll (φ), pitch (θ), and yaw (ψ), or α, β, and γ, respectively, representing the angles of rotation around the x, y, and z axes. Therefore, pose contains a total of 6 degrees of freedom: three positional degrees of freedom and three rotational degrees of freedom.
[0062] Contrastive Language-Image Pre-Training (CLIP) models are a type of multimodal pre-trained neural network that learns the alignment relationships between images and text through pre-training on a large amount of paired image and text data. The core idea of CLIP models is to use a text encoder and an image encoder to convert text and images into vector representations, and then generate predictions by calculating the cosine similarity between these vectors. This model is particularly suitable for zero-shot learning tasks, meaning the model can make predictions without seeing new training examples of images or text. CLIP models have performed well in multiple domains, such as image-text retrieval and image-text generation. The structure of a CLIP model includes an image encoder, a text encoder, and a multimodal fusion module. The image encoder typically uses models such as Convolutional Neural Networks (CNNs) or Variational Autoencoders (VAEs) to encode images into vectors. The text encoder typically uses models such as Recurrent Neural Networks (RNNs) or Transformers to encode text into vectors. The multimodal fusion module typically uses models such as attention mechanisms or fully connected layers to fuse image and text vectors.
[0063] Accuracy: How close the observed value is to the true value.
[0064] The following section introduces the camera pose adjustment system that may be applied to the camera pose adjustment method provided in the embodiments of the present invention. Figure 1a This diagram illustrates an example architecture of a camera pose adjustment system according to an embodiment of the present invention. The camera pose adjustment method provided by this embodiment can be applied to, for example... Figure 1a The system architecture diagram shown is as follows. Figure 1aAs shown, the camera pose adjustment system includes an object conveying device 101, an object acquisition device 102, a camera 103, a support 104, a control center 110, and a terminal 120. The control center 110 communicates with the camera device 103, the object acquisition device 102, and the terminal 120 via a network. The network can be a wired network or a wireless network. For example, a wired network can be a cable network, a fiber optic network, a Digital Data Network (DDN), etc., while a wireless network can be an internal network, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Public Service Telephone Network (PSTN), a Bluetooth network, a ZigBee network, a Global System for Mobile Communications (GSM), a CDMA (Code Division Multiple Access) network, a CPRS (General Packet Radio Service) network, etc., or any combination thereof. Understandably, a network can use any known network communication protocol to enable communication between different client layers and gateways. These network communication protocols can be various wired or wireless communication protocols, such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), New Radio (NR), Bluetooth, Wireless Fidelity (Wi-Fi), and other communication protocols.
[0065] The object conveying device 101 is used to convey objects. In some possible implementations, the object conveying device 101 includes a conveying device and a driving device. The area on the conveying device where objects can be placed is called the conveying area 1011. The conveying device is used to place objects, for example, it can be a conveyor belt. The driving device is used to drive the conveying device to move, so that the objects on the conveying device move, thereby changing the position and type of the objects on the conveying area 1011.
[0066] The object acquisition device 102 is used to acquire objects on the transfer area 1011 of the object transfer device 101, and can be a robot. In some possible implementations, the object acquisition device 102 may include a control device 1021 and an acquisition device 1022. The control device 1021 controls the operation of the acquisition device 1022, enabling the acquisition device 1022 to acquire the object. The acquisition device 1022 may include... Figure 1a The robotic arm 1023 shown and / or Figure 1b The suction cup 1024 shown can acquire objects by either gripping with a robotic arm 1023 or sucking with a suction cup 1024. The acquisition device 1022 has a specific acquisition area 1025, which includes at least a portion of the transfer area 1011.
[0067] The camera 103 is used for taking pictures. For example, the camera 103 can take pictures according to a shooting cycle, which is used to indicate the duration between two consecutive shooting moments.
[0068] The bracket 104 is used to mount the camera 103.
[0069] The terminal 120 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and other devices with a display screen 121. Exemplary embodiments of the terminal 120 involved in this solution include, but are not limited to, electronic devices running iOS, Android, Windows, Harmony OS, or other operating systems. This embodiment of the invention does not specifically limit the type of electronic device. In this embodiment, the terminal 120 is used to input camera parameters from the object conveying device 101, the object acquisition device 102, and the camera 103, and upload the working parameters to the control center 110.
[0070] It should be noted that, in this embodiment of the invention, the control center 110 can be the control device 1021, or a device other than the control device 1021. In some possible implementations, the control center 110 can be a cluster of several electronic devices. The electronic devices can be terminals, computers, or servers. The server can be a hardware server or embedded in a virtualization environment. For example, the server involved in this solution can be a virtual machine running on a hardware server that includes one or more other virtual machines. In one example, the electronic device can be used to provide cloud services. It can be a server or a super terminal that can establish communication connections with other devices and provide computing and / or storage functions for other devices.
[0071] In this embodiment of the invention, the control center 110 is used to determine the first pose of the camera 103; the image captured by the camera 103 in the first pose is input into the image processing model, and the control center 110 adjusts the first pose of the camera 103 based on the accuracy evaluation of the image processing model in the first pose to obtain the second pose of the camera 103; subsequently, the camera 103 can take pictures according to the second pose. In this scheme, the camera takes pictures of the object from a single perspective to save shooting time; at the same time, the camera pose is adjusted according to the accuracy evaluation of the image processing model in the camera pose, so that when the image processing model processes the image captured from the single perspective of the camera, the output result of the image processing model can guide the object acquisition device to successfully acquire the object. This is only an overview of the camera pose adjustment scheme; for details, please refer to the description below, which will not be repeated here.
[0072] Next, in conjunction with the camera pose adjustment system provided above, a camera pose adjustment method provided by an embodiment of the present invention will be described in detail.
[0073] Figure 2 This is a schematic flowchart of the camera pose adjustment method provided in an embodiment of the present invention. This embodiment can be applied to electronic devices, specifically servers or general computers.
[0074] like Figure 2 As shown, the camera pose adjustment method provided in this embodiment of the invention includes at least the following steps:
[0075] Step 201: Control center 110 determines the first pose of camera 103; camera 103 takes pictures of objects on transport area 1011 in the first pose.
[0076] In some possible embodiments, the first pose can be the current pose of the camera 103.
[0077] In some possible embodiments, the first pose can be determined based on job information, which indicates the transport area 1011 of the object transport device 101, the acquisition area 1025 of the object acquisition device 102, and the camera parameters of the camera 103; the acquisition area 1025 includes at least a portion of the transport area 1011. In a specific implementation, the display screen 121 of the terminal 120 can display an interface provided by the control center 110, where the user uploads job information, and the terminal 120 sends the job information to the control center 110.
[0078] The operational information may include operational parameters of object acquisition device 101, operational parameters of object acquisition device 102, and camera parameters of camera 103. For example, the operational parameters of object acquisition device 101 may be the location of the transport area 1011 used to transport objects, such as the length, width, and height of the transport device. For example, the operational parameters of object acquisition device 102 may be the acquisition area 1025, such as the activity drive of acquisition device 1022. For example, the camera parameters of camera 103 may be intrinsic parameters, extrinsic parameters, and field of view; wherein, intrinsic parameters are parameters related to the characteristics of camera 103 itself, including focal length, principal point (optical center) coordinates, distortion coefficients, etc. These parameters are usually determined during camera calibration because they are fixed for a specific camera model and do not change over time. Once the camera intrinsic parameters are determined, they generally remain unchanged during camera use. The intrinsic parameter matrix K is a 3×3 matrix, where fx and fy are the focal length parameters, and cx and cy are the coordinates of the origin of the image plane. The distortion parameter matrix D is a 1×5 matrix, including radial and tangential distortion coefficients. Extrinsic parameters describe the position and orientation of camera 103 in the world coordinate system, typically including rotation matrices and translation vectors. Extrinsic parameters may change at different camera positions or shooting times. The extrinsic parameter matrix is usually a 3×4 matrix, where the rotation matrix R is a 3×3 matrix representing the camera's rotational orientation, and the translation vector T is a 3×1 vector representing the camera's position offset in the world coordinate system. The field of view is the maximum angle between light rays entering camera 103. It should be noted that the height of camera 103, shooting direction, and field of view determine the size of the camera 103's viewing angle.
[0079] In one possible implementation of this embodiment, the first pose can be an initial pose selected based on the job information.
[0080] In one example of this implementation, the control center 110 can determine the initial pose of the camera 103 based on the job information. Starting from the initial pose of the camera 103, it performs a pose search according to the second scale step size and / or the second angle step size to obtain multiple second candidate poses of the camera 103. From the multiple second candidate poses, the first pose of the camera is determined. It should be noted that considering the job information for the initial pose is merely an example and does not constitute a specific limitation. In other possible scenarios, the display screen 121 of the terminal 120 can display the interface provided by the control center 110. The user uploads the initial pose of the camera 103 on the interface, and the terminal 120 sends the initial pose of the camera 103 to the control center 110. Correspondingly, the control center 110 can directly obtain the initial pose of the camera 103 and, starting from the initial pose of the camera 103, perform a pose search according to the second scale step size and / or the second angle step size to obtain multiple second candidate poses of the camera. From the multiple second candidate poses, the first pose of the camera is determined. Furthermore, the initial pose is merely an example; in some possible scenarios, the initial pose can also be the current pose of camera 103. In one possible implementation of this embodiment, the control center 110 determines the score of the image captured by camera 103 when it is in each of the plurality of second candidate poses; subsequently, the second candidate pose with the highest score is used as the first pose of camera 103.
[0081] The control center 110 determines the image captured when the camera 103 is in the second candidate pose by: constructing a digital work scene based on the work parameters of the object conveying device 101, the work parameters of the object acquisition device 102, and the camera parameters of the camera 103 in the acquired work information; deploying the camera 103 in the work scene according to the second candidate pose; then simulating the object transmission process involved in the actual scene on the object conveying device 101, and controlling the camera 103 to take pictures to obtain the image captured when the camera 103 is in the second candidate pose. This image is not an image obtained by taking pictures of the real world, but an image obtained by taking pictures of the digital world.
[0082] In some possible scenarios, the camera 103 can move flexibly on the support 104. The control center 110 can determine the image captured when the camera 130 is in the second candidate pose in another way: For each second candidate pose, control the camera 103 to move on the support 104 and / or adjust the shooting angle at a certain position so that the camera 103 is in the second candidate pose. Then, place the object into the transport area 1011 of the object transport device 101 to simulate the actual object transport process, and control the camera 103 to capture the image in the second candidate pose, thus obtaining the image captured when the camera 103 is in the second candidate pose. This method requires a significant amount of time and is typically used to determine the first pose from multiple second candidate poses. The multiple second candidate poses are obtained by searching for poses starting from the initial pose of the camera 103, according to a second scale step and / or a second angle step.
[0083] The method for determining multiple second candidate poses of camera 103 is as follows: starting from an initial pose, traversing three positional degrees of freedom according to the second scale step size to obtain multiple second candidate poses. For each second candidate pose, traversing three rotational degrees of freedom according to the second angle step size to continue obtaining multiple second candidate poses.
[0084] The scoring method for the image captured by camera 103 when it is in the second candidate pose, as determined by control center 110, can be as follows: terminal 120 sends text to control center 110, the text describing the shooting features and / or object features of the object captured by camera 103. Shooting features can be the viewing direction, and object features describe the characteristics of the object. The similarity between the image captured by camera 103 when it is in the first candidate pose and the text is determined. The similarity indicates whether the image satisfies the description in the text, such as whether the shooting features match, whether the image contains a significant amount of object features, or whether the image clearly represents the object features. A score for the image captured by camera 103 when it is in the first candidate pose is determined based on the similarity score; for example, the similarity score can be used as the score for the image captured by camera 103 when it is in the first candidate pose. For example, the similarity between the image captured by camera 103 when it is in the first candidate pose is determined. In a specific implementation, the display screen 121 of the terminal 120 can display the interface provided by the control center 110. The user can input text on the interface according to the characteristics of the object to be grasped. Subsequently, the control center 110 can input the text and image into the CLIP model to obtain the similarity between the text and image output by the CLIP model.
[0085] It should be noted that if the camera 103 is not yet in the first pose, the camera 103 can be installed manually according to the first pose, or the control center 110 can send the first pose of the camera 103 to the camera 103, and the camera 103 can adjust its own pose.
[0086] Step 202: The control center 110 inputs the image captured by the camera 103 in the first pose into the image processing model to obtain the output result of the image processing model. The output result is used to indicate the three-dimensional reconstruction result of the object and / or the acquisition position point of the object. The output result is used to guide the object acquisition device 102 to acquire the object on the transfer area 1011.
[0087] It should be noted that the objects indicated in the output are those identified by the image processing model.
[0088] The object acquisition location point describes the contact point when the object is acquired by the object acquisition device 102. For example, the acquisition location point can be the gripping point of the robotic arm 1023 or the suction point of the suction cup 1024. Correspondingly, the output result can include position description information of the acquisition location point; for example, the position description information can be the position of the contact point on the object on the identified object; for example, the position description information can be the coordinates of the acquisition location point on the object in 3D space (acquisition area 1025) and its position on the identified object.
[0089] The three-dimensional reconstruction result of an object is a mathematical model for the computer representation and processing of the object.
[0090] In other possible scenarios, the output may also include a score for the object's location point and / or the object type.
[0091] The image input to the image processing model can be a single image captured by the camera 103 in the first pose, or multiple images captured continuously by the camera 103 in the first pose.
[0092] In one possible embodiment, the control center 110 can deploy an image processing model, inputting the image captured by the camera 103 in the first pose into the image processing model to obtain the output result of the image processing model. The image processing model can identify the object indicated by the image, perform 3D reconstruction on the object indicated by the image, obtain the 3D reconstruction result, and predict the acquisition position point of the identified object based on the 3D reconstruction result. In some possible scenarios, the object acquisition device 102 adopts a robotic arm 1023, and the control center 110 has a motion database of the acquisition device 1022. Then, the method of predicting the acquisition position point of the identified object based on the 3D reconstruction result can be: based on the 3D reconstruction result, sorting and filtering the motion database, and selecting the acquisition position point of the object acquisition device 102 from the filtered list. In some possible scenarios, the object acquisition device 102 adopts a suction cup 1024. Then, the method of predicting the acquisition position point of the identified object based on the 3D reconstruction result can be: based on the 3D reconstruction result, determining the possible acquisition position points of the identified object and scoring them, and using the acquisition position point with the highest score as the acquisition position point of the identified object.
[0093] In some possible embodiments, the image processing model can be trained in a supervised manner, and the embodiments of the present invention are not intended to limit the structure of the image processing model. Specifically, the training method of the image processing model is as follows:
[0094] Multiple image samples are identified, each with a label indicating the acquisition location or 3D reconstruction result of each object. For each image sample, the image sample is input into the image processing model to be trained, and the output of the image processing model is determined. The image processing model is trained based on the error between the labels of multiple image samples and the corresponding output of the image processing model.
[0095] Step 203: Based on the output of the image processing model, the control center 110 determines the evaluation information, which is used to indicate the accuracy evaluation of the image processing model in the first pose.
[0096] In this embodiment of the invention, the evaluation information includes at least one of the following parameters: acquisition success rate, accuracy of the acquisition location points predicted by the image processing model, and accuracy of the 3D reconstruction result predicted by the image processing model. The acquisition success rate indicates the proportion of successful object acquisitions out of the total number of acquisitions during the object acquisition process performed by the object acquisition device 102 based on the output results. It is worth noting that the above evaluation information is merely a possible example and does not constitute a specific limitation. In some possible implementations, the evaluation information can be a weighted fusion of multiple parameters, including at least any two of the following: acquisition success rate, accuracy of the acquisition location points predicted by the image processing model, and accuracy of the 3D reconstruction result predicted by the image processing model. In the weighted fusion process, the weight of each parameter can be determined in conjunction with the importance of the parameter, for example, they can be the same.
[0097] In one possible scenario, the output of the image processing model includes the 3D reconstruction results of the identified objects. Correspondingly, the evaluation information can include the accuracy of the 3D reconstruction results predicted by the image processing model. In specific implementation, a 3D model library can be constructed, which includes 3D models of moving objects on the object conveying device 101. For the identified object, the control center 110 compares the similarity between the 3D reconstruction results output by the image processing model and the 3D models of the object in the 3D model library to obtain the 3D reconstruction error corresponding to the identified object. Then, the 3D reconstruction errors corresponding to the identified objects over a period of time are statistically analyzed. Based on the statistical 3D reconstruction errors, the accuracy of the 3D reconstruction results predicted by the image processing model is determined. For example, the accuracy can be the mean, mean square error, etc. of the statistical 3D reconstruction errors, and the specific accuracy calculation formula can be designed according to actual needs.
[0098] In one possible scenario, the output of the image processing model includes the acquisition location points of the identified objects. Correspondingly, the evaluation information can include the accuracy of the acquisition location points predicted by the image processing model. In specific implementations, the acquisition location point errors corresponding to the identified objects over a period of time are statistically analyzed. Based on these statistical errors over a period of time, the accuracy of the acquisition location points predicted by the image processing model is obtained. For example, the accuracy can be the mean or mean square error of the statistical acquisition location point errors. The specific accuracy calculation formula can be designed according to actual needs.
[0099] There are two methods for determining the location point error:
[0100] In one possible implementation of this scenario, a 3D model library can be constructed, which includes 3D models of moving objects on the object conveying device 101 and the optimal acquisition position points of the 3D models. For the identified object, the control center 110 compares the similarity between the acquisition position points output by the image processing model and the optimal acquisition position points of the object in the 3D model library to obtain the acquisition position point error corresponding to the identified object.
[0101] In another possible implementation of this scenario, the output of the image processing model includes a score for the acquisition location point of the identified object in the acquisition area 1025, and the evaluation information may include the accuracy of the acquisition location point predicted by the image processing model. In a specific implementation, the control center 110 determines the acquisition result corresponding to the acquisition location point based on the acquisition location point output by the image processing model. The acquisition result is the result of the object acquisition device 102 acquiring the object according to the identified acquisition location point. If the acquisition result indicates successful capture, the acquisition location point error indicates the difference between the full score and the score of the acquisition location point; if the acquisition result indicates capture failure, the acquisition location point error is the score of the acquisition location point.
[0102] In one possible scenario of this implementation, the output of the image processing model may include an acquisition success rate. The acquisition success rate is the ratio of the number of times the object acquisition device 102 successfully acquires an object to the total number of acquisitions, based on the output of the image processing model. In specific implementation, the control center 110 determines the acquisition result corresponding to the output result based on the image processing model's output; it then statistically analyzes the acquisition results to determine the acquisition success rate. For example, based on the acquisition location point output by the image processing model and the corresponding acquisition result, if the acquisition result indicates successful acquisition, the number of successful acquisitions is incremented by 1; if the acquisition result indicates failed acquisition, the number of successful acquisitions is incremented by 0.
[0103] The control center 110 determines the acquisition result corresponding to the output result based on the output result of the image processing model in the following ways: the control center 110 sends the location description information of the acquisition location point of the identified object to the object acquisition device 102; the object acquisition device 102 performs actions according to the acquisition location point of the identified object to determine the acquisition result; the object acquisition device 102 sends the acquisition result to the control center 110; and the control center 110 establishes the association between the acquisition result and the acquisition location point of the identified object.
[0104] In this embodiment of the invention, the control device 1021 of the object acquisition device 102 sends a control command to the acquisition device 1022 when it determines that an object exists at the identified acquisition location point. The acquisition device 1022 then performs an action according to the control command to acquire the object at the acquisition location point. In some possible implementations, the acquisition device 1022, such as the robotic arm 1023 or suction cup 1024, may be equipped with a sensor to determine the position of the object acquisition. The sensor is connected to the control device 1021; for example, the sensor can be a pressure sensor. Based on the signal sent by the sensor, the control device 1021 can determine whether the acquisition device 1022 has grasped the object and obtain the acquisition result.
[0105] In practical applications, the object acquisition device 102 is equipped with an object detection sensor, such as an activated radar. The object detection sensor is connected to the control device 1021. Based on the information sent by the object detection sensor, when the control device 1021 predicts that an object exists at the acquisition location point, it controls the acquisition device 1022 to perform actions based on the acquisition location point.
[0106] The above-described method for determining the acquisition result is merely an example and does not constitute a specific limitation. In other possible cases, the control center 110 sends the object type and location description information of the acquisition location point to the object acquisition device 102. Based on the identified acquisition location point and object type, the control device 1021 of the object acquisition device 102 sends a control command to the acquisition device 1022 when it determines that an object of the object type exists at the acquisition location point. The acquisition device 1022 then performs an action according to the control command and grabs the object according to the acquisition location point.
[0107] Step 204: Based on the evaluation information, the control center 110 adjusts the first pose of the camera 103 to obtain the second pose of the camera 103.
[0108] In some possible embodiments, the control center 110 adjusts the first pose of the camera 103 based on evaluation information when it determines that the first pose needs to be adjusted, thereby obtaining the second pose of the camera 103.
[0109] In one implementation of this embodiment, the control center 110 determines that the first pose needs adjustment when the accuracy of the image processing model is low, based on evaluation information. Exemplarily, there are several ways to determine that the accuracy of the image processing model is low based on evaluation information:
[0110] In implementation method 1, the evaluation information consists of several parameters. If each parameter in the evaluation information is greater than or equal to a preset threshold, the accuracy of the image processing model will be low.
[0111] In implementation method 2, the evaluation information can be a weighted fusion of multiple parameters. However, when the evaluation information is greater than or equal to the threshold, the accuracy of the image processing model is relatively low.
[0112] In some possible embodiments, the control center 110 adjusts the first pose of the camera 103 to obtain the second pose of the camera 103 by: starting from the first pose, performing pose search according to the first scale step and / or the first angle step to determine multiple first candidate poses of the camera 103; and determining the second pose of the camera 103 from the multiple first candidate poses.
[0113] In one example of this embodiment, the first scale step and the first angle step are preset step sizes. In a specific implementation, the control center 110 can set a scale step reduction and an angle step reduction. The first scale step can be the difference between the second scale step and the scale step reduction, and the first angle step can be the difference between the second angle step and the angle step reduction.
[0114] In one example of this implementation, the control center 110 determines the first scale step and / or the first angle step based on the evaluation information. In a specific implementation, the control center 110 can set up a mapping table, with the table entries being the evaluation information, scale step reduction, and angle step reduction. The scale step reduction indicates the difference between any two scale steps, and the angle step reduction indicates the difference between any two angle steps. Subsequently, the control center 110 can match rows in the mapping table based on the evaluation information to determine the scale step reduction and angle step reduction in the row. Based on the scale step reduction and angle step reduction, a fine-grained search for the pose of the camera 103 is performed. For example, in a scenario where the initial pose of the camera 103 is taken as the starting point and the pose search is performed according to the second scale step and / or the second angle step to obtain the first pose, the difference between the second scale step and the scale step reduction is used as the first scale step, and the difference between the second angle step and the angle step reduction is used as the first angle step.
[0115] In one possible scenario, in step 201, the control center 110 starts from the initial pose of the camera 103 and performs pose search according to the second scale step and / or the second angle step to obtain multiple second candidate poses of the camera 103; from the multiple second candidate poses, the first pose of the camera 103 is determined; then the first scale step is smaller than the second scale step, and the first angle step is smaller than the second angle step, thereby realizing fine-grained pose search. For example, in step 201, the control center 110 starts from the initial pose of the camera 103 and performs pose search according to the second scale step and the second angle step; then in step 204, the control center 110 starts from the first pose of the camera 103 and performs pose search according to the second scale step and the second angle step. In one possible case, the first scale step is less than the second scale step, and the first angle step is less than the second angle step; in another possible case, the first scale step is equal to the second scale step, and the first angle step is less than the second angle step; in yet another possible case, the first scale step is less than the second scale step, and the first angle step is equal to the second angle step.
[0116] In one possible implementation of this embodiment, the method for determining multiple first candidate poses of camera 103 is as follows: starting from the first pose, traversing three positional degrees of freedom according to a first scale step size to obtain multiple first candidate poses; for each first candidate pose, traversing three rotational degrees of freedom according to a first angle step size to continue obtaining multiple first candidate poses.
[0117] In one possible implementation of this embodiment, the control center 110 determines the score of the image captured by the camera 103 when it is in each of the plurality of first candidate poses; and uses the first candidate pose with the highest score as the second pose of the camera 103. For example, the method for determining the score of the image captured by the camera 103 when it is in the first candidate pose is as follows: determining text, which describes the shooting features and / or object features of the object captured by the camera 103; determining the similarity between the image captured by the camera when it is in the first candidate pose and the text; and determining the score of the image captured by the camera 103 when it is in the first candidate pose based on the similarity. For details, please refer to the relevant content in step 201 regarding determining the first pose of the camera 103 from the plurality of second candidate poses, the difference being that the second candidate pose is replaced by the first candidate pose, and the first pose is replaced by the second pose, which will not be repeated here.
[0118] Subsequently, the camera 103 is manually installed according to the second pose, or the control center 110 sends the second pose of the camera 103 to the camera 103, and the camera 103 adjusts its own pose.
[0119] In this scheme, the camera takes a single-view shot of the object to save shooting time; at the same time, the camera pose is adjusted according to the accuracy evaluation of the camera pose by the image processing model, so that when the image processing model processes the single-view image taken by the camera, the output of the image processing model can guide the object acquisition device to successfully acquire the object.
[0120] Figure 2 The embodiments shown are merely basic embodiments of the method of the present invention. Other preferred embodiments of the method can be obtained by making certain optimizations and extensions based on them.
[0121] In one embodiment, this embodiment, based on the foregoing embodiments, provides a more detailed description and a certain degree of optimization of the camera pose adjustment process. For example... Figure 3 As shown, the method in this embodiment includes:
[0122] Step 301: Terminal 120 sends job information to control center 110. The job information is used to indicate the transmission area 1011 of object transmission device 101, the acquisition area 1025 of object acquisition device 102, and the camera parameters of camera 103.
[0123] The operational information may include operational parameters of object acquisition device 101, operational parameters of object acquisition device 102, and camera parameters of camera 103. For example, the operational parameters of object acquisition device 101 may be the location of the transport area 1011 used to transport objects, such as the length, width, and height of the transport device. For example, the operational parameters of object acquisition device 102 may be the acquisition area 1025, such as the activity drive of acquisition device 1022. For example, the camera parameters of camera 103 may be intrinsic parameters, extrinsic parameters, and field of view.
[0124] Step 302: Based on the operation information, the control center 110 determines the initial pose of the camera 103.
[0125] In this embodiment of the invention, an initial pose is selected based on job information and text. The text is used to describe the shooting features and / or object features of the object captured by camera 103.
[0126] Step 303: Starting from the initial pose of camera 103, control center 110 performs pose search according to the second scale step size and / or the second angle step size to obtain multiple second candidate poses of camera 103; from multiple second candidate poses, determine the first pose of camera 103.
[0127] For details, please refer to the description in step 201 above, which will not be repeated here.
[0128] Step 304: The control center 110 inputs the image captured by the camera 103 in the first pose into the image processing model to obtain the output result of the image processing model. The output result is used to indicate the three-dimensional reconstruction result of the identified object and / or the acquisition position point of the identified object.
[0129] For details, please refer to the description of step 202, which will not be repeated here.
[0130] Step 305: Based on the output of the image processing model, the control center 110 determines the evaluation information, which is used to indicate the accuracy evaluation of the image processing model in the first pose.
[0131] For details, please refer to the description of step 203, which will not be repeated here.
[0132] Step 306: Based on the evaluation information, the control center 110 determines the first scale step size and / or the first angle step size, wherein the first scale step size is smaller than the second scale step size and the first angle step size is smaller than the second angle step size.
[0133] In one possible embodiment, the control center 110 can determine a first scale step size and / or a first angle step size when the evaluation information indicates that the accuracy of the image processing model is low. Specifically, the control center 110 can set a scale step size reduction and an angle step size reduction, whereby the first scale step size can be the difference between the second scale step size and the scale step size reduction, and the first angle step size can be the difference between the second angle step size and the angle step size reduction.
[0134] In one possible embodiment, the control center 110 can determine the first scale step size and / or the first angle step size based on the evaluation information. In a specific implementation, the control center 110 can set up a mapping table, the entries of which are the evaluation information, the scale step size reduction, and the angle step size reduction. Subsequently, the control center 110 can match the rows in the mapping table based on the evaluation information, determine the scale step size reduction and the angle step size reduction in the row, and take the difference between the second scale step size and the scale step size reduction as the first scale step size, and take the difference between the second angle step size and the angle step size reduction as the first angle step size.
[0135] Step 307: Starting from the first pose, the control center 110 performs pose search according to the first scale step and / or the first angle step to determine multiple first candidate poses of the camera; from the multiple first candidate poses, the second pose of the camera 103 is determined.
[0136] The control center 110 determines multiple first candidate poses of camera 103 in the following way: starting from the first pose, it traverses three positional degrees of freedom according to the first scale step size to obtain multiple first candidate poses. For each first candidate pose, it traverses three rotational degrees of freedom according to the first angle step size to continue obtaining multiple first candidate poses.
[0137] The control center 110 determines the second pose of the camera 103 from multiple first candidate poses by: for each first candidate pose, determining the score of the image captured by the camera 103 when it is in that first candidate pose; and using the first candidate pose with the highest score as the second pose of the camera 103. For example, the score of the image captured by the camera 103 when it is in the first candidate pose is determined by: determining text, which describes the shooting features and / or object features of the object captured by the camera 103; determining the similarity between the image captured by the camera when it is in the first candidate pose and the text; and determining the score of the image captured by the camera 103 when it is in the first candidate pose based on the similarity. For details, please refer to step 201, which describes determining the first pose of the camera 103 from multiple second candidate poses, the difference being that the second candidate pose is replaced by the first candidate pose, and the first pose is replaced by the second pose; this will not be repeated here.
[0138] Step 308: The control center 110 inputs the image captured by the camera 103 in the second pose into the image processing model to obtain the output result of the image processing model. The output result is used to indicate the three-dimensional reconstruction result of the identified object and / or the acquisition position point of the identified object.
[0139] For details, please refer to the description of step 202, which will not be repeated here. The difference is that the first pose is replaced with the second pose.
[0140] Step 309: Based on the output of the image processing model, the control center 110 determines the evaluation information, which is used to indicate the accuracy evaluation of the image processing model in the second pose.
[0141] For details, please refer to the description of step 203, which will not be repeated here. The difference is that the first pose is replaced with the second pose.
[0142] Step 310: Based on the evaluation information, the control center 110 determines the third scale step size and / or the third angle step size, wherein the third scale step size is smaller than the first scale step size and the third angle step size is smaller than the first angle step size.
[0143] In one possible embodiment, the control center 110 can determine a third scale step and / or a third angle step when the evaluation information indicates that the accuracy of the image processing model is low. Specifically, the control center 110 can set a scale step reduction and an angle step reduction, the third scale step can be the difference between the first scale step and the scale step reduction, and the third angle step can be the difference between the first angle step and the angle step reduction.
[0144] In one possible embodiment, the control center 110 can determine the third scale step size and / or the third angle step size based on the evaluation information. In a specific implementation, the control center 110 can set up a mapping table, the entries of which are the evaluation information, the scale step size reduction, and the angle step size reduction. Subsequently, the control center 110 can match the rows in the mapping table based on the evaluation information, determine the scale step size reduction and the angle step size reduction in the row, and take the difference between the first scale step size and the scale step size reduction as the third scale step size, and take the difference between the first angle step size and the angle step size reduction as the third angle step size.
[0145] Step 311: Starting from the second pose, the control center 110 performs pose search according to the third scale step and / or the third angle step to determine multiple third candidate poses of the camera 103; from the multiple third candidate poses, the third pose of the camera 103 is determined.
[0146] For details, please refer to the description of step 307. The difference is that the second candidate pose is replaced with the third candidate pose, and the first pose is replaced with the second pose. This will not be elaborated further.
[0147] Repeat steps 308 to 311 to narrow the pose search range of camera 103 until the pose of camera 103 converges.
[0148] In this scheme, the camera takes a single-view shot of the object to save shooting time; at the same time, the camera pose is adjusted according to the accuracy evaluation of the camera pose by the image processing model, so that when the image processing model processes the single-view image taken by the camera, the output of the image processing model can guide the object acquisition device to successfully acquire the object.
[0149] Based on the camera pose adjustment method provided above, the specific application of the camera pose adjustment method will be explained. Figure 4 This is a schematic diagram illustrating a specific application of a camera pose adjustment method provided for the implementation of this invention. For example... Figure 4 As shown, the specific content includes:
[0150] The control center 110 acquires the operation information and determines the initial pose of the camera 103 based on the operation information; according to the current scale step size and the current angle step size, it performs pose search starting from the initial pose of the camera 103 to obtain multiple candidate camera poses; for the image under each candidate camera pose, it calculates the similarity between the image and the text to obtain a score for each candidate camera pose; and it takes the candidate camera pose with the highest score as the current pose of the camera 103.
[0151] Camera 103 captures an image in the current pose; control center 110 inputs the image captured by camera 103 in the current pose into the image processing model to obtain the output result of the image processing model. The output result is used to determine the accuracy evaluation of the image processing model (accuracy of 3D reconstruction results, accuracy of obtaining position points, and success rate of obtaining points); based on the accuracy evaluation of the image processing model, the current scale step size and the current angle step size are reduced to obtain the updated current scale step size and the updated current angle step size.
[0152] Based on the updated current scale step size and the updated current angle step size, a pose search is performed starting from the current pose of camera 103 to obtain multiple candidate camera poses; for each candidate camera pose, the similarity between the image and the text is calculated to obtain a score for each candidate camera pose; the candidate camera pose with the highest score replaces the current pose of camera 103.
[0153] The process is repeated, continuously narrowing the pose search range of camera 103 until the current pose of camera 103 converges.
[0154] Based on the camera pose adjustment method provided above, the architecture of the camera pose adjustment method is explained in conjunction with actual application scenarios.
[0155] Application scenario: Object acquisition device 102: a robotic arm, object conveying device 101: a moving conveyor belt, a single camera 103. The robotic arm uses a suction cup 1204 to pick up objects on the moving conveyor belt and remove them from the conveyor belt.
[0156] System Architecture: Figure 5a and Figure 5b This is a schematic diagram illustrating a specific application of a camera pose adjustment method provided for the implementation of this invention. For example... Figure 5a and Figure 5b As shown, the specific content is as follows:
[0157] On the data upload interface, users first need to upload environmental parameters such as conveyor belt parameters, robotic arm working area, and camera parameters. Specifically, users need to first measure the parameters in their user scene: conveyor belt parameters (length, width, and height), robotic arm working 3D area, and camera parameters (intrinsic and extrinsic parameters, field of view). Then, users need to upload these parameters to the data upload interface.
[0158] The camera module, based on user-uploaded data (including conveyor belt parameters, robotic arm working area, and camera parameters), calculates a suitable camera pose for camera 103, ensuring that camera 103 can observe the most objects on the conveyor belt with the most comprehensive detail. Specifically, the camera module calculates the initial pose of camera 103 based on the user-uploaded data, and then searches for camera poses using scale and angle steps to obtain a list of camera poses. This list indicates the scores of multiple camera poses, and the user selects one based on the actual deployment difficulty. Alternatively, the camera module can directly output the camera pose with the highest score. The score indicates the rating of the image captured by camera 103 when it is in the specified camera pose; for example, it could be a similarity score between an image and text, where the text describes the shooting features and / or object features captured by camera 103. Specifically, the camera pose includes a 3D position vector and 3D Euler angles. Starting from an initial pose, a traversal based on scale and angle steps is performed to obtain multiple candidate camera poses. For each candidate camera pose, the CLIP visual feature obtained from the image captured by the camera at that pose is compared with the text prompt to calculate a similarity score, which is then used as the score for that candidate camera pose. It should be noted that users need to input text prompts according to their needs and actual situation. For example, if the objects the user needs to capture are mostly regular-shaped objects such as cylinders or boxes, they can input "sideways observation of the conveyor belt" to better observe the object geometry. The CLIP model can obtain the CLIP visual feature based on this text prompt and the image captured by the camera at the candidate camera pose, and then calculate the similarity score between the CLIP visual feature and the text prompt, which is used as the score for the candidate camera pose.
[0159] After deploying the existing image processing model at the given camera pose, the accuracy evaluation of the image processing model (accuracy of 3D reconstruction results, accuracy of location point acquisition, and acquisition success rate) is input into the camera module. Based on the accuracy evaluation of the image processing model, the camera module continuously refines the search range (i.e., continuously reduces the scale step and / or angular step), improving the accuracy of the image processing model by fine-tuning the camera pose, ultimately obtaining the optimal camera pose. Specifically, after the image processing model is deployed, it is tested to obtain its accuracy and acquisition success rate. Using the accuracy and acquisition success rate of the image processing model as input to the camera module, the module then initializes the current pose of camera 103 and performs a fine-grained search (smaller scale step and / or angular step) to obtain multiple candidate camera poses. The similarity score of the candidate camera poses is calculated using the same method described above, resulting in a score for each candidate camera pose. The camera pose with the highest score is then used as the latest camera pose for camera 103. This process is repeated, continuously fine-tuning the camera pose to achieve the best image processing model results.
[0160] This invention also provides a management platform, which can be found below. Figure 6 , Figure 6 This is a schematic diagram of a camera pose adjustment device provided in an embodiment of the present invention. The camera pose adjustment device may include at least:
[0161] The pose determination module 601 is used to determine the first pose of the camera; the camera captures an object in the conveying area of the object conveying device in the first pose.
[0162] The prediction module 602 is used to input the image captured by the camera in the first pose into the image processing model and determine the output result of the image processing model; the output result is used to indicate the three-dimensional reconstruction result of the object and / or the acquisition position point of the object; the output result is used to guide the object acquisition device to acquire the object on the transmission area;
[0163] Evaluation module 603 is used to determine evaluation information based on the output of the image processing model, the evaluation information being used to indicate the accuracy evaluation of the image processing model in the first pose;
[0164] The pose adjustment module 604 is used to adjust the first pose based on the evaluation information to obtain the second pose of the camera.
[0165] It is worth noting that all of the above modules can achieve the corresponding technical functions, and the embodiments of the present invention do not limit them.
[0166] The pose determination module 601, prediction module 602, evaluation module 603, and pose adjustment module 604 can all be implemented in software or hardware. For example, the implementation of the pose determination module 601 will be described below. Similarly, the implementation of the prediction module 602, evaluation module 603, and pose adjustment module 604 can refer to the implementation of the pose determination module 601.
[0167] As an example of a software functional unit, the pose determination module 601 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, the pose determination module 601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone or in different availability zones, each availability zone including one data center or multiple geographically proximate data centers. Typically, a region may include multiple availability zones.
[0168] As an example of a hardware functional unit, the pose determination module 601 may include at least one computing device, such as a server. Alternatively, the pose determination module 601 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0169] It should be noted that, in other embodiments, the pose determination module 601 can be used to execute any step in the camera pose adjustment method, the prediction module 602 can be used to execute any step in the camera pose adjustment method, and the evaluation module 603 and pose adjustment module 604 can be used to execute any step in the camera pose adjustment method. The steps implemented by the pose determination module 601, prediction module 602, evaluation module 603, and pose adjustment module 604 can be specified as needed. By implementing different steps in the camera pose adjustment method through the pose determination module 601, prediction module 602, evaluation module 603, and pose adjustment module 604, all functions of the management platform can be realized.
[0170] The methods, management platforms, and systems of the embodiments of the present invention have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of the present invention, relevant equipment for cooperating in implementing the above solutions is also provided below.
[0171] This invention provides a computing device, as detailed below. Figure 7 , Figure 7 This is a schematic diagram of a computing device for a camera pose adjustment method provided in an embodiment of the present invention. The computing device 700 includes: a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other via the bus 702. The computing device 700 can be a server or a terminal device. It should be understood that the present invention does not limit the number of processors and memories in the computing device 700.
[0172] The 702 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus 702 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 702 may include a path for transmitting information between various components of the computing device 700 (e.g., memory 706, processor 704, communication interface 708).
[0173] Processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0174] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0175] The memory 706 stores executable program code, which the processor 704 executes to implement the functions of the pose determination module 601 and the prediction module 602, thereby realizing the camera pose adjustment method. In other words, the memory 706 stores instructions from the management platform for executing the camera pose adjustment method.
[0176] The communication interface 708 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.
[0177] This invention also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0178] Please see below Figure 8 , Figure 8 This is a schematic diagram of a computing device cluster for a camera pose adjustment method according to an embodiment of the present invention. Figure 8 As shown, the computing device cluster includes at least one computing device 700, and the memory 706 of one or more computing devices 700 in the computing device cluster may store the same management platform for executing camera pose adjustment methods.
[0179] In some possible implementations, one or more computing devices 700 in the computing device cluster can also be used to execute some of the instructions used by the management platform to perform the camera pose adjustment method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions used by the management platform to perform the camera pose adjustment method.
[0180] It should be noted that the memory 706 in different computing devices 700 within the computing device cluster can store different instructions for executing some functions of the management platform. That is, the instructions stored in the memory 706 of different computing devices 700 can implement the functions of one or more modules among the pose determination module 601, prediction module 602, evaluation module 603, and pose adjustment module 604.
[0181] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the camera pose adjustment method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the camera pose adjustment method.
[0182] Please see below Figure 8 , Figure 8 This is another schematic diagram of the computing device cluster for the camera pose adjustment method according to an embodiment of the present invention. For example... Figure 8 As shown, two computing devices 700A and 700B are connected via a communication interface 709. The memory in computing device 700A stores instructions for executing the pose determination module 601. The memory in computing device 700B stores instructions for executing the prediction module 602, evaluation module 603, and pose adjustment module 604. In other words, the memory 706 of computing devices 700A and 700B jointly stores the instructions used by the management platform to execute the camera pose adjustment method.
[0183] Figure 9 The connection method between the computing device clusters shown can be based on the consideration that the camera pose adjustment method provided by this invention may simultaneously process multiple job scenarios, requiring a large amount of data transmission to the pose determination module 601. Considering the amount of data transmission, in order to avoid overloading the computing device 700A, the functions of the prediction module 602, the evaluation module 603, and the pose adjustment module 604 are delegated to the computing device 700B.
[0184] It should be understood that Figure 9 The functions of the computing device 700A shown can also be performed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be performed by multiple computing devices 700.
[0185] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the camera pose adjustment method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the camera pose adjustment method.
[0186] This invention also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to perform the above-described method for performing camera pose adjustment applied to a management platform.
[0187] This invention also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-described method applied to a management platform for performing camera pose adjustment.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
[0189] Those skilled in the art will clearly understand that the specific working process of the system, management platform or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0190] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0191] This invention also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to perform the above-described method for performing camera pose adjustment applied to a management platform.
[0192] This invention also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-described method applied to a management platform for performing camera pose adjustment.
[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
[0194] Those skilled in the art will clearly understand that the specific working process of the system, management platform or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0195] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for adjusting camera pose, characterized in that, The method includes: Determine the first pose of the camera; the camera in the first pose photographs the object on the conveying area of the object conveying device; The image captured by the camera in the first pose is input into the image processing model, and the output result of the image processing model is determined; the output result is used to indicate the three-dimensional reconstruction result of the object and / or the acquisition position point of the object; the output result is used to guide the object acquisition device to acquire the object on the transmission area; Based on the output of the image processing model, evaluation information is determined, which is used to indicate the accuracy evaluation of the image processing model in the first pose; Based on the evaluation information, the first pose is adjusted to obtain the second pose of the camera.
2. The method according to claim 1, characterized in that, The step of adjusting the first pose based on the evaluation information to obtain the second pose of the camera includes: Based on the evaluation information, if it is determined that the first pose needs to be adjusted, the first pose is adjusted to obtain the second pose of the camera.
3. The method according to claim 1 or 2, characterized in that, The step of adjusting the first pose to obtain the second pose of the camera includes: Starting from the first pose, pose search is performed according to the first scale step size and / or the first angle step size to determine multiple first candidate poses of the camera; The second pose of the camera is determined from the plurality of first candidate poses.
4. The method according to claim 3, characterized in that, The first scale step size and / or the first angle step size are determined by the evaluation information.
5. The method according to claim 3 or 4, characterized in that, Determining the second pose of the camera from the plurality of first candidate poses includes: For each of the plurality of first candidate poses, determine the score of the image captured by the camera when it is in each first candidate pose; The first candidate pose with the highest score is used as the second pose of the camera.
6. The method according to claim 5, characterized in that, The scoring of the image captured when the camera is in the first candidate pose includes: The text is defined as describing the shooting features and / or object features captured by the camera. Determine the similarity between the image captured by the camera when it is in the first candidate pose and the text; The score of the image captured by the camera when it is in the first candidate pose is determined based on the similarity.
7. The method according to any one of claims 3 to 6, characterized in that, Determining the first pose of the camera includes: Starting from the initial pose of the camera, pose search is performed according to the second scale step size and / or the second angle step size to obtain multiple second candidate poses of the camera; the second angle step size is greater than the first angle step size; the second scale step size is greater than the first scale step size. The first pose of the camera is determined from the plurality of second candidate poses.
8. The method according to any one of claims 1 to 7, characterized in that, The first pose or the initial pose is determined based on job information, which is used to indicate the transport area of the object transport device, the acquisition area of the object acquisition device, and the camera parameters of the camera; the acquisition area includes at least a portion of the transport area.
9. The method according to any one of claims 1 to 8, characterized in that, The evaluation information includes at least one of the following: acquisition success rate, accuracy of the acquisition location points predicted by the image processing model, and accuracy of the three-dimensional reconstruction results predicted by the image processing model; the acquisition success rate is used to indicate the proportion of the number of times the object acquisition device successfully acquires an object during the process of acquiring an object based on the output results.
10. A camera pose adjustment device, characterized in that, The device includes: The pose determination module is used to determine the first pose of the camera; the camera captures an object in the transport area of the object transport device in the first pose. The prediction module is used to input the image captured by the camera in the first pose into the image processing model and determine the output result of the image processing model; the output result is used to indicate the three-dimensional reconstruction result of the object and / or the acquisition position point of the object; the output result is used to guide the object acquisition device to acquire the object on the transmission area; An evaluation module is used to determine evaluation information based on the output of the image processing model, wherein the evaluation information is used to indicate the accuracy evaluation of the image processing model in the first pose; The pose adjustment module is used to adjust the first pose based on the evaluation information to obtain the second pose of the camera.
11. The apparatus according to claim 10, characterized in that, The pose adjustment module is used to adjust the first pose based on the evaluation information, when it is determined that the first pose needs to be adjusted, to obtain the second pose of the camera.
12. The apparatus according to claim 10 or 11, characterized in that, The pose adjustment module is used to perform pose search starting from the first pose, according to the first scale step and / or the first angle step, to determine multiple first candidate poses of the camera. The second pose of the camera is determined from the plurality of first candidate poses.
13. The apparatus according to claim 12, characterized in that, The first scale step size and / or the first angle step size are determined by the evaluation information.
14. The apparatus according to claim 12 or 13, characterized in that, The pose adjustment module is used to determine a score for the image captured by the camera when it is in each of the plurality of first candidate poses. The first candidate pose with the highest score is used as the second pose of the camera.
15. The apparatus according to claim 14, characterized in that, The pose adjustment module is used to determine text, which describes the shooting features and / or object features captured by the camera; determine the similarity between the image captured by the camera in the first candidate pose and the text; and determine a score for the image captured by the camera in the first candidate pose based on the similarity.
16. The apparatus according to any one of claims 12 to 15, characterized in that, The pose determination module is used to perform pose search starting from the initial pose of the camera, according to the second scale step and / or the second angle step, to obtain multiple second candidate poses of the camera. The second angle step size is greater than the first angle step size; The second scale step size is greater than the first scale step size; the first pose of the camera is determined from the plurality of second candidate poses.
17. The method according to any one of claims 10 to 16, characterized in that, The first pose or the initial pose is determined based on job information, which is used to indicate the transport area of the object transport device, the acquisition area of the object acquisition device, and the camera parameters of the camera; the acquisition area includes at least a portion of the transport area.
18. The method according to any one of claims 10 to 17, characterized in that, The evaluation information includes at least one of the following: acquisition success rate, accuracy of the acquisition location points predicted by the image processing model, and accuracy of the three-dimensional reconstruction results predicted by the image processing model; the acquisition success rate is used to indicate the proportion of the number of times the object acquisition device successfully acquires an object during the process of acquiring an object based on the output results.
19. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 9.
20. A computer program product containing instructions, characterized in that, When the instruction is executed by a cluster of computer devices, the cluster of computer devices causes the cluster of computer devices to perform the method as described in any one of claims 1 to 9.
21. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 9.