Image acquisition position and orientation determination device, image acquisition position and orientation determination system, robot, and image acquisition position and orientation determination method
The imaging position and orientation determination device addresses the challenge of robotic imaging in diverse environments by using operator data and advanced algorithms to determine optimal positions and orientations, enhancing efficiency and safety in robotic inspections.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HITACHI LTD
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Existing technologies do not provide a method for a robot to accurately image an object in an environment different from that of an operator, lacking a systematic approach to determine optimal imaging positions and orientations.
An imaging position and orientation determination device that acquires imaging information from an operator and robot environment data, generates interference conditions, calculates candidate positions and orientations, evaluates them based on similarity, and determines the optimal imaging position and orientation for the robot using Novel View Synthesis and path planning algorithms.
Enables accurate imaging by a robot in a different environment, reducing the need for operator intervention and inspection time, and eliminating the need for hazardous manual inspections.
Smart Images

Figure 2026071919000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an imaging position and orientation determination device, an imaging position and orientation determination system, a robot, and an imaging position and orientation determination method.
Background Art
[0002] As a technology that utilizes a robot for inspection work of an object, for example, the technology described in Patent Document 1 is known. That is, Patent Document 1 describes "estimating an optimal imaging position at which an image capable of inspecting the inspection target can be taken based on the position of the facility to be inspected and the specifications of the inspection camera."
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Patent Document 1 describes the estimation of the optimal imaging position when a robot takes an image with a camera. However, for example, it does not describe a technology applicable to the case where a robot takes an image of an object in an environment different from the imaging environment of an operator based on the imaging information when the operator takes an image of the object with a camera.
[0005] Therefore, an object of the present disclosure is to provide an imaging position and orientation determination device or the like that enables an object to be appropriately imaged using a robot.
Means for Solving the Problems
[0006] To solve the aforementioned problems, the imaging position and posture determination device according to this disclosure comprises: a data acquisition unit that acquires imaging information when an operator images an object with a first camera and robot environment information including three-dimensional information of the surroundings when a robot images the object with a second camera; an interference condition generation unit that generates interference conditions between the robot and the environment when the robot images the object with the second camera based on the robot environment information; and a position and posture calculation unit that calculates candidate imaging position and candidate imaging posture when the robot images the object with the second camera based on the imaging information and the interference conditions, wherein the imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera is imaging; an evaluation unit that evaluates the candidate imaging position and candidate imaging posture based on the data; and a determination unit that determines the imaging position and imaging posture when the robot images the object with the second camera based on the evaluation results of the evaluation unit. [Effects of the Invention]
[0007] According to this disclosure, it is possible to provide an imaging position and orientation determination device, etc., that uses a robot to appropriately image an object. [Brief explanation of the drawing]
[0008] [Figure 1A] This is an explanatory diagram showing an example of the environment in which an operator has previously photographed a railway vehicle, the target object, with respect to the imaging position and orientation determination device according to the first embodiment. [Figure 1B] This is an explanatory diagram showing an example of the environment in which a robot images a railway vehicle, which is an object, with respect to the imaging position and orientation determination device according to the first embodiment. [Figure 2] This is a functional block diagram of an imaging position and orientation determination system including an imaging position and orientation determination device according to the first embodiment. [Figure 3] This figure shows the hardware configuration of the imaging position and orientation determination device according to the first embodiment. [Figure 4]This is an explanatory diagram showing an example of an imaging position candidate set by the position and orientation calculation unit of the imaging position and orientation determination device according to the first embodiment. [Figure 5A] This is an explanatory diagram showing the position of an object when an operator images it in a first environment, according to the imaging position and orientation determination device according to the first embodiment. [Figure 5B] This is an explanatory diagram showing candidate imaging positions when a robot images an object in a second environment, according to the imaging position and orientation determination device according to the first embodiment. [Figure 6] This is an explanatory diagram illustrating another example showing candidate imaging positions when a robot images an object in a second environment, in the imaging position and orientation determination device according to the first embodiment. [Figure 7] This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the first embodiment. [Figure 8] This is a functional block diagram of an imaging position and orientation determination system including an imaging position and orientation determination device according to the second embodiment. [Figure 9] This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the second embodiment. [Figure 10] This is a functional block diagram of an imaging position and orientation determination system including an imaging position and orientation determination device according to the third embodiment. [Figure 11] This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the third embodiment. [Figure 12] This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the fourth embodiment. [Figure 13] This is a functional block diagram of a robot equipped with an imaging position and orientation determination device according to a modified example. [Modes for carrying out the invention]
[0009] ≪First Embodiment≫ Below, we will first briefly describe the imaging environment of the object to be inspected, and then describe in detail the imaging position and orientation determination device 10 (see Figure 2) according to the first embodiment. Hereinafter, as an example, the case where the object to be inspected is a railway vehicle V1 (see FIGS. 1A and 1B) will be described, but it is not limited thereto. For example, the devices and facilities to be inspected may be, in addition to ships and aircraft, power transmission facilities, power distribution facilities, substation facilities, communication facilities, air conditioning facilities, refrigeration facilities, medical facilities, gas facilities, water supply facilities, railway tracks and electric wires, roads, plants, and the like. Examples of the above-mentioned plants include power generation plants, manufacturing plants, chemical plants, water treatment plants, and the like.
[0010] Note that the image obtained by imaging the object to be inspected may be a still image or a moving image. In the following description, still images and moving images are collectively referred to as "images".
[0011] <Examples of objects> FIG. 1A is an explanatory diagram showing an example of the environment when an operator M1 previously images the railway vehicle V1, which is the object, with respect to the imaging position and orientation determination device according to the first embodiment. In the example of FIG. 1A, a state where the railway vehicle V1 is carried into the inspection facility is shown. Also, in FIG. 1A, the illustration of the vehicle body of the railway vehicle V1 is omitted, and the wheels W1 and axles X1 of the railway vehicle V1 are illustrated (the same applies to FIG. 1B). The wheels W1 of the railway vehicle V1 are placed on a pair of rails R1. The rail R1 is supported by a support member E1 extending along the rail R1, and further, the support member E1 is supported by a plurality of support bases F1.
[0012] Below the railway vehicle V1 in the inspection facility, a pit P1, which is the working space of the operator M1, is provided. Then, the operator M1 enters the pit P1 and performs the inspection by looking up at the railway vehicle V1 from below. The operator M1 visually checks, for example, whether there is any looseness or damage in the fastening parts or welded parts of the carriage supporting the vehicle body of the railway vehicle V1.
[0013] In addition, the number of inspection facilities is limited, and the number of railway vehicles that can be brought into one inspection facility is also limited. Therefore, in the past, a great deal of time and effort has been required for the inspection of railway vehicles, and as a result, the inspection work in the pit has become a bottleneck. Therefore, in the first embodiment, first, the operator M1 images the railway vehicle V1, which is the object to be inspected, with the first camera 20. The image obtained by imaging with the first camera 20 is used as a kind of model when the robot 30 (see FIG. 1B) images the railway vehicle V1 with the second camera 31 (see FIG. 1B). Hereinafter, the environment when the operator M1 images the object (for example, the railway vehicle V1) with the first camera 20 is referred to as the "first environment".
[0014] In the example of FIG. 1A, the first camera 20 is installed on the helmet of the operator M1, but the imaging mode can be changed as appropriate. For example, the first camera 20 may be installed on a detachable device such as a head-mounted display. Also, the operator M1 may image while holding the first camera 20 by hand. The image obtained by imaging with the first camera 20 may be a still image such as an RGB color image, or may be a moving image. When the operator M1 images the inspection location of the railway vehicle V1 with the first camera 20, the position (imaging position) and orientation (imaging posture) of the first camera 20 at the time of imaging are also recorded as appropriate.
[0015] FIG. 1B is an explanatory diagram showing an example of the environment when the robot 30 images the railway vehicle V1, which is the object. In addition, in FIG. 1B, the case where the robot 30 images the railway vehicle V1 with the second camera 31 in the yard Y1, which is the operation yard of the railway vehicle V1, is shown. In FIG. 1B, the case where an articulated snake-shaped robot is used as the robot 30 is shown, but the type of the robot 30 can be changed as appropriate. For example, in addition to robots equipped with moving means such as walking legs, wheels, or crawlers (tracks), unmanned flying bodies such as drones may be used as the robot 30.
[0016] As shown in Figure 1B, the robot 30 is equipped with a second camera 31. Based on the imaging position and orientation determined by the imaging position and orientation determination device 10 (see Figure 2), which will be described later, the robot 30 takes images of the railway vehicle V1. The environment in which the robot 30 takes images of the target object (for example, the railway vehicle V1) with the second camera 31 is referred to as the "second environment".
[0017] Thus, the first environment in which worker M1 (see Figure 1A) images the railway vehicle V1 (for example, pit P1 in Figure 1A) and the second environment in which robot 30 images the railway vehicle V1 (for example, yard Y1 in Figure 1B) are different. However, the first embodiment is not limited to cases where the first and second environments are different; it can also be applied when the first and second environments are similar. Furthermore, the railway vehicle V1 in the first environment (see Figure 1A) and the railway vehicle V1 in the second environment (see Figure 1B) do not necessarily need to be completely identical; they can be of the same type (or have a similar structure).
[0018] The general procedure for inspecting railway vehicle V1 is as follows: First, in pit P1 (first environment) as shown in Figure 1A, worker M1 uses the first camera 20 to image the inspection points of railway vehicle V1. Then, using the images obtained by worker M1 as a guide, the imaging position and orientation determination device 10 (see Figure 2), described later, determines the imaging position and orientation for robot 30 to perform imaging in yard Y1 (second environment) as shown in Figure 1B. Next, based on the imaging position and orientation determined by the imaging position and orientation determination device 10, robot 30 uses the second camera 31 to image the inspection points of railway vehicle V1 in yard Y1 (second environment). The images captured by robot 30 are then displayed on a remote terminal (tablet, etc.), and the worker inspects railway vehicle V1 by viewing the screen of that terminal. Note that the worker may view the imaging results of the second camera 31 in real time, or they may view them later.
[0019] <Configuration of the imaging position and orientation determination system> Figure 2 is a functional block diagram of the imaging position and orientation determination system 100, including the imaging position and orientation determination device 10. The imaging position and orientation determination system 100 shown in Figure 2 is a system for determining the imaging position and orientation when a robot 30 (see Figure 1B) images an object (for example, the railway vehicle V1 in Figure 1B) with a second camera 31 (see Figure 1B), and consists of the robot 30 and the imaging position and orientation determination device 10. For example, a computer can be used as such an imaging position and orientation determination device 10. Alternatively, the functions of the imaging position and orientation determination device 10 may be distributed across multiple computers, such as cloud servers or edge servers.
[0020] As shown in Figure 2, the imaging position and orientation determination device 10 comprises a storage unit 11 and a processing unit 12. The storage unit 11 has predetermined programs and data stored in it beforehand. In addition, the storage unit 11 stores robot environment information and imaging information, which will be described later, as well as the calculation results of the processing unit 12 as appropriate. The processing unit 12 executes predetermined processes based on the programs and data stored in the storage unit 11.
[0021] As shown in Figure 2, the processing unit 12 includes a data acquisition unit 121, an interference condition generation unit 122, a position and orientation calculation unit 123, an evaluation unit 124, and a determination unit 125. The data acquisition unit 121 acquires robot environment information and imaging information. Here, "robot environment information" refers to data that includes 3D information of the surroundings (around the robot 30) when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B). Such robot environment information is created in advance based on predetermined 3D surveying and stored in the memory of another computer (not shown).
[0022] In the example shown in Figure 1B, the 3D information of the second environment, yard Y1, is used as robot environment information. The 3D shape and arrangement of the railway vehicle V1, which is the object to be inspected, are also included in the robot environment information. Such robot environment information may include 3D maps, mesh models, or point cloud data as appropriate. The aforementioned 3D map is created based on the well-known SLAM (Simultaneous Localization and Mapping) technique. The mesh model is generated by dividing the object's shape into simple shapes such as triangles and quadrilaterals. Point cloud data is a collection of points acquired by a 3D laser scanner.
[0023] The "imaging information" shown in Figure 2 includes, for example, work-time images, position and orientation information, and worker environment information. Here, "work-time images" refer to images obtained when worker M1 (see Figure 1A) images the target object with the first camera 20 (see Figure 1A). These work-time images are used as data to identify objects (or objects within the field of view of worker M1) that are included in the field of view of the first camera 20 when it takes images. It is assumed that the object to be inspected is captured in the work-time images. The data used to identify objects within the field of view of the first camera 20 may include, for example, line-of-sight information that can be represented on a specific image and image coordinates, from the viewpoint of being able to identify which point in the image the first camera 20 is looking at. Furthermore, the data used to identify objects within the worker M1's field of view may include, for example, the worker M1's position and orientation, as well as line-of-sight information that can be represented in three-dimensional space, from the perspective of identifying the three-dimensional coordinates of the object being gazed upon. Incidentally, by using AR goggles or the like, it is possible to make the field of view of worker M1 and the field of view of the first camera 20 approximately the same range. Therefore, based on the data obtained from the AR goggles or the like, it is possible to identify objects within the field of view of the first camera 20 (i.e., objects within the field of view of worker M1).
[0024] Furthermore, the "position and orientation information" included in the imaging information refers to information indicating the position and orientation of the first camera 20 (see Figure 1A) when operator M1 (see Figure 1A) images an object with the first camera 20 (see Figure 1A). Here, the "position" of the first camera 20 is the three-dimensional position of the first camera 20, and is expressed in predetermined three-dimensional coordinates. The "orientation" of the first camera 20 refers to the orientation of the first camera 20. For example, the orientation of the first camera 20 can be expressed using the well-known Euler angle, which represents the optical axis direction of the lens of the first camera 20. Such position and orientation information is generated, for example, when operator M1 (see Figure 1A) images an object with the first camera 20 (see Figure 1A). For example, a gyro sensor (not shown) may be provided on the first camera 20, and the position and orientation of the first camera 20 may be measured using this gyro sensor.
[0025] Alternatively, instead of positional information, the gaze information of worker M1 (see Figure 1A) may be used. Here, "gaze information" refers to information indicating the starting point and direction of worker M1's gaze. Furthermore, the "worker environment information" included in the imaging information refers to the three-dimensional information surrounding worker M1 (see Figure 1A) when worker M1 images an object with the first camera 20 (see Figure 1A). This type of worker environment information is created in advance based on predetermined three-dimensional surveying.
[0026] The aforementioned imaging information, including images taken during operation, position and orientation information, line of sight information, and worker environment information, is provided as appropriate from, for example, a predetermined computer (not shown) that stores this data. In Figure 2, the information acquired by the data acquisition unit 121 (robot environment information and imaging information) is shown to be output directly to the interference condition generation unit 122, the position and orientation calculation unit 123, and the evaluation unit 124. However, this information may also be stored in the storage unit 11 first and then read out from the storage unit 11 as appropriate.
[0027] The interference condition generation unit 122 shown in Figure 2 generates interference conditions between the robot 30 (see Figure 1B) and the environment when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B), based on robot environment information. Here, "interference conditions" refer to information indicating whether surrounding structures interfere with the robot 30 (whether they come into contact with or collide with the robot 30), as well as whether other structures obstruct the robot 30 when it images an object with the second camera 31. For example, the interference conditions may be expressed as a predetermined numerical range or mathematical formula that specifies the range of positions the robot 30 can take.
[0028] Furthermore, when interference conditions are generated, mesh data or polygon data of the environment may be used as appropriate. For example, the interference condition generation unit 122 determines whether or not there is interference between the robot 30 and the environment based on geometric algorithms such as Bounding Volume Hierarchy or Separating Axis Theorem. Alternatively, the processing unit 12 may divide the 3D space into 3D boxes (voxels) based on the voxel grid method and register whether or not there is an object in each voxel. In this case, the presence or absence of interference between the robot 30 and the environment is determined based on whether or not the position of the robot 30 overlaps with a voxel.
[0029] The position and orientation calculation unit 123 shown in Figure 2 calculates candidate imaging positions and imaging orientations for when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B) based on the imaging information and interference conditions described above. Here, "candidate imaging position" refers to a candidate position for the second camera 31 when the robot 30 images an object with the second camera 31. Also, "candidate imaging orientation" refers to a candidate orientation (in the optical axis direction of the lens) of the second camera 31 when the robot 30 images an object with the second camera 31. Note that the number of candidate imaging positions and imaging orientations may be one pair or multiple pairs.
[0030] The evaluation unit 124 shown in Figure 2 evaluates the candidate imaging position and candidate imaging posture according to a predetermined criteria. Specifically, the evaluation unit 124 first generates candidate images assuming that the robot 30 (see Figure 1B) performs imaging at the predetermined candidate imaging position and posture. The generation of such candidate images is performed based on Novel View Synthesis technology, which generates an image from a new viewpoint from images from multiple viewpoints. More specifically, techniques such as Gaussian Splatting and Nerf (Neural Radiance Fields) are used as appropriate. It is assumed that one candidate image is generated for each pair of candidate imaging position and candidate imaging posture.
[0031] The evaluation unit 124 then calculates the similarity between the work image included in the imaging information (the image obtained by imaging worker M1) and the candidate image obtained when imaging at a predetermined candidate imaging position and imaging posture. The more similar the candidate image is to the work image, the higher the similarity score.
[0032] The evaluation unit 124 calculates the similarity between, for example, an operational image obtained when worker M1 images the railway vehicle V1 in the first environment, pit P1 (see Figure 1A), and a candidate image assuming that robot 30 images the railway vehicle V1 in the second environment, yard Y1 (see Figure 1B). If there are multiple pairs of candidate imaging positions and imaging postures, the similarity between each candidate image corresponding to each pair of candidate imaging positions and imaging postures and the operational image is calculated.
[0033] The determination unit 125 shown in Figure 2 determines the imaging position and orientation when the robot 30 (see Figure 1B) images the object with the second camera 31 (see Figure 1B) based on the similarity evaluation result from the evaluation unit 124. In other words, the determination unit 125 determines the candidate imaging position and orientation corresponding to the candidate image with the highest similarity to the work image as the imaging position and orientation when the robot 30 images the object. Here, "imaging position" refers to the three-dimensional position of the second camera 31 when the robot 30 images the object with the second camera 31. Also, "imaging orientation" refers to the orientation of the second camera 31 when the robot 30 images the object with the second camera 31.
[0034] In this way, the imaging position and orientation determination device 10 determines the imaging position and orientation for the robot 30 (see Figure 1B) when it images the target object, based on the robot environment information and imaging information acquired by the data acquisition unit 121. Information indicating the determination result of the imaging position and orientation determination device 10 is transmitted to the robot 30 by wired or wireless connection. The robot 30 is equipped with a control device (not shown) that performs control based on the imaging position and orientation determined by the imaging position and orientation determination device 10.
[0035] Figure 3 shows the hardware configuration of the imaging position and orientation determination device 10. As shown in Figure 3, the imaging position and orientation determination device 10 has a hardware configuration comprising a processor 10a, RAM 10b (Random Access Memory), ROM 10c (Read Only Memory), HDD 10d (Hard Disk Drive), a communication interface 10e, and an input / output interface 10f, all of which are predeterminedly connected via an internal bus 10g.
[0036] The processor 10a is hardware that functions as the processing unit 12 (see Figure 2) described above. The RAM 10b, ROM 10c, and HDD 10d are hardware that functions as the storage unit 11 (see Figure 2) described above. The processor 10a reads a predetermined program stored in the ROM 10c or HDD 10d and loads it into the RAM 10b, thereby executing a predetermined process.
[0037] The communication interface 10e is used for communication with external devices and networks. For example, the communication interface 10e sends and receives predetermined information to and from the robot 30. The input / output interface 10f is an interface used for data input from the input device 40 and data output to the display device 50. The input device 40 is, for example, a keyboard or mouse, and is used when the user inputs data.
[0038] The display device 50 is, for example, a display that shows the calculation results of the imaging position and orientation determination device 10. Alternatively, a touch-panel type mobile terminal, such as a smartphone or tablet, which combines the functions of the input device 4 and the display device 5 with predetermined calculation functions, may be used. Furthermore, the hardware configuration shown in Figure 3 is merely an example and is not limited thereto. While details will be explained in the modified examples, for instance, the robot 30 may have a built-in imaging position and orientation determination device 10.
[0039] Figure 4 is an explanatory diagram showing an example of an imaging position candidate set by the position and orientation calculation unit. Note that the position K1, indicated by a star in Figure 4, represents the position of the object in the second environment (for example, yard Y1 in Figure 1B). Circle C1 is a circle (a sphere in three dimensions) centered at the object's position K1. Multiple imaging position candidates, such as those indicated by position K2, are set inside this circle C1. The radius of circle C1 is set appropriately so that a clear image can be obtained when imaging the object (position K1) from an imaging position candidate (position K2). In the example in Figure 4, the multiple imaging position candidates indicated by position K2 are set to be uniformly distributed inside circle C1 (sphere).
[0040] When the robot 30 (see Figure 1B) images an object, the candidate imaging postures, which are the orientations of the second camera 31 (see Figure 1B), are set, for example, in the direction of the line segment connecting the candidate imaging position (position K2) and the object (position K1) shown in Figure 4. These candidate imaging position and imaging postures are set by the position and posture calculation unit 123 (see Figure 2) described above.
[0041] Furthermore, among the multiple positions K2 shown in Figure 4, those where interference occurs between the robot 30 (see Figure 1B) and the structures of the second environment, or where the structures of the second environment obstruct imaging, may be excluded from the imaging position candidates. The process of identifying the optimal position from multiple imaging position candidates (and corresponding imaging pose candidates) will be described later. Next, another example will be explained using Figures 5A and 5B.
[0042] Figure 5A is an explanatory diagram showing the position of the object when the operator took an image of it in the first environment. In Figure 5A, the dotted area represents a two-dimensional representation of the voxel that shows the three-dimensional shape of the structure (occupied object) in the first environment (for example, pit P1 in Figure 1A). The position K3, indicated by the white circle, indicates the position of the worker when the object (for example, the railway vehicle V1 in Figure 1A) was imaged in the first environment. The position K4, indicated by the star in Figure 5A, indicates the position of the object in the first environment. The line segment L1 indicates the optical axis direction (or the worker's line of sight direction) of the lens of the first camera 20 (see Figure 1A) when the worker at position K3 imaged the object at position K4.
[0043] The operator images the object at position K4 from a position where the object is visible without being obscured by other structures. The imaging information obtained by the operator is transmitted to the imaging position and orientation determination device 10 (see Figure 2). The position and orientation calculation unit 123 (see Figure 2) first identifies the position K4 of the object in the first environment based on the imaging information. To give a specific example, the position and orientation calculation unit 123 identifies a structure located on the optical axis of the lens of the first camera 20 (see Figure 1A) (in the operator's line of sight) as the object, based on the position and orientation information and operator environment information included in the imaging information.
[0044] In the example shown in Figure 1A, the connection between the wheel W1 and axle X1 in pit P1 lies on the optical axis of the first camera 20, and this portion is identified as the object. Furthermore, the position and orientation calculation unit 123 (see Figure 2) also identifies the positional relationship between positions K3 and K4 shown in Figure 5A, based on the position of the object and the position and orientation information of the worker. This positional relationship information is used to set the candidate imaging positions, which will be described later.
[0045] Figure 5B is an explanatory diagram showing candidate imaging positions when the robot images an object in the second environment. Furthermore, in terms of the positional relationship between the worker and the object, Figure 5B corresponds to Figure 5A. The dotted area shown in Figure 5B represents a two-dimensional representation of voxels that show the three-dimensional shape of structures (occupied objects) in the second environment (for example, yard Y1 in Figure 1B). Since this second environment is different from the first environment (see Figure 5A), the area of voxels shown by the dots is also different from that of the first environment. The position K4a, indicated by a star in Figure 5B, indicates the position of an object (for example, railway vehicle V1 in Figure 1B) in the second environment. This position K4a is the same as position K4 in Figure 5A in that it represents the position of an object, but because it is a position in the second environment, it uses a different symbol than position K4 in the first environment.
[0046] The position and orientation calculation unit 123 (see Figure 2) determines the position K4a of the object identified in the first environment if that object exists in the second environment. In the example in Figure 1B, the three-dimensional position of the connection between the wheel W1 and the axle X1 in the second environment is determined. It is assumed that the structure of the object and its arrangement in the second environment are known.
[0047] The position and orientation calculation unit 123 (see Figure 2) then applies the positional relationship between the worker's position K3 and the object's position K4 in the first environment to the second environment, thereby determining the position K3a assuming the worker is present in the second environment. In other words, the position and orientation calculation unit 123 determines the worker's position K3a in the second environment based on the positional relationship between the worker's position K3 (see Figure 5A) and the object's position K4 (see Figure 5A) in the first environment (the length of the line segment L1 and the angle relative to the object), and the object's position K4a in the second environment.
[0048] Furthermore, the position and orientation calculation unit 123 (see Figure 2) sets the nodes obtained by dividing the line segment L2 connecting the operator's position K3a and the object's position K4a during imaging as candidate imaging positions. This allows multiple candidate imaging positions to be set along the operator's line of sight (on line segment L2 in Figure 5B) during imaging. In the example in Figure 5B, line segment L2 is divided into five equal parts, and the positions K11, K12, K13, and K14 of the four nodes are set as candidate imaging positions. The position and orientation calculation unit 123 (see Figure 2) also sets the orientation when imaging the object along line segment L2 from the candidate imaging positions (positions K11, K12, K13, K14) as a candidate imaging orientation.
[0049] Furthermore, the position and orientation calculation unit 123 (see Figure 2) may exclude any of the multiple pairs of imaging position candidates and imaging orientation candidates that would cause interference between the robot 30 (see Figure 1B) and the environment, and may also exclude any that would cause the object to be obscured (obstructed from imaging by other structures) when the robot 30 images the object with the second camera 31 (see Figure 1B).
[0050] The determination of whether or not occlusion occurs is performed, for example, by comparing the feature points of an image of the object taken from a candidate imaging position with the feature points of another image of the same object taken from a different angle. Here, "feature points" refer to points on the image whose brightness value or color can be distinguished from their surroundings. If the feature points match between the aforementioned images, the position and orientation calculation unit 123 (see Figure 2) determines that no occlusion occurs when imaging the object. If the feature points do not match between the images, the position and orientation calculation unit 123 determines that occlusion occurs when imaging the object. Alternatively, the determination of whether or not occlusion occurs during imaging may be based on a deep learning model such as a Vision Language Model.
[0051] In the example in Figure 5B, positions K11 and K12 are excluded from the list of potential imaging locations because they interfere with environmental structures (voxels). Similarly, position K13 is also excluded from the list of potential imaging locations because environmental structures would obstruct (hide) the image of the target object from that position. Of the remaining positions, K14 does not interfere with the environment and does not cause any hilt, so K14 remains as a potential imaging location.
[0052] Figure 6 is an explanatory diagram illustrating another example showing candidate imaging positions when the robot images an object in the second environment. In the example shown in Figure 6, the robot 30 (see Figure 1B) is treated as an autonomous mobile object, and the position and orientation calculation unit 123 (see Figure 2) solves a path planning problem for when the autonomous mobile object moves from a predetermined starting point to an ending point, thereby generating the optimal path. The position and orientation calculation unit 123 sets a position near the worker's position K3a that does not interfere with structures in the second environment (i.e., a position that meets the interference conditions) as the starting point for the path planning problem. In the example shown in Figure 6, position K21 is set as the starting point.
[0053] Furthermore, the position and orientation calculation unit 123 (see Figure 2) sets the position K4a of the object as the endpoint of the path planning problem. Then, by solving the aforementioned path planning problem, the position and orientation calculation unit 123 generates the optimal path for the autonomous mobile object to move from the starting point (position K21) to the endpoint (position K4a) while avoiding interference with the environment. An algorithm for solving such a path planning problem is, for example, A * Algorithms and RRT * An algorithm is used.
[0054] Furthermore, the position and orientation calculation unit 123 sets each of the multiple points included in the optimal path as a candidate imaging position. For example, the nodes obtained by dividing the path length of the optimal path into equal parts may be set as candidate imaging positions. In the example in Figure 6, in addition to position K21, which is the starting point of the path planning problem, positions K22 to K25 are set as candidate imaging positions. Note that the starting and ending points of the path planning problem may be swapped so that the position K4a of the object is set as the starting point and position K21 is set as the ending point.
[0055] In this way, the position and orientation calculation unit 123 (see Figure 2) derives a path (for example, an optimal path) which is the solution to a path planning problem when one of the following is used as the starting point: position K21, which is near the operator's position when imaging the object and meets predetermined interference conditions, and position K4a, which is the object's position, and the other is the ending point. The unit sets multiple points included in this path as imaging position candidates. The position and orientation calculation unit 123 also sets the orientation when imaging the object from the imaging position candidates as an imaging orientation candidate.
[0056] The method shown in Figure 6 can generate candidate imaging positions and orientations that avoid occlusion or interference, even if occlusion or interference occurs at any point on the line segment connecting the worker and the object (for example, on line segment L2 in Figure 5B).
[0057] Figure 7 is a flowchart of the processes executed by the processing unit (see also Figure 2 as appropriate). Note that at the "START" point in Figure 7, it is assumed that the worker has already completed the task of imaging the target object with the first camera 20 (see Figure 1A) in the first environment. Also, it is assumed that the process of the robot 30 (see Figure 1B) imaging the target object with the second camera 31 (see Figure 1B) in the second environment has not yet been performed.
[0058] In step S101, the processing unit 12 acquires robot environment information and imaging information using the data acquisition unit 121 (data acquisition process). As described above, robot environment information is information that indicates the surrounding environment when the robot 30 (see Figure 1B) images the target object. Imaging information is information obtained when the operator images the target object with the first camera 20 (see Figure 1A), and includes work images, position and orientation information, and operator environment information.
[0059] In step S102, the processing unit 12 generates predetermined interference conditions using the interference condition generation unit 122 (interference condition generation process). That is, based on robot environment information, the processing unit 12 generates interference conditions between the robot 30 (see Figure 1B) and the environment when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B).
[0060] In step S103, the processing unit 12 calculates candidate imaging positions and candidate imaging positions using the position and orientation calculation unit 123 (position and orientation calculation processing). That is, the processing unit 12 calculates candidate imaging positions and candidate imaging positions for when the robot 30 (see Figure 1B) images the target object with the second camera 31 (see Figure 1B) based on the imaging information and interference conditions. For example, the processing unit 12 calculates candidate imaging positions and candidate imaging positions based on the method described using Figure 4, as well as Figures 5A, 5B, and 6.
[0061] In step S104, the processing unit 12 generates candidate images of the object when it is imaged at the candidate imaging position and imaging orientation using the evaluation unit 124. For example, the processing unit 12 generates candidate images of the object when it is imaged at the candidate imaging position and imaging orientation using the Novel View Synthesis technique described above. As described above, one candidate image is generated for each pair of candidate imaging position and imaging orientation.
[0062] In step S105, the processing unit 12 calculates the similarity between the working image and the candidate image using the evaluation unit 124. That is, the processing unit 12 calculates the similarity between the working image captured by the operator with the first camera 20 (see Figure 1A) and the candidate image generated in step S104. For calculating the similarity, methods such as histograms, hashes, and feature point matching are used as appropriate. The higher the similarity between the working image and the candidate image, the higher the evaluation of the candidate imaging position and imaging pose corresponding to that candidate image. In this way, the processing unit 2 evaluates the candidate imaging position and imaging pose corresponding to the candidate image based on a comparison between the working image and the candidate image.
[0063] Furthermore, the "evaluation process," in which the evaluation unit 124 evaluates candidate imaging positions and candidate imaging postures based on data that identifies objects included in the field of view of the first camera 20 (or objects included in the operator's field of view) when the first camera 20 (see Figure 1A) is capturing images, includes steps S104 and S105 in Figure 7. In the first embodiment, the data used is the image taken during operation.
[0064] In step S106, the processing unit 12 determines the imaging position and imaging orientation based on the similarity using the determination unit 125. That is, the processing unit 12 determines the candidate imaging position and imaging orientation corresponding to the candidate image with the highest similarity to the work image as the imaging position and imaging orientation for when the robot 30 (see Figure 1B) images the target object in the second environment. After performing the processing in step S106, the processing unit 12 terminates the series of processes (END). Although omitted in Figure 7, the processing unit 12 may also display the determination results of the imaging position and imaging orientation on the display device 50 (see Figure 3) in a predetermined manner.
[0065] Furthermore, if the image used during the process is a video, for example, one or more frames (still images) containing the target object may be extracted from the numerous frames (still images) included in the video. In this case, the processing unit 12 calculates the similarity between the frame (still image) extracted from the video and the candidate image (S105).
[0066] <Effects> According to the first embodiment, the imaging position and orientation determination device 10 determines the imaging position and orientation of the robot 30 based on imaging information and robot environment information. Therefore, when determining the imaging position and orientation of the robot 30, there is no particular need for the operator to operate the robot 30, nor is there any particular need for the operator to teach the robot 30 the imaging position and orientation, thus saving the operator time. In addition, since the robot 30 performs imaging in the second environment on behalf of the operator, the inspection time for the operator can be greatly reduced. Furthermore, based on the imaging information obtained during the operator's inspection, the processing unit 12 can identify an imaging position and orientation that does not interfere with the environment and the robot 30, and that does not obstruct the imaging.
[0067] Furthermore, for example, in the inspection of railway vehicle V1, tasks that were previously performed by workers in pit P1 (first environment: see Figure 1A) can now be performed by robot 30 in yard Y1 (second environment: see Figure 1B). This significantly reduces the man-hours and costs required for inspections in pit P1, which has been a bottleneck. Additionally, by having robot 30 take images in yard Y1, it becomes possible to eliminate the need for inspections in pit P1 (see Figure 1A). Moreover, since there is no particular need for workers to enter yard Y1 (see Figure 1B), where high-voltage power lines are present, and teach robot 30, the burden on workers can be reduced.
[0068] ≪Second Embodiment≫ The second embodiment differs from the first embodiment in that the operator's position and posture during imaging of the object are compared with candidate imaging positions and postures. Other aspects are the same as the first embodiment. Therefore, the differences from the first embodiment will be explained, and the explanation of overlapping parts will be omitted.
[0069] Figure 8 is a functional block diagram of the imaging position and orientation determination system 100A, which includes the imaging position and orientation determination device 10A according to the second embodiment. As shown in Figure 8, the processing unit 12A of the imaging position and orientation determination device 10A includes a data acquisition unit 121, an interference condition generation unit 122, a position and orientation calculation unit 123, an evaluation unit 124A, and a determination unit 125A.
[0070] The data acquisition unit 121 acquires robot environment information and imaging information. The imaging information acquired by the data acquisition unit 121 includes at least position and orientation information. Here, "position and orientation information" refers to information indicating the position (work position) and orientation (work posture) of the first camera 20 (see Figure 1A) when the operator images the target object with the first camera 20. Such position and orientation information is used as data to identify what was in the field of view of the first camera 20 when the first camera 20 was imaging.
[0071] For recording positional and orientation information, devices such as head-mounted displays that can be worn by the worker are used as appropriate. Alternatively, positional and orientation information may be generated by placing predetermined markers on the worker's head and calculating the position and orientation of the markers using a motion capture system.
[0072] Alternatively, instead of positional information, line-of-sight information indicating the worker's gaze direction at the time of imaging may be used. In this case, the line-of-sight information is used to identify what was in the worker's field of view when the first camera 20 was capturing images. Other information, such as work images and worker environment information, may also be included in the imaging information, but these can be omitted as appropriate. For example, even without work images, it is possible to identify what was in the field of view of the first camera 20 (or the worker's field of view) from the positional information at the time the worker captured the image.
[0073] As shown in Figure 8, the position and orientation information (information included in the imaging information) obtained by the data acquisition unit 121 is output to the evaluation unit 124A. The evaluation unit 124A evaluates the imaging position candidates based on a comparison between the position of the first camera 20 (see Figure 1A) at the time of imaging by the operator and the imaging position candidates, and also evaluates the imaging posture candidates based on a comparison between the orientation of the first camera 20 and the imaging posture candidates.
[0074] The determination unit 125A selects the imaging position and imaging posture from among the candidate imaging position and imaging posture candidates that receive the highest evaluation from the evaluation unit 124A as the imaging position and imaging posture. The determination result of the determination unit 125A is stored in the storage unit 11 and used when the robot 30 (see Figure 1B) images the target object in the second environment.
[0075] Figure 9 is a flowchart of the processes executed by the processing unit (see also Figure 8 as appropriate). Steps S101 to S103 in Figure 9 are the same as those described in the first embodiment (see Figure 7), so their explanation will be omitted. After calculating the candidate imaging position and orientation in step S103, the processing unit 12A proceeds to step S204. In step S204, the processing unit 12A calculates a first similarity by comparing the working position and the candidate imaging position using the evaluation unit 124A. Here, the "working position" is the three-dimensional position of the first camera 20 when the operator images the object with the first camera 20. More specifically, the relative positional relationship between the first camera 20 and the object in the first environment is applied to the second environment, and the "working position" in the second environment is determined based on the position of the object in the second environment and the aforementioned positional relationship.
[0076] Furthermore, the "first similarity" in step S204 is the similarity between the working position (the position of the first camera 20) and the candidate imaging position. In calculating the first similarity, the processing unit 12A calculates, for example, the Euclidean distance between the working position and the candidate imaging position. The processing unit 12A then increases the first similarity as the smaller the Euclidean distance between the working position and the candidate imaging position. For example, the reciprocal of the Euclidean distance may be used as the first similarity, or other formulas may be used as appropriate. In short, the closer the candidate imaging position is to the working position, the higher the first similarity is set to be.
[0077] In step S205, the processing unit 12A calculates a second similarity by comparing the working posture with the candidate imaging posture using the evaluation unit 124A. Here, "working posture" refers to the orientation of the first camera 20 when the operator images the object with the first camera 20. The "second similarity" is the similarity between the working posture (the posture of the first camera 20) and the candidate imaging posture.
[0078] In calculating the second similarity, the processing unit 12A calculates the second similarity based, for example, on the sum of the absolute values of the errors in each angle of the Euler angles between the working posture and the candidate imaging posture. Note that quaternion errors may be used instead of errors in each angle of the Euler angles. The formula for calculating the second similarity is set appropriately so that the closer the candidate imaging posture is to the working posture, the higher the second similarity.
[0079] In step S206, the processing unit 12A calculates an overall evaluation value based on the first similarity and the second similarity using the evaluation unit 124A. For example, the evaluation unit 124A calculates the overall evaluation value by taking the sum of the first similarity and the second similarity. In this way, the evaluation unit 124A evaluates the candidate imaging position and candidate imaging posture based on the first similarity and the second similarity. The higher the overall evaluation value, the more similar the imaging position and posture of the robot 30 (see Figure 1B) will be to the imaging position and posture of the operator, and as a result, an image similar to the image taken during the operation can be obtained.
[0080] In step S207, the processing unit 12A determines the imaging position and imaging orientation based on the overall evaluation value using the determination unit 125A. That is, the processing unit 12 determines the imaging position candidate and imaging orientation candidate with the highest overall evaluation value as the imaging position and imaging orientation when the robot 30 (see Figure 1B) images the target object in the second environment. After performing the processing in step S207, the processing unit 12 terminates the series of processes (END).
[0081] <Effects> According to the second embodiment, the processing unit 12A compares the work position with candidate imaging positions and compares the work posture with candidate imaging postures. In this process, there is no particular need to use work images (images obtained by the worker), so the computational load on the processing unit 12A can be reduced.
[0082] ≪Third Embodiment≫ The third embodiment differs from the first embodiment in that it calculates an evaluation value indicating the quality of the image of the object for multiple candidate images corresponding to multiple pairs of candidate imaging positions and candidate imaging poses. Other aspects are the same as the first embodiment. Therefore, the differences from the first embodiment will be explained, and the explanation of overlapping parts will be omitted.
[0083] Figure 10 is a functional block diagram of the imaging position and attitude determination system 100B, which includes the imaging position and attitude determination device 10B according to the third embodiment. As shown in Figure 10, the processing unit 12B of the imaging position and orientation determination device 10B includes a data acquisition unit 121, an interference condition generation unit 122, a position and orientation calculation unit 123, an evaluation unit 124B, and a determination unit 125B.
[0084] The position and orientation calculation unit 123 calculates multiple pairs of imaging position candidates and imaging orientation candidates. The calculation results of the position and orientation calculation unit 123 are output to the evaluation unit 124B. The evaluation unit 124B calculates an evaluation value for a candidate image based on data that identifies what was in the field of view of the first camera 20 (or what was in the field of view of the worker M1) when the first camera 20 was capturing images. The aforementioned data is information that can identify the position of the object, such as work images, position and posture information, line of sight information, and worker environment information when the worker captured images of the object with the first camera 20.
[0085] Specifically, the evaluation unit 124B first generates candidate images for each pair of multiple candidate imaging positions and imaging postures, representing the images that would be captured at those positions and postures. This generation of candidate images is performed based on techniques such as Novel View Synthesis. Then, the evaluation unit 124B calculates an evaluation value for each pair of multiple candidate imaging positions and imaging postures.
[0086] For example, when imaging is performed from predetermined candidate imaging positions and orientations, the evaluation value may be set so that the closer the center of the object captured in the candidate image is to the center of the candidate image, the higher the evaluation value. Alternatively, the evaluation value may be set so that the larger the number of pixels occupied by the object in the candidate image, the higher the evaluation value. In this case, a predetermined upper limit may be set for the number of pixels occupied by the object. This is because if the number of pixels occupied by the object is too large, it may actually become difficult to see the entire object. In short, the evaluation unit 124B calculates a predetermined evaluation value as an index indicating the quality of the image of the object in the candidate image. The formula for calculating such an evaluation value is set in advance as appropriate.
[0087] Furthermore, while evaluation values may be appropriately assigned to candidate imaging positions and orientations in which the object fits within the image's field of view (the field of view of the second camera 31), candidate imaging positions and orientations in which the object does not fit within the image's field of view may be excluded from evaluation. The determination unit 125B shown in Figure 10 determines the imaging position and imaging orientation when the robot 30 (see Figure 1B) images the target object in the second environment, based on the evaluation results of the evaluation unit 124B.
[0088] Figure 11 is a flowchart of the processes executed by the processing unit (see also Figure 10 as appropriate). Steps S101 to S103 in Figure 11 are the same as those described in the first embodiment (see Figure 7), so their explanation will be omitted. After calculating the candidate imaging position and imaging orientation in step S103, the processing unit 12B proceeds to step S304. In step S304, the processing unit 12B generates candidate images corresponding to each pair of multiple pairs of imaging position candidates and imaging pose candidates using the evaluation unit 124B. That is, for each pair in the multiple pairs of imaging position candidates and imaging pose candidates, the evaluation unit 124B generates a candidate image which is the image of the object when it is imaged at that imaging position candidate and imaging pose candidate.
[0089] In step S305, the processing unit 12B calculates an evaluation value for each candidate image using the evaluation unit 124B. That is, as described above, the processing unit 12b calculates an evaluation value for each of the multiple candidate images that indicates how well the object is captured in that candidate image.
[0090] In step S306, the processing unit 12B determines the imaging position and imaging orientation based on the evaluation value using the determination unit 125B. That is, the processing unit 12B determines the imaging position and imaging orientation for the robot 30 (see Figure 1B) when it images the target object in the second environment, based on the evaluation value of the candidate image among multiple pairs of imaging position candidates and imaging orientation candidates. After performing the processing in step S306, the processing unit 12B terminates the series of processes (END).
[0091] <Effects> According to the third embodiment, the processing unit 12B identifies the candidate image that best captures the object from among multiple pairs of candidate imaging positions and imaging orientations. As a result, when the robot 30 images the object in the second environment, it can acquire an image that captures the object well, thus enabling smooth inspection using the imaging results from the robot 30 (visual inspection by a remote worker).
[0092] ≪Fourth Embodiment≫ The fourth embodiment differs from the third embodiment in that it calculates evaluation values based on the positional relationship between the candidate imaging position and candidate imaging orientation, and the object. Other aspects (such as the configuration of the imaging position and orientation determination device 10B: see Figure 10) are the same as in the third embodiment. Therefore, the differences from the third embodiment will be explained, and the explanation of overlapping parts will be omitted.
[0093] Figure 12 is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the fourth embodiment (see also Figure 10 as appropriate). Steps S101 to S103 in Figure 12 are the same as those described in the first embodiment (see Figure 7), so their explanation will be omitted. After calculating the candidate imaging position and imaging orientation in step S103, the processing unit 12B proceeds to step S404.
[0094] In step S404, the processing unit 12B calculates an evaluation value for each pair of multiple candidate imaging positions and candidate imaging poses. That is, the evaluation unit 124B calculates an evaluation value for each pair of the multiple candidate imaging positions and candidate imaging poses, based on the positional relationship between the candidate imaging position and candidate imaging pose and the object. The evaluation value is a value that indicates the quality of the image when the object is imaged at the predetermined candidate imaging position and candidate imaging pose.
[0095] For example, to ensure that even the fine details of the object are clearly captured when imaging is performed, it is desirable to assign a higher evaluation value to imaging position candidates that are closer to the object. Alternatively, an evaluation value may be assigned to candidates where the positional relationship between the imaging position candidate and the object (the orientation of the second camera 31 in Figure 1B) is similar to the positional relationship between the first camera 20 (see Figure 1A) and the object when an operator takes images in the first environment (the orientation of the first camera 20).
[0096] In this way, the evaluation unit 124B calculates evaluation values for candidate imaging position and candidate imaging posture based on data that identifies what was in the field of view of the first camera 20 (or what was in the field of view of worker M1) when the first camera 20 was capturing images. The aforementioned data is information that can identify the position of the object, such as work images, position and posture information, line of sight information, and worker environment information when the worker captured images of the object with the first camera 20 (see Figure 1A).
[0097] In step S405, the processing unit 12B determines the imaging position and imaging orientation based on the evaluation value. That is, the processing unit 12B selects the one with the highest evaluation value from among multiple pairs of imaging position and imaging orientation candidates as the imaging position and imaging orientation for the robot 30 (see Figure 1B) when it images the target object in the second environment. After performing the process in step S405, the processing unit 12 terminates the series of processes (END).
[0098] Furthermore, a predetermined evaluation value may be assigned to candidate imaging positions and orientations such that the object fits within the field of view of the image (the field of view of the second camera 31) when imaging is performed, while candidate imaging positions and orientations such that the object does not fit within the field of view of the image may be excluded from evaluation.
[0099] Furthermore, a predetermined evaluation value may be assigned to imaging position candidates (and corresponding imaging pose candidates) where the Euclidean distance between the imaging position candidate and the object is greater than or equal to a predetermined value, while imaging position candidates where the Euclidean distance is less than the predetermined value may be excluded from evaluation. This is because if the robot 30 is too close to the object, the resulting image may become difficult to see.
[0100] Furthermore, when imaging, a predetermined evaluation value may be assigned to candidate imaging positions and orientations that do not obstruct (i.e., do not cause occlusion) of the target object by structures or other objects in the environment, and candidate imaging positions and orientations that cause occlusion may be excluded from the evaluation. This allows candidate imaging positions and orientations that capture the target object in the image to remain as evaluation targets.
[0101] <Effects> According to the fourth embodiment, the processing unit 12B identifies from among multiple pairs of imaging position candidates and imaging pose candidates that result in a good image of the object. This allows the robot 30 to acquire a good image when imaging the object in the second environment.
[0102] ≪Variations≫ Although embodiments of the imaging position and orientation determination system 100 etc. related to this disclosure have been described above, this disclosure is not limited to these descriptions and various modifications can be made.
[0103] Figure 13 is a functional block diagram of a robot 200 equipped with an imaging position and orientation determination device 10 according to a modified example. In the modified example shown in Figure 13, the robot 200 is configured to include a second camera 60, an imaging position and orientation determination device 10, and a control device 70. The configuration of the imaging position and orientation determination device 10 may be any of the first to fourth embodiments. The control device 70 of the robot 200 performs control based on the imaging position and orientation determined by the imaging position and orientation determination device 10. That is, the robot 70 uses the second camera 60 to image the target object at the aforementioned imaging position and orientation. The same effects as in each embodiment are achieved with this configuration as well.
[0104] Furthermore, in the first embodiment, a case in which work images are used as "data" to identify what was in the field of view of the first camera 20 or in the field of view of the worker when the first camera 20 was capturing images was described, and in the second embodiment, a case in which positional orientation information (or line of sight information) is used was described, but the invention is not limited to these. That is, one or more of the above-mentioned "data" may be used, including work images, positional orientation information, and line of sight information.
[0105] Furthermore, while the first embodiment described a case where the "imaging information" includes an image taken during work, position and orientation information, and worker environment information, it is not limited to this. For example, the worker's line of sight information may be used instead of position and orientation information. Also, one or more of the work image, position and orientation information, line of sight information, and worker environment information may be used as the "imaging information."
[0106] Furthermore, when generating the worker's line of sight information, the following method may be used: that is, the worker's line of sight information may be recorded in a virtual space such as a metaverse, which incorporates environmental information (including the position and structure of objects) surrounding the worker.
[0107] Furthermore, in the first embodiment, the determination unit 125 (see Figure 2) may retain, among multiple pairs of imaging position candidates and imaging pose candidates, those whose similarity (similarity between the working image and the candidate image) is equal to or greater than a predetermined value, and exclude those whose similarity is less than the predetermined value from the determination. If the similarity of all imaging position candidates and imaging pose candidates is less than the predetermined value, the evaluation unit 124 (see Figure 2) may change the conditions as appropriate, such as changing the imaging position candidates and imaging pose candidates, and then recalculate the imaging position candidates and imaging pose candidates. The same applies to the second to fourth embodiments.
[0108] Furthermore, while the second embodiment described a case in which the evaluation unit 124A (see Figure 8) evaluates both the imaging position candidate and the imaging position candidate, the invention is not limited to this. For example, the evaluation unit 124A may evaluate the imaging position candidate (and the corresponding imaging posture candidate) based on a comparison with the working position. Alternatively, the evaluation unit 124A may evaluate the imaging posture candidate (and the corresponding imaging position candidate) based on a comparison with the working posture.
[0109] Furthermore, in the first embodiment, the case in which the determination result of the imaging position and orientation determination device 10 is displayed on the display device 50 (see Figure 3) was described, but it is also possible to omit the display device 50 as appropriate. Furthermore, in the first embodiment, when the work image is a video, a representative frame included in the video is used to compare this frame with the candidate image, but the system is not limited to this. For example, the processing unit 12 (see Figure 2) may generate moment-by-moment candidate images (videos) based on the moment-by-moment candidate imaging position and imaging posture of the worker, and the evaluation unit 124 may compare the work image (video) with the candidate images (videos). In this case, the evaluation unit 124 (see Figure 2) calculates the similarity between the frame of the work image and the frame of the candidate image at the corresponding time, and further performs an overall evaluation by summing the similarities at multiple time points.
[0110] Furthermore, each embodiment can be combined as appropriate. For example, the first embodiment (see Figure 7) and the second embodiment (see Figure 9) can be combined to give a higher evaluation the greater the similarity between the work image and the candidate image, and also to give a higher evaluation the greater the similarity between the position and orientation of the first camera 20 during image acquisition of the worker and the candidate imaging position and imaging orientation. Furthermore, for example, the first embodiment (see Figure 7) and the third embodiment (see Figure 11) may be combined to give a higher evaluation the greater the similarity between the working image and the candidate image, and also to give a higher evaluation the better the image of the object in the candidate image is captured. Furthermore, for example, the third embodiment (see Figure 11) and the fourth embodiment (see Figure 12) may be combined to give a higher evaluation to candidate images where the quality of the object is better, and a higher evaluation to candidate images where the distance between the candidate imaging position and the object is appropriate. Many other combinations are also possible.
[0111] Furthermore, the processing performed by the imaging position and orientation determination system 100 and the imaging position and orientation determination device 10 (such as the imaging position and orientation determination method) may be executed as a predetermined program on a computer. The aforementioned program can be provided via a communication line, or it can be written to a recording medium such as a CD-ROM and distributed.
[0112] Furthermore, this disclosure is not limited to the embodiments and includes various modifications. For example, the embodiments are described in detail for illustrative purposes and are not necessarily limited to having all the configurations described. Also, some of the configurations of the embodiments can be added, deleted, or replaced with other configurations.
[0113] Furthermore, each of the aforementioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the aforementioned configurations, functions, etc., may be implemented in software by having the processor interpret and execute programs that realize each function. Information such as programs, tables, and files that realize each function can be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0114] Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it can be assumed that almost all components are interconnected. [Explanation of Symbols]
[0115] 10, 10A, 10B Imaging position and orientation determination device 11 Storage section 12, 12A, 12B Processing Unit 20. Camera 1 30,200 robots 31,60 Second Camera 40 Input devices 50 Display device 70 Control device 100, 100A, 100B Imaging Position and Attitude Determination System 121 Data Acquisition Unit 122 Interference Condition Generation Unit 123 Position and orientation calculation section 124, 124A, 124B Evaluation Unit 125,125A,125B Decision section K11, K12, K13, K14 positions (candidate imaging positions) K21, K22, K23, K24, K25 positions (candidate imaging positions) P1 Pit (Environment 1) S101 Step (Data Acquisition Process) S102 Step (Interference Condition Generation Process) S103 Step (Position and orientation calculation processing) S104, S105, S204, S205, S206, S304, S305, S404 Step (Evaluation Process) S106, S207, S306, S405 Step (Decision Processing) V1 Railway vehicles (objects) Y1 Yard (Second Environment)
Claims
1. A data acquisition unit that acquires imaging information when an operator images an object with a first camera, and robot environment information including 3D information of the surroundings when the robot images the object with a second camera. An interference condition generation unit generates interference conditions between the robot and the environment when the robot images the target object with the second camera, based on the robot environment information, The system includes a position and orientation calculation unit that calculates candidate imaging position and candidate imaging orientation when the robot images the target object with the second camera, based on the imaging information and interference conditions. The imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera is imaging. An evaluation unit evaluates the candidate imaging position and candidate imaging orientation based on the aforementioned data, An imaging position and orientation determination device further comprising: a determination unit that determines the imaging position and orientation when the robot images the target object with the second camera based on the evaluation results of the evaluation unit.
2. The aforementioned data is an image taken during work, which is an image obtained when the worker captures the object with the first camera. The evaluation unit generates candidate images, which are images of the object when it is photographed using the candidate imaging position and candidate imaging posture, and evaluates the candidate imaging position and candidate imaging posture corresponding to the candidate images based on a comparison between the work image and the candidate images. The imaging position and orientation determination device according to claim 1, characterized by the above.
3. The evaluation unit performs the comparison based on the similarity between the work image and the candidate image. The imaging position and orientation determination device according to claim 2, characterized by the above.
4. The aforementioned data represents the position and orientation of the first camera when the operator captured an image of the object with the first camera. The evaluation unit evaluates the candidate imaging position based on a comparison of the position of the first camera with the candidate imaging position, and evaluates the candidate imaging posture based on a comparison of the posture of the first camera with the candidate imaging posture. The imaging position and orientation determination device according to claim 1, characterized by the above.
5. The evaluation unit evaluates the candidate imaging position based on the similarity between the position of the first camera and the candidate imaging position, and evaluates the candidate imaging posture based on the similarity between the posture of the first camera and the candidate imaging posture. The imaging position and orientation determination device according to claim 4, characterized by the above.
6. The position and orientation calculation unit calculates a plurality of pairs of imaging position candidates and imaging orientation candidates, The evaluation unit generates a candidate image for each pair of the multiple pairs of candidate imaging positions and candidate imaging postures, which is an image of the object when it is imaged at the candidate imaging position and posture, and further calculates an evaluation value indicating the quality of the image of the object in the candidate image. The determination unit determines the imaging position and imaging orientation based on the evaluation value. The imaging position and orientation determination device according to claim 1, characterized by the above.
7. The position and orientation calculation unit calculates a plurality of pairs of imaging position candidates and imaging orientation candidates, The evaluation unit calculates an evaluation value for each pair of the plurality of candidate imaging positions and candidate imaging postures, based on the positional relationship between the candidate imaging position and candidate imaging posture and the object. The determination unit determines the imaging position and imaging orientation based on the evaluation value. The imaging position and orientation determination device according to claim 1, characterized by the above.
8. The imaging information includes one or more of the following: an image taken during work, which is an image taken by the worker with the first camera of the object; position and orientation information indicating the position and orientation of the first camera; line of sight information indicating the starting point position and direction of the worker's line of sight; and worker environment information, which is three-dimensional information of the worker's surroundings when the worker took an image of the object with the first camera. The imaging position and orientation determination device according to claim 1, characterized by the above.
9. The position and orientation calculation unit sets the nodes obtained by dividing the line segment connecting the operator's position and the object's position at the time of imaging the object into multiple segments as candidate imaging positions, and sets the orientation when imaging the object along the line segment from the candidate imaging positions as candidate imaging orientations. The imaging position and orientation determination device according to claim 1, characterized by the above.
10. The position and orientation calculation unit derives a path that is the solution to a path planning problem when one of the following is taken as the starting point and the other as the ending point: a position near the operator's position when the object is being imaged that fits the interference conditions, and the position of the object. It sets a plurality of points included in this path as candidate imaging positions and sets the orientation when the object is imaged from the candidate imaging positions as a candidate imaging orientation. The imaging position and orientation determination device according to claim 1, characterized by the above.
11. The position and orientation calculation unit excludes from the plurality of pairs of imaging position candidates and imaging orientation candidates those that cause interference between the robot and the environment, and also excludes those that cause the object to be obscured when the robot images the object with the second camera. The imaging position and orientation determination device according to claim 1, characterized by the above.
12. The first environment, which is the environment in which the operator images the object with the first camera, and the second environment, which is the environment in which the robot images the object with the second camera, are different. An imaging position and orientation determination device according to any one of claims 1 to 11, characterized by the above.
13. Robots and, An imaging position and orientation determination system comprising an imaging position and orientation determination device as described in claim 1, The robot is an imaging position and orientation determination system comprising a control device that performs control based on the imaging position and orientation determined by the imaging position and orientation determination device.
14. A robot comprising an imaging position and orientation determination device as described in claim 1, A robot comprising a control device that performs control based on the imaging position and imaging orientation determined by the imaging position and orientation determination device.
15. A data acquisition process that acquires imaging information when an operator images an object with a first camera, and robot environment information including 3D information of the surroundings when the robot images the object with a second camera. Interference condition generation process that generates interference conditions between the robot and the environment when the robot images the target object with the second camera, based on the robot environment information, The process includes a position and orientation calculation process that calculates candidate imaging position and candidate imaging orientation when the robot images the object with the second camera, based on the imaging information and interference conditions. The imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera is imaging. Based on the above data, an evaluation process is performed to evaluate the candidate imaging position and the candidate imaging posture. A method for determining imaging position and orientation, further comprising: a determination process for determining the imaging position and orientation when the robot images the object with the second camera based on the evaluation results of the evaluation process.
Citation Information
Patent Citations
Inspection time photographing position estimation system and program of them
JP2023179844A