Imaging position and orientation determination device, imaging position and orientation determination system, robot, and imaging position and orientation determination method

The imaging position and orientation determination device addresses the challenge of robotic imaging in diverse environments by using operator data and environmental information to determine optimal positions and orientations, enhancing efficiency and safety in inspections.

WO2026083629A1PCT designated stage Publication Date: 2026-04-23HITACHI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HITACHI LTD
Filing Date
2025-05-30
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing technologies fail to provide a method for a robot to capture optimal imaging positions and orientations in environments different from the operator's imaging environment, limiting the efficiency of inspections.

Method used

An imaging position and orientation determination device that acquires imaging information from an operator and robot environment data, generates interference conditions, calculates candidate positions and orientations, evaluates them based on similarity, and determines the optimal imaging position and orientation for the robot to image objects accurately.

Benefits of technology

Enables efficient and accurate imaging by robots in environments distinct from operators, reducing inspection time and costs by allowing robots to perform tasks previously done by humans, thereby minimizing human exposure to hazardous conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025019754_23042026_PF_FP_ABST
    Figure JP2025019754_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an imaging position and orientation determination device and the like that make it possible to appropriately image a subject by using a robot. An imaging position and orientation determination device (10) comprises: a data acquisition unit (121) that acquires robot environment information and imaging information obtained when an operator imaged a subject by means of a first camera; an interference condition generation unit (122) that generates, on the basis of the robot environment information, an interference condition between a robot and an environment when the robot images the subject by means of a second camera; a position and orientation computation unit (123) that computes, on the basis of the imaging information and the interference condition, imaging position candidates and imaging orientation candidates when the robot images the subject by means of the second camera; an evaluation unit (124) that evaluates the imaging position candidates and the imaging orientation candidates; and a determination unit (125) that determines, on the basis of an evaluation result from the evaluation unit (124), an imaging position and an imaging orientation when the robot images the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Imaging Position and Orientation Determination Device, Imaging Position and Orientation Determination System, Robot, and Imaging Position and Orientation Determination Method

[0001] The present disclosure relates to an imaging position and orientation determination device, an imaging position and orientation determination system, a robot, and an imaging position and orientation determination method.

[0002] As a technology that utilizes a robot for inspection work of an object, for example, the technology described in Patent Document 1 is known. That is, Patent Document 1 describes "estimating an optimal imaging position at which an image capable of inspecting the inspection target can be captured based on the position of the facility to be inspected and the specifications of the inspection camera."

[0003] Japanese Patent Application Laid-Open No. 2023-179844

[0004] Although Patent Document 1 describes the estimation of the optimal imaging position when a robot captures an image with a camera, for example, it does not describe a technology applicable to the case where a robot captures an object in an environment different from the operator's imaging environment based on the imaging information when the operator captures an object with a camera.

[0005] Therefore, an object of the present disclosure is to provide an imaging position and orientation determination device or the like that enables an object to be appropriately imaged using a robot.

[0006] To solve the aforementioned problems, the imaging position and orientation determination device according to this disclosure comprises: a data acquisition unit that acquires imaging information when an operator images an object with a first camera and robot environment information including three-dimensional information of the surroundings when a robot images the object with a second camera; an interference condition generation unit that generates interference conditions between the robot and the environment when the robot images the object with the second camera based on the robot environment information; and a position and orientation calculation unit that calculates candidate imaging position and candidate imaging orientation when the robot images the object with the second camera based on the imaging information and the interference conditions, wherein the imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera is imaging; an evaluation unit that evaluates the candidate imaging position and candidate imaging orientation based on the data; and a determination unit that determines the imaging position and imaging orientation when the robot images the object with the second camera based on the evaluation results of the evaluation unit.

[0007] According to this disclosure, it is possible to provide an imaging position and orientation determination device, etc., that uses a robot to appropriately image an object.

[0008] This is an explanatory diagram showing an example of the environment when an operator has previously imaged a railway vehicle, the target object, with respect to the imaging position and orientation determination device according to the first embodiment. This is an explanatory diagram showing an example of the environment when a robot images a railway vehicle, the target object, with respect to the imaging position and orientation determination device according to the first embodiment. This is a functional block diagram of the imaging position and orientation determination system including the imaging position and orientation determination device according to the first embodiment. This is a diagram showing the hardware configuration of the imaging position and orientation determination device according to the first embodiment. This is an explanatory diagram showing an example of an imaging position candidate set by the position and orientation calculation unit of the imaging position and orientation determination device according to the first embodiment. This is an explanatory diagram showing the position when an operator images an object in the first environment with respect to the imaging position and orientation determination device according to the first embodiment. This is an explanatory diagram showing an imaging position candidate when a robot images an object in the second environment with respect to the imaging position and orientation determination device according to the first embodiment. This is an explanatory diagram of another example showing an imaging position candidate when a robot images an object in the second environment with respect to the imaging position and orientation determination device according to the first embodiment. This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the first embodiment. This is a functional block diagram of the imaging position and orientation determination system including the imaging position and orientation determination device according to the second embodiment. This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the second embodiment. This is a functional block diagram of an imaging position and orientation determination system including an imaging position and orientation determination device according to the third embodiment. This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the third embodiment. This is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the fourth embodiment. This is a functional block diagram of a robot equipped with an imaging position and orientation determination device according to a modified example.

[0009] <<First Embodiment>> Below, we will first briefly describe the imaging environment of the object to be inspected, and then describe in detail the imaging position and attitude determination device 10 (see Figure 2) according to the first embodiment. Below, as an example, we will describe the case in which the object to be inspected is a railway vehicle V1 (see Figures 1A and 1B), but it is not limited to this. For example, the equipment and facilities to be inspected may be ships and aircraft, as well as power transmission equipment, power distribution equipment, substation equipment, communication equipment, air conditioning equipment, refrigeration equipment, medical equipment, gas equipment, water supply equipment, railway tracks and power lines, roads, and plants. As mentioned above, plants include power plants and manufacturing plants, as well as chemical plants and water treatment plants.

[0010] The images obtained from photographing the object being inspected may be still images or videos. In the following explanation, still images and videos will be collectively referred to as "images."

[0011] <Example of an object> Figure 1A is an explanatory diagram showing an example of the environment when an operator M1 has previously imaged a railway vehicle V1, the object of the imaging position and orientation determination device according to the first embodiment. In the example in Figure 1A, the state in which the railway vehicle V1 has been brought into the inspection facility is shown. Also, in Figure 1A, the body of the railway vehicle V1 is not shown, and only the wheels W1 and axles X1 are shown (the same applies to Figure 1B). The wheels W1 of the railway vehicle V1 are placed on a pair of rails R1. The rails R1 are supported by support members E1 that extend along the rails R1, and the support members E1 are further supported by a plurality of support bases F1.

[0012] In the inspection facility, a pit P1, which serves as a workspace for worker M1, is provided directly beneath the railway vehicle V1. Worker M1 enters pit P1 and inspects the railway vehicle V1 by looking up at it from below. Worker M1 visually checks, for example, whether there is any looseness or damage to the fastenings and welds of the bogies that support the body of the railway vehicle V1.

[0013] Furthermore, there is a limit to the number of inspection facilities, as well as a limit to the number of railway vehicles that can be brought into a single inspection facility. Consequently, until now, inspecting railway vehicles has required a great deal of time and effort, and as a result, inspection work in the pits has become a bottleneck. Therefore, in the first embodiment, the worker M1 first takes an image of the railway vehicle V1, which is the object to be inspected, with the first camera 20. The image obtained by the first camera 20 is used as a kind of model when the robot 30 (see Figure 1B) takes an image of the railway vehicle V1 with the second camera 31 (see Figure 1B). Hereafter, the environment in which worker M1 takes an image of the object (for example, the railway vehicle V1) with the first camera 20 will be referred to as the "first environment".

[0014] In the example shown in Figure 1A, the first camera 20 is mounted on the helmet of worker M1, but the imaging method can be changed as appropriate. For example, the first camera 20 may be mounted on a detachable device such as a head-mounted display. Alternatively, worker M1 may hold the first camera 20 in their hand while imaging. The images obtained by the first camera 20 may be still images such as RGB color images, or they may be moving images. When worker M1 images the inspection points of the railway vehicle V1 with the first camera 20, the position (imaging position) and orientation (imaging posture) of the first camera 20 at the time of imaging are also recorded as appropriate.

[0015] Figure 1B is an explanatory diagram illustrating an example of the environment in which robot 30 images a railway vehicle V1. In Figure 1B, robot 30 images the railway vehicle V1 with a second camera 31 in a yard Y1, which is a marshalling yard for the railway vehicle V1. In Figure 1B, a multi-jointed snake-like robot is used as robot 30, but the type of robot 30 can be changed as appropriate. For example, robots equipped with means of movement such as walking legs, wheels, or crawlers (tracks), or unmanned aerial vehicles such as drones may be used as robot 30.

[0016] As shown in Figure 1B, the robot 30 is equipped with a second camera 31. Based on the imaging position and orientation determined by the imaging position and orientation determination device 10 (see Figure 2), which will be described later, the robot 30 takes images of the railway vehicle V1. The environment in which the robot 30 takes images of the target object (for example, the railway vehicle V1) with the second camera 31 is referred to as the "second environment".

[0017] Thus, the first environment in which worker M1 (see Figure 1A) images the railway vehicle V1 (for example, pit P1 in Figure 1A) and the second environment in which robot 30 images the railway vehicle V1 (for example, yard Y1 in Figure 1B) are different. However, the first embodiment is not limited to cases where the first and second environments are different; it can also be applied when the first and second environments are similar. Furthermore, the railway vehicle V1 in the first environment (see Figure 1A) and the railway vehicle V1 in the second environment (see Figure 1B) do not necessarily need to be completely identical; they can be of the same type (or have a similar structure).

[0018] The general procedure for inspecting the railway vehicle V1 is as follows: First, in pit P1 (first environment) as shown in Figure 1A, worker M1 uses the first camera 20 to image the inspection points of the railway vehicle V1. Then, using the images obtained by worker M1 as a guide, the imaging position and orientation determination device 10 (see Figure 2), described later, determines the imaging position and orientation for robot 30 to perform imaging in yard Y1 (second environment) as shown in Figure 1B. Next, based on the imaging position and orientation determined by the imaging position and orientation determination device 10, robot 30 uses the second camera 31 to image the inspection points of railway vehicle V1 in yard Y1 (second environment). The images captured by robot 30 are then displayed on a remote terminal (tablet, etc.), and the railway vehicle V1 is inspected by viewing the screen of that terminal. Note that the operator may view the imaging results of the second camera 31 in real time, or they may view them later.

[0019] <Configuration of the Image Capture Position and Attitude Determination System> Figure 2 is a functional block diagram of the image capture position and attitude determination system 100, including the image capture position and attitude determination device 10. The image capture position and attitude determination system 100 shown in Figure 2 is a system for determining the imaging position and orientation when a robot 30 (see Figure 1B) images an object (for example, the railway vehicle V1 in Figure 1B) with a second camera 31 (see Figure 1B), and is composed of the robot 30 and the image capture position and attitude determination device 10. For example, a computer can be used as such an image capture position and attitude determination device 10. The functions of the image capture position and attitude determination device 10 may be distributed among multiple computers such as cloud servers or edge servers.

[0020] As shown in Figure 2, the imaging position and orientation determination device 10 comprises a storage unit 11 and a processing unit 12. The storage unit 11 has predetermined programs and data stored in it beforehand. In addition, the storage unit 11 stores robot environment information and imaging information, which will be described later, as well as the calculation results of the processing unit 12 as appropriate. The processing unit 12 executes predetermined processes based on the programs and data stored in the storage unit 11.

[0021] As shown in Figure 2, the processing unit 12 includes a data acquisition unit 121, an interference condition generation unit 122, a position and orientation calculation unit 123, an evaluation unit 124, and a determination unit 125. The data acquisition unit 121 acquires robot environment information and imaging information. Here, "robot environment information" refers to data including three-dimensional information of the surroundings (around the robot 30) when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B). Such robot environment information is created in advance based on predetermined three-dimensional surveying and stored in the memory of another computer (not shown).

[0022] In the example shown in Figure 1B, the three-dimensional information of the second environment, yard Y1, is used as robot environment information. The three-dimensional shape and arrangement of the railway vehicle V1, which is the object to be inspected, are also included in the robot environment information. As appropriate, three-dimensional maps, mesh models, and point cloud data are used as such robot environment information. The aforementioned three-dimensional maps are created based on the well-known SLAM (Simultaneous Localization and Mapping) technology. Mesh models are generated by dividing the shape of the object into simple shapes such as triangles and quadrilaterals. Point cloud data is a collection of points acquired by a 3D laser scanner.

[0023] The "imaging information" shown in Figure 2 includes, for example, work images, position and orientation information, and worker environment information. Here, "work images" refer to images obtained when worker M1 (see Figure 1A) images an object with the first camera 20 (see Figure 1A). These work images are used as data to identify objects included in the field of view of the first camera 20 (or objects included in the field of view of worker M1) when the first camera 20 is imaging. It is assumed that the work images include the object to be inspected. As data to identify objects included in the field of view of the first camera 20, for example, it may include a specific image and line-of-sight information that can be represented on the image coordinates, from the viewpoint that it is possible to identify which point the first camera 20 is looking at in the image. Also, as data to identify objects included in the field of view of worker M1, for example, it may include the position and orientation of worker M1 and line-of-sight information that can be represented in three-dimensional space, from the viewpoint that it is possible to identify the three-dimensional coordinates of the object being gazed at. Incidentally, by using AR goggles or similar devices, it is possible to make the field of view of worker M1 and the field of view of the first camera 20 approximately the same. Therefore, based on the data obtained from the AR goggles or similar devices, it is possible to identify objects that are within the field of view of the first camera 20 (i.e., objects that are within the field of view of worker M1).

[0024] Furthermore, the "position and orientation information" included in the imaging information refers to information indicating the position and orientation of the first camera 20 (see Figure 1A) when operator M1 (see Figure 1A) images an object with the first camera 20 (see Figure 1A). Here, the "position" of the first camera 20 is the three-dimensional position of the first camera 20, and is expressed in predetermined three-dimensional coordinates. The "orientation" of the first camera 20 refers to the orientation of the first camera 20. For example, the orientation of the first camera 20 can be expressed using the well-known Euler angle, which represents the optical axis direction of the lens of the first camera 20. Such position and orientation information is generated, for example, when operator M1 (see Figure 1A) images an object with the first camera 20 (see Figure 1A). For example, a gyro sensor (not shown) may be provided on the first camera 20, and the position and orientation of the first camera 20 may be measured using this gyro sensor.

[0025] Alternatively, instead of positional information, the line of sight information of worker M1 (see Figure 1A) may be used. Here, "line of sight information" refers to information indicating the starting point position and direction of worker M1's line of sight. Furthermore, "worker environment information" included in the imaging information refers to the three-dimensional information of the area around worker M1 (see Figure 1A) when worker M1 (see Figure 1A) images the target object with the first camera 20 (see Figure 1A). Such worker environment information is created in advance based on predetermined three-dimensional surveying.

[0026] The aforementioned imaging information, including images taken during operation, position and orientation information, line of sight information, and worker environment information, is provided as appropriate from, for example, a predetermined computer (not shown) that stores this data. In Figure 2, the information acquired by the data acquisition unit 121 (robot environment information and imaging information) is shown to be output directly to the interference condition generation unit 122, the position and orientation calculation unit 123, and the evaluation unit 124. However, this information may be stored in the storage unit 11 first and then read out from the storage unit 11 as appropriate.

[0027] The interference condition generation unit 122 shown in Figure 2 generates interference conditions between the robot 30 (see Figure 1B) and the environment when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B), based on robot environment information. Here, "interference conditions" refer to information indicating whether surrounding structures interfere with the robot 30 (whether they come into contact with or collide with the robot 30), as well as whether other structures obstruct the robot 30 when it images an object with the second camera 31. For example, the interference conditions may be expressed as a predetermined numerical range or mathematical formula that specifies the range of positions the robot 30 can take.

[0028] Furthermore, when interference conditions are generated, mesh data or polygon data of the environment may be used as appropriate. For example, the interference condition generation unit 122 determines whether or not there is interference between the robot 30 and the environment based on geometric algorithms such as Bounding Volume Hierarchy or Separating Axis Theorem. Alternatively, the processing unit 12 may divide the three-dimensional space into three-dimensional boxes (voxels) based on the voxel grid method and register whether or not there is an object in each voxel. In this case, the presence or absence of interference between the robot 30 and the environment is determined based on whether or not the position of the robot 30 overlaps with a voxel.

[0029] The position and orientation calculation unit 123 shown in Figure 2 calculates candidate imaging positions and imaging orientations for when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B) based on the imaging information and interference conditions described above. Here, "candidate imaging position" refers to a candidate position for the second camera 31 when the robot 30 images an object with the second camera 31. Also, "candidate imaging orientation" refers to a candidate orientation (in the optical axis direction of the lens) of the second camera 31 when the robot 30 images an object with the second camera 31. Note that the number of candidate imaging positions and imaging orientations may be one pair or multiple pairs.

[0030] The evaluation unit 124 shown in Figure 2 evaluates the candidate imaging position and candidate imaging posture to a predetermined degree. Specifically, the evaluation unit 124 first generates candidate images assuming that the robot 30 (see Figure 1B) performs imaging at predetermined candidate imaging position and posture. The generation of such candidate images is performed based on Novel View Synthesis technology, which generates an image from a new viewpoint from images from multiple viewpoints. More specifically, techniques such as Gaussian Splatting and Nerf (Neural Radiance Fields) are used as appropriate. It is assumed that one candidate image is generated for each pair of candidate imaging position and candidate imaging posture.

[0031] The evaluation unit 124 then calculates the similarity between the work image included in the imaging information (the image obtained by imaging worker M1) and the candidate image obtained when imaging at a predetermined candidate imaging position and imaging posture. The more similar the candidate image is to the work image, the higher the similarity score.

[0032] The evaluation unit 124 calculates the similarity between, for example, an image taken during work by worker M1 in the first environment, pit P1 (see Figure 1A), of the railway vehicle V1, and a candidate image assuming that robot 30 takes an image of the railway vehicle V1 in the second environment, yard Y1 (see Figure 1B). If there are multiple pairs of candidate imaging positions and imaging orientations, the similarity between each candidate image corresponding to each pair of candidate imaging positions and imaging orientations and the image taken during work is calculated.

[0033] The determination unit 125 shown in Figure 2 determines the imaging position and imaging orientation when the robot 30 (see Figure 1B) images the target object with the second camera 31 (see Figure 1B) based on the similarity evaluation result from the evaluation unit 124. In other words, the determination unit 125 determines the candidate imaging position and imaging orientation corresponding to the candidate image with the highest similarity to the work image as the imaging position and imaging orientation when the robot 30 images the target object. Here, "imaging position" refers to the three-dimensional position of the second camera 31 when the robot 30 images the target object with the second camera 31. Also, "imaging orientation" refers to the orientation of the second camera 31 when the robot 30 images the target object with the second camera 31.

[0034] In this way, the imaging position and orientation determination device 10 determines the imaging position and orientation for the robot 30 (see Figure 1B) when it images the target object, based on the robot environment information and imaging information acquired by the data acquisition unit 121. Information indicating the determination result of the imaging position and orientation determination device 10 is transmitted to the robot 30 by wire or wireless. The robot 30 is equipped with a control device (not shown) that performs control based on the imaging position and orientation determined by the imaging position and orientation determination device 10.

[0035] Figure 3 shows the hardware configuration of the imaging position and orientation determination device 10. As shown in Figure 3, the imaging position and orientation determination device 10 has a hardware configuration that includes a processor 10a, RAM 10b (Random Access Memory), ROM 10c (Read Only Memory), HDD 10d (Hard Disk Drive), a communication interface 10e, and an input / output interface 10f, which are predeterminedly connected via an internal bus 10g.

[0036] The processor 10a is hardware that functions as the processing unit 12 (see Figure 2). The RAM 10b, ROM 10c, and HDD 10d are hardware that functions as the storage unit 11 (see Figure 2). The processor 10a reads a predetermined program stored in the ROM 10c or HDD 10d and loads it into the RAM 10b, thereby executing a predetermined process.

[0037] The communication interface 10e is used for communication with external devices and networks. For example, the communication interface 10e sends and receives predetermined information with the robot 30. The input / output interface 10f is an interface used for data input from the input device 40 and data output to the display device 50. The input device 40 is, for example, a keyboard or mouse, and is used when the user inputs data.

[0038] The display device 50 is, for example, a display that shows the calculation results of the imaging position and orientation determination device 10. Alternatively, a touch-panel type mobile terminal, such as a smartphone or tablet, which combines the functions of the input device 4 and the display device 5 with predetermined calculation functions, may be used. Furthermore, the hardware configuration shown in Figure 3 is merely an example and is not limited thereto. While further details will be explained in the modified examples, for instance, the robot 30 may have a built-in imaging position and orientation determination device 10.

[0039] Figure 4 is an explanatory diagram showing an example of imaging position candidates set by the position and orientation calculation unit. The position K1 indicated by the star in Figure 4 represents the position of the object in the second environment (for example, yard Y1 in Figure 1B). Circle C1 is a circle (a sphere in three dimensions) centered at the object's position K1. Multiple imaging position candidates, such as those indicated by position K2, are set inside this circle C1. The radius of circle C1 is set appropriately so that a clear image is obtained when imaging the object (position K1) from the imaging position candidate (position K2). In the example in Figure 4, the multiple imaging position candidates indicated by position K2 are set to be uniformly distributed inside circle C1 (the sphere).

[0040] When the robot 30 (see Figure 1B) images an object, the candidate imaging postures, which are the orientations of the second camera 31 (see Figure 1B), are set, for example, in the direction of the line segment connecting the candidate imaging position (position K2) and the object (position K1) shown in Figure 4. These candidate imaging position and imaging postures are set by the position and posture calculation unit 123 (see Figure 2) described above.

[0041] In addition, among the plurality of positions K2 shown in FIG. 4, those where interference occurs between the robot 30 (see FIG. 1B) and the structures in the second environment, or where the structures in the second environment obstruct imaging, may be excluded from the imaging position candidates. The process of specifying the optimal one from the plurality of imaging position candidates (and the corresponding imaging pose candidates) will be described later. Next, another example will be described using FIGS. 5A and 5B.

[0042] FIG. 5A is an explanatory diagram showing the position when an operator images an object in the first environment. Note that the dotted area in FIG. 5A is a two-dimensional representation of voxels showing the three-dimensional shape of the structure (occupied object) in the first environment (for example, the pit P1 in FIG. 1A). Also, the position K3 indicated by the white circle shows the position of the operator when imaging an object (for example, the railway vehicle V1 in FIG. 1A) in the first environment. The position K4 indicated by the star in FIG. 5A shows the position of the object in the first environment. Also, the line segment L1 indicates the optical axis direction of the lens of the first camera 20 (see FIG. 1A) (or the line of sight direction of the operator) when the operator at the position K3 images the object at the position K4.

[0043] The operator images the object from a position where the object at the position K4 can be seen without being hidden by other structures. The imaging information obtained by the operator's imaging is transmitted to the imaging position and pose determination device 10 (see FIG. 2). First, the position and pose calculation unit 123 (see FIG. 2) identifies the position K4 of the object in the first environment based on the imaging information. For a specific example, the position and pose calculation unit 123 identifies the structure existing on the optical axis of the lens of the first camera 20 (see FIG. 1A) (at the tip of the operator's line of sight) as the object based on the position and pose information and the operator environment information included in the imaging information.

[0044] In the example of FIG. 1A, since the connection part between the wheel W1 and the axle X1 existing in the pit P1 is on the optical axis of the first camera 20, this part is identified as the object. Further, the position and pose calculation unit 123 (see FIG. 2) also identifies the positional relationship between the positions K3 and K4 shown in FIG. 5A based on the position of the object and the position and pose information of the operator. The information indicating this positional relationship is used for setting the imaging position candidates described later.

[0045] FIG. 5B is an explanatory diagram showing candidate imaging positions when the robot images an object in the second environment. Note that in the positional relationship between the operator and the object, FIG. 5B corresponds to FIG. 5A. The dot area shown in FIG. 5B is a two-dimensional representation of voxels showing the three-dimensional shape of the structure (occupied object) in the second environment (for example, the yard Y1 in FIG. 1B). Since this second environment is different from the above-described first environment (see FIG. 5A), the voxel area indicated by dots is also different from that of the first environment. The position K4a indicated by a star in FIG. 5B indicates the position of the object (for example, the railway vehicle V1 in FIG. 1B) in the second environment. This position K4a is common with the position K4 in FIG. 5A in terms of the point representing the position of the object, but is given a different reference from the position K4 in the first environment because it is a position in the second environment.

[0046] The position and orientation calculation unit 123 (see FIG. 2) specifies the position K4a when the object identified in the first environment exists in the second environment for the object. In the example of FIG. 1B, the three-dimensional position of the connection part between the wheel W1 and the axle X1 in the second environment is specified. Note that the structure of the object and the arrangement of the object in the second environment are assumed to be known.

[0047] Then, the position and orientation calculation unit 123 (see FIG. 2) applies the positional relationship between the position K3 of the operator in the first environment and the position K4 of the object as it is to the second environment, thereby specifying the position K3a when it is assumed that the operator exists in the second environment. That is, the position and orientation calculation unit 123 determines the positional relationship (the length of the line segment L1 and the angle with respect to the object) between the position K3 (see FIG. 5A) of the operator in the first environment and the position K4 (see FIG. 5A) of the object, and the position K4a of the object in the second environment, and specifies the position K3a of the operator in the second environment.

[0048] Furthermore, the position and orientation calculation unit 123 (see Figure 2) sets the nodes obtained by dividing the line segment L2 connecting the operator's position K3a and the position K4a of the object during imaging as candidate imaging positions. This allows multiple candidate imaging positions to be set along the operator's line of sight (on the line segment L2 in Figure 5B) during imaging. In the example in Figure 5B, the line segment L2 is divided into five equal parts, and the positions K11, K12, K13, and K14 of the four nodes are set as candidate imaging positions. The position and orientation calculation unit 123 (see Figure 2) also sets the orientation when imaging the object along the line segment L2 from the candidate imaging positions (positions K11, K12, K13, K14) as a candidate imaging orientation.

[0049] Furthermore, the position and orientation calculation unit 123 (see Figure 2) may exclude any of the multiple pairs of imaging position candidates and imaging orientation candidates that would cause interference between the robot 30 (see Figure 1B) and the environment, and may also exclude any that would cause the object to be obscured (obstructed from imaging by other structures) when the robot 30 images the object with the second camera 31 (see Figure 1B).

[0050] The determination of whether or not occlusion occurs is performed, for example, by comparing the feature points of an image of the object taken from a candidate imaging position with the feature points of another image of the same object taken from a different angle. Here, "feature points" refer to points on the image whose brightness value or color can be distinguished from their surroundings. If the feature points match between the images, the position and orientation calculation unit 123 (see Figure 2) determines that no occlusion occurs when imaging the object. If the feature points do not match between the images, the position and orientation calculation unit 123 determines that occlusion occurs when imaging the object. Alternatively, the determination of whether or not occlusion occurs during imaging may be based on a deep learning model such as a Vision Language Model.

[0051] In the example in Figure 5B, positions K11 and K12 are excluded from the list of imaging location candidates because they interfere with environmental structures (voxels). Also, when imaging the object from position K13, environmental structures obstruct the view (causing occlusion), so position K13 is also excluded from the list of imaging location candidates. At the remaining position K14, there is no interference with the environment and no occlusion occurs, so position K14 remains as a candidate for imaging location.

[0052] Figure 6 is an explanatory diagram illustrating another example showing candidate imaging positions when a robot images an object in the second environment. In the example in Figure 6, the robot 30 (see Figure 1B) is treated as an autonomous mobile object, and the position and orientation calculation unit 123 (see Figure 2) solves a path planning problem for when the autonomous mobile object moves from a predetermined starting point to an ending point, thereby generating the optimal path. The position and orientation calculation unit 123 sets a position near the operator's position K3a that does not interfere with structures in the second environment (i.e., a position that meets the interference conditions) as the starting point for the path planning problem. In the example in Figure 6, position K21 is set as the starting point.

[0053] Furthermore, the position and orientation calculation unit 123 (see Figure 2) sets the position K4a of the object as the endpoint of the path planning problem. Then, by solving the aforementioned path planning problem, the position and orientation calculation unit 123 generates the optimal path for the autonomous mobile object to move from the starting point (position K21) to the endpoint (position K4a) while avoiding interference with the environment. An algorithm for solving such a path planning problem is, for example, A * Algorithms and RRT * An algorithm is used.

[0054] Furthermore, the position and orientation calculation unit 123 sets each of the multiple points included in the optimal path as a candidate imaging position. For example, the nodes obtained by dividing the path length of the optimal path into equal parts may be set as candidate imaging positions. In the example in Figure 6, in addition to position K21, which is the starting point of the path planning problem, positions K22 to K25 are set as candidate imaging positions. Note that the starting and ending points of the path planning problem may be swapped so that the position K4a of the object is set as the starting point and position K21 is set as the ending point.

[0055] In this way, the position and orientation calculation unit 123 (see Figure 2) derives a path (for example, an optimal path) which is the solution to a path planning problem when one of the following is taken as the starting point: a position K21 near the operator's position when imaging the object that meets predetermined interference conditions, and the position K4a of the object, with the other as the ending point, and sets a plurality of points included in this path as imaging position candidates. The position and orientation calculation unit 123 also sets the orientation when imaging the object from the imaging position candidates as an imaging orientation candidate.

[0056] The method shown in Figure 6 can generate candidate imaging positions and orientations that avoid occlusion or interference, even if occlusion or interference occurs at any point on the line segment connecting the worker and the object (for example, on line segment L2 in Figure 5B).

[0057] Figure 7 is a flowchart of the process executed by the processing unit (see also Figure 2 as appropriate). Note that at the "START" stage in Figure 7, it is assumed that the worker has already completed the task of imaging the target object with the first camera 20 (see Figure 1A) in the first environment. Also, it is assumed that the robot 30 (see Figure 1B) has not yet performed the process of imaging the target object with the second camera 31 (see Figure 1B) in the second environment.

[0058] In step S101, the processing unit 12 acquires robot environment information and imaging information using the data acquisition unit 121 (data acquisition process). As described above, robot environment information is information that indicates the surrounding environment when the robot 30 (see Figure 1B) images the target object. Imaging information is information when the operator images the target object with the first camera 20 (see Figure 1A), and includes work images, position and orientation information, and operator environment information.

[0059] In step S102, the processing unit 12 generates predetermined interference conditions using the interference condition generation unit 122 (interference condition generation process). That is, based on robot environment information, the processing unit 12 generates interference conditions between the robot 30 (see Figure 1B) and the environment when the robot 30 (see Figure 1B) images an object with the second camera 31 (see Figure 1B).

[0060] In step S103, the processing unit 12 calculates candidate imaging positions and candidate imaging positions using the position and orientation calculation unit 123 (position and orientation calculation processing). That is, the processing unit 12 calculates candidate imaging positions and candidate imaging positions for when the robot 30 (see Figure 1B) images the target object with the second camera 31 (see Figure 1B) based on the imaging information and interference conditions. For example, the processing unit 12 calculates candidate imaging positions and candidate imaging positions based on the method described using Figure 4, as well as Figures 5A, 5B, and 6.

[0061] In step S104, the processing unit 12 generates candidate images of the object when it is imaged at the candidate imaging position and imaging orientation using the evaluation unit 124. For example, the processing unit 12 generates candidate images of the object when it is imaged at the candidate imaging position and imaging orientation using the Novel View Synthesis technique described above. As described above, one candidate image is generated for each pair of candidate imaging position and imaging orientation.

[0062] In step S105, the processing unit 12 calculates the similarity between the working image and the candidate image using the evaluation unit 124. That is, the processing unit 12 calculates the similarity between the working image captured by the operator with the first camera 20 (see Figure 1A) and the candidate image generated in step S104. For calculating the similarity, methods such as histograms, hashes, and feature point matching are used as appropriate. The higher the similarity between the working image and the candidate image, the higher the evaluation of the candidate imaging position and imaging orientation corresponding to that candidate image. In this way, the processing unit 2 evaluates the candidate imaging position and imaging orientation corresponding to the candidate image based on a comparison between the working image and the candidate image.

[0063] Furthermore, the "evaluation process," in which the evaluation unit 124 evaluates candidate imaging positions and candidate imaging postures based on data that identifies objects included in the field of view of the first camera 20 (or objects included in the operator's field of view) when the first camera 20 (see Figure 1A) is capturing images, includes steps S104 and S105 in Figure 7. In the first embodiment, the data used is the image taken during operation.

[0064] In step S106, the processing unit 12 determines the imaging position and imaging orientation based on the similarity using the determination unit 125. That is, the processing unit 12 determines the candidate imaging position and imaging orientation corresponding to the candidate image with the highest similarity to the work image as the imaging position and imaging orientation for when the robot 30 (see Figure 1B) images the target object in the second environment. After performing the processing in step S106, the processing unit 12 terminates the series of processes (END). Although not shown in Figure 7, the processing unit 12 may also display the determination results of the imaging position and imaging orientation on the display device 50 (see Figure 3) in a predetermined manner.

[0065] Furthermore, if the image used during the process is a video, for example, one or more frames (still images) containing the target object may be extracted from the numerous frames (still images) included in the video. In this case, the processing unit 12 calculates the similarity between the frame (still image) extracted from the video and the candidate image (S105).

[0066] <Effects> According to the first embodiment, the imaging position and orientation determination device 10 determines the imaging position and orientation of the robot 30 based on the imaging information and the robot environment information. Therefore, when determining the imaging position and orientation of the robot 30, there is no particular need for the operator to operate the robot 30, nor is there any particular need for the operator to teach the robot 30 the imaging position and orientation, thus saving the operator time. In addition, since the robot 30 performs imaging in the second environment on behalf of the operator, the inspection time for the operator can be greatly reduced. Furthermore, based on the imaging information obtained during the operator's inspection, the processing unit 12 can identify an imaging position and orientation that does not interfere with the environment and the robot 30, and that does not obstruct the image during imaging.

[0067] Furthermore, for example, in the inspection of a railway vehicle V1, the robot 30 can perform the work that was previously done by an operator in pit P1 (first environment: see Figure 1A) in yard Y1 (second environment: see Figure 1B). This significantly reduces the man-hours and costs required for inspections in pit P1, which has been a bottleneck. In addition, by having the robot 30 take images in yard Y1, it becomes possible to eliminate the need for inspections in pit P1 (see Figure 1A). Moreover, since there is no particular need for an operator to enter yard Y1 (see Figure 1B), where high-voltage power lines are present, and teach the robot 30, the burden on the operator can be reduced.

[0068] ≪Second Embodiment≫ The second embodiment differs from the first embodiment in that the operator's position and posture during imaging of the object are compared with candidate imaging positions and postures. Other aspects are the same as the first embodiment. Therefore, the differences from the first embodiment will be explained, and the explanation of overlapping parts will be omitted.

[0069] Figure 8 is a functional block diagram of the imaging position and orientation determination system 100A, which includes the imaging position and orientation determination device 10A according to the second embodiment. As shown in Figure 8, the processing unit 12A of the imaging position and orientation determination device 10A includes a data acquisition unit 121, an interference condition generation unit 122, a position and orientation calculation unit 123, an evaluation unit 124A, and a determination unit 125A.

[0070] The data acquisition unit 121 acquires robot environment information and imaging information. The imaging information acquired by the data acquisition unit 121 includes at least position and orientation information. Here, "position and orientation information" refers to information indicating the position (work position) and orientation (work posture) of the first camera 20 (see Figure 1A) when the operator images the target object with the first camera 20. Such position and orientation information is used as data to identify what was in the field of view of the first camera 20 when the first camera 20 was imaging.

[0071] For recording positional and orientation information, devices such as head-mounted displays that can be worn by the worker are used as appropriate. Alternatively, positional and orientation information may be generated by placing predetermined markers on the worker's head and calculating the position and orientation of the markers using a motion capture system.

[0072] Alternatively, instead of positional information, line-of-sight information indicating the worker's gaze direction at the time of imaging may be used. In this case, the line-of-sight information is used as data to identify what was in the worker's field of view when the first camera 20 was capturing images. In addition, work images and worker environment information may be included in the imaging information, but these can be omitted as appropriate. For example, even without work images, what was in the field of view of the first camera 20 (or the worker's field of view) can be identified from the positional information at the time the worker captured the image.

[0073] As shown in Figure 8, the position and orientation information (information included in the imaging information) obtained by the data acquisition unit 121 is output to the evaluation unit 124A. The evaluation unit 124A evaluates the imaging position candidates based on a comparison between the position of the first camera 20 (see Figure 1A) at the time of imaging by the operator and the imaging position candidates, and also evaluates the imaging orientation candidates based on a comparison between the orientation of the first camera 20 and the imaging orientation candidates.

[0074] The determination unit 125A selects the imaging position and imaging posture from among the candidate imaging position and imaging posture candidates that receive the highest evaluation from the evaluation unit 124A as the imaging position and imaging posture. The determination result of the determination unit 125A is stored in the storage unit 11 and used when the robot 30 (see Figure 1B) images the target object in the second environment.

[0075] Figure 9 is a flowchart of the processing performed by the processing unit (see also Figure 8 as appropriate). Steps S101 to S103 in Figure 9 are the same as those described in the first embodiment (see Figure 7), so their explanation is omitted. After calculating the candidate imaging position and orientation and the candidate imaging orientation in step S103, the processing unit 12A proceeds to step S204. In step S204, the processing unit 12A calculates the first similarity by comparing the working position and the candidate imaging position using the evaluation unit 124A. Here, "working position" is the three-dimensional position of the first camera 20 when the operator images the object with the first camera 20. More specifically, the relative positional relationship between the first camera 20 and the object in the first environment is applied to the second environment, and the "working position" in the second environment is determined based on the position of the object in the second environment and the aforementioned positional relationship.

[0076] Furthermore, the "first similarity" in step S204 is the similarity between the working position (the position of the first camera 20) and the candidate imaging position. In calculating the first similarity, the processing unit 12A calculates, for example, the Euclidean distance between the working position and the candidate imaging position. The processing unit 12A then increases the first similarity as the smaller the Euclidean distance between the working position and the candidate imaging position. For example, the reciprocal of the Euclidean distance may be used as the first similarity, or other formulas may be used as appropriate. In short, the setting is such that the closer the candidate imaging position is to the working position, the higher the first similarity.

[0077] In step S205, the processing unit 12A calculates a second similarity score by comparing the working posture with the candidate imaging posture using the evaluation unit 124A. Here, "working posture" refers to the orientation of the first camera 20 when the operator images the object with the first camera 20. The "second similarity score" is the similarity between the working posture (the posture of the first camera 20) and the candidate imaging posture.

[0078] In calculating the second similarity, the processing unit 12A calculates the second similarity based, for example, on the sum of the absolute values ​​of the errors in each angle of the Euler angles between the working posture and the candidate imaging posture. Note that quaternion errors may be used instead of errors in each angle of the Euler angles. The formula for calculating the second similarity is set appropriately so that the closer the candidate imaging posture is to the working posture, the higher the second similarity.

[0079] In step S206, the processing unit 12A calculates an overall evaluation value based on the first similarity and the second similarity using the evaluation unit 124A. For example, the evaluation unit 124A calculates the overall evaluation value by taking the sum of the first similarity and the second similarity. In this way, the evaluation unit 124A evaluates the candidate imaging position and candidate imaging posture based on the first similarity and the second similarity. The higher the overall evaluation value, the more similar the imaging position and posture of the robot 30 (see Figure 1B) will be to the imaging position and posture of the operator, and as a result, an image similar to the image taken during the operation can be obtained.

[0080] In step S207, the processing unit 12A determines the imaging position and imaging orientation based on the overall evaluation value using the determination unit 125A. That is, the processing unit 12 determines the imaging position candidate and imaging orientation candidate with the highest overall evaluation value as the imaging position and imaging orientation when the robot 30 (see Figure 1B) images the target object in the second environment. After performing the process in step S207, the processing unit 12 terminates the series of processes (END).

[0081] <Effects> According to the second embodiment, the processing unit 12A compares the work position with candidate imaging positions, and also compares the work posture with candidate imaging postures. In this process, there is no particular need to use work images (images obtained by the worker taking images), so the computational load on the processing unit 12A can be reduced.

[0082] ≪Third Embodiment≫ The third embodiment differs from the first embodiment in that it calculates an evaluation value indicating the quality of the image of the object for multiple candidate images corresponding to multiple pairs of candidate imaging positions and candidate imaging poses. Other aspects are the same as the first embodiment. Therefore, the differences from the first embodiment will be explained, and the explanation of overlapping parts will be omitted.

[0083] Figure 10 is a functional block diagram of an imaging position and attitude determination system 100B including an imaging position and attitude determination device 10B according to the third embodiment. As shown in Figure 10, the processing unit 12B of the imaging position and attitude determination device 10B includes a data acquisition unit 121, an interference condition generation unit 122, a position and attitude calculation unit 123, an evaluation unit 124B, and a determination unit 125B.

[0084] The position and orientation calculation unit 123 calculates multiple pairs of candidate imaging positions and candidate imaging orientations. The calculation results of the position and orientation calculation unit 123 are output to the evaluation unit 124B. The evaluation unit 124B calculates an evaluation value for the candidate images based on data that identifies what was in the field of view of the first camera 20 (or what was in the field of view of the worker M1) when the first camera 20 was capturing images. The aforementioned data is information that can identify the position of the object, such as work images, position and orientation information, line of sight information, and worker environment information when the worker captured images of the object with the first camera 20.

[0085] Specifically, the evaluation unit 124B first generates candidate images for each pair of multiple candidate imaging positions and imaging postures, representing the images that would appear if captured at those positions and postures. This generation of candidate images is performed based on techniques such as Novel View Synthesis. Then, the evaluation unit 124B calculates an evaluation value for each pair of multiple candidate imaging positions and imaging postures.

[0086] For example, when imaging is performed from predetermined candidate imaging positions and orientations, the evaluation value may be set so that the closer the center of the object captured in the candidate image is to the center of the candidate image, the higher the evaluation value. Alternatively, the evaluation value may be set so that the more pixels occupied by the object in the candidate image, the higher the evaluation value. In this case, a predetermined upper limit may be set for the number of pixels occupied by the object. This is because if the number of pixels occupied by the object is too large, it may actually become difficult to see the entire object. In short, the evaluation unit 124B calculates a predetermined evaluation value as an index indicating the quality of the image of the object in the candidate image. The formula for calculating such an evaluation value is set in advance as appropriate.

[0087] Furthermore, evaluation values ​​may be appropriately assigned to candidate imaging positions and imaging postures such that the object fits within the field of view of the image (the field of view of the second camera 31), while candidate imaging positions and imaging postures such that the object does not fit within the field of view of the image may be excluded from evaluation. The determination unit 125B shown in Figure 10 determines the imaging position and imaging posture when the robot 30 (see Figure 1B) images the object in the second environment, based on the evaluation results of the evaluation unit 124B.

[0088] Figure 11 is a flowchart of the processing performed by the processing unit (see also Figure 10 as appropriate). Steps S101 to S103 in Figure 11 are the same as those described in the first embodiment (see Figure 7), so their explanation is omitted. After calculating the candidate imaging position and candidate imaging posture in step S103, the processing unit 12B proceeds to step S304. In step S304, the processing unit 12B generates candidate images corresponding to each pair of multiple pairs of candidate imaging position and candidate imaging posture using the evaluation unit 124B. That is, for each pair in the multiple pairs of candidate imaging position and candidate imaging posture, the evaluation unit 124B generates a candidate image which is the image obtained when the object is imaged at that candidate imaging position and candidate imaging posture.

[0089] In step S305, the processing unit 12B calculates an evaluation value for each candidate image using the evaluation unit 124B. That is, as described above, the processing unit 12b calculates an evaluation value for each of the multiple candidate images that indicates how well the object is captured in that candidate image.

[0090] In step S306, the processing unit 12B determines the imaging position and imaging orientation based on the evaluation value using the determination unit 125B. That is, the processing unit 12B determines the imaging position and imaging orientation for the robot 30 (see Figure 1B) when it images the target object in the second environment, based on the evaluation value of the candidate image among the multiple pairs of candidate imaging position and imaging orientation. After performing the process in step S306, the processing unit 12B terminates the series of processes (END).

[0091] <Effects> According to the third embodiment, the processing unit 12B identifies the candidate image that best captures the object from among multiple pairs of candidate imaging positions and imaging orientations. As a result, when the robot 30 images the object in the second environment, it can acquire an image that captures the object well, thus enabling smooth inspection using the imaging results of the robot 30 (visual inspection by a remote worker).

[0092] ≪Fourth Embodiment≫ The fourth embodiment differs from the third embodiment in that it calculates evaluation values ​​based on the positional relationship between the candidate imaging position and candidate imaging orientation, and the object. Other aspects (such as the configuration of the imaging position and orientation determination device 10B: see Figure 10) are the same as in the third embodiment. Therefore, the parts that differ from the third embodiment will be explained, and the explanation of overlapping parts will be omitted.

[0093] Figure 12 is a flowchart of the processing performed by the processing unit of the imaging position and orientation determination device according to the fourth embodiment (see also Figure 10 as appropriate). Steps S101 to S103 in Figure 12 are the same as those described in the first embodiment (see Figure 7), so their explanation is omitted. After calculating the candidate imaging position and candidate imaging orientation in step S103, the processing unit 12B proceeds to step S404.

[0094] In step S404, the processing unit 12B calculates an evaluation value for each pair of multiple pairs of imaging position candidates and imaging pose candidates. That is, the evaluation unit 124B calculates an evaluation value for each pair in the multiple pairs of imaging position candidates and imaging pose candidates based on the positional relationship between the said imaging position candidate and the said imaging pose candidate and the object. The aforementioned evaluation value is a value that indicates the quality of the image when the object is imaged at the predetermined imaging position candidate and imaging pose candidate.

[0095] For example, to ensure that even the fine details of the object are clearly captured when imaging is performed, it is desirable to assign a higher evaluation value to imaging position candidates that are closer to the object. Alternatively, an evaluation value may be assigned to candidates where the positional relationship between the imaging position candidate and the object (the orientation of the second camera 31 in Figure 1B) is similar to the positional relationship between the first camera 20 (see Figure 1A) and the object when an operator takes images in the first environment (the orientation of the first camera 20).

[0096] In this way, the evaluation unit 124B calculates evaluation values ​​for candidate imaging position and candidate imaging posture based on data that identifies what was in the field of view of the first camera 20 (or what was in the field of view of the worker M1) when the first camera 20 was capturing images. The aforementioned data is information that can identify the position of the object, such as work images, position and posture information, line of sight information, and worker environment information when the worker captured images of the object with the first camera 20 (see Figure 1A).

[0097] In step S405, the processing unit 12B determines the imaging position and imaging orientation based on the evaluation value. That is, the processing unit 12B selects the one with the highest evaluation value from among multiple pairs of imaging position candidate and imaging orientation candidate as the imaging position and imaging orientation when the robot 30 (see Figure 1B) images the target object in the second environment. After performing the process in step S405, the processing unit 12 terminates the series of processes (END).

[0098] Furthermore, a predetermined evaluation value may be assigned to candidate imaging positions and orientations such that the object fits within the field of view of the image (the field of view of the second camera 31) when imaging is performed, while candidate imaging positions and orientations such that the object does not fit within the field of view of the image may be excluded from evaluation.

[0099] Furthermore, a predetermined evaluation value may be assigned to imaging position candidates (and corresponding imaging pose candidates) where the Euclidean distance between the imaging position candidate and the object is greater than or equal to a predetermined value, while imaging position candidates where the Euclidean distance is less than the predetermined value may be excluded from evaluation. This is because if the robot 30 is too close to the object, the resulting image may become difficult to see.

[0100] Furthermore, when imaging, a predetermined evaluation value may be assigned to candidate imaging positions and orientations that do not obstruct (i.e., do not cause occlusion) of the target object by structures or other objects in the environment, and candidate imaging positions and orientations that cause occlusion may be excluded from the evaluation. This allows candidate imaging positions and orientations that capture the target object in the image to remain as evaluation targets.

[0101] <Effects> According to the fourth embodiment, the processing unit 12B identifies from among multiple pairs of imaging position candidates and imaging pose candidates that result in a good image of the object. As a result, when the robot 30 images the object in the second environment, it can acquire an image with good quality.

[0102] <Modifications> The imaging position and orientation determination system 100 etc. related to this disclosure have been described above in terms of embodiments, but this disclosure is not limited to these descriptions and various modifications can be made.

[0103] Figure 13 is a functional block diagram of a robot 200 equipped with an imaging position and orientation determination device 10 according to a modified example. In the modified example shown in Figure 13, the robot 200 is configured to include a second camera 60, an imaging position and orientation determination device 10, and a control device 70. The configuration of the imaging position and orientation determination device 10 may be any of the first to fourth embodiments. The control device 70 of the robot 200 performs control based on the imaging position and orientation determined by the imaging position and orientation determination device 10. That is, the robot 70 uses the second camera 60 to image the target object at the aforementioned imaging position and orientation. The same effects as in each embodiment are achieved even with this configuration.

[0104] Furthermore, in the first embodiment, a case in which an image taken during work is used as "data" to identify what was in the field of view of the first camera 20 or in the field of view of the worker when the first camera 20 was capturing images was described, and in the second embodiment, a case in which positional orientation information (or line of sight information) is used was described, but the invention is not limited to these. That is, one or more of the above-mentioned "data" may be used, including the image taken during work, positional orientation information, and line of sight information.

[0105] Furthermore, while the first embodiment described a case where the "imaging information" includes an image taken during work, position and orientation information, and worker environment information, it is not limited to this. For example, the worker's line of sight information may be used instead of position and orientation information. Also, one or more of the work image, position and orientation information, line of sight information, and worker environment information may be used as the "imaging information."

[0106] Furthermore, when generating the worker's line of sight information, the following method may be used: that is, the worker's line of sight information may be recorded in a virtual space such as a metaverse, which incorporates environmental information surrounding the worker (including information such as the position and structure of objects).

[0107] Furthermore, in the first embodiment, the determination unit 125 (see Figure 2) may retain, among multiple pairs of imaging position candidates and imaging pose candidates, those whose similarity (similarity between the working image and the candidate image) is equal to or greater than a predetermined value, and exclude those whose similarity is less than the predetermined value from the determination. If the similarity of all imaging position candidates and imaging pose candidates is less than the predetermined value, the evaluation unit 124 (see Figure 2) may change the conditions as appropriate, such as changing the imaging position candidates and imaging pose candidates, and then recalculate the imaging position candidates and imaging pose candidates. The same applies to the second to fourth embodiments.

[0108] Furthermore, in the second embodiment, a case was described in which the evaluation unit 124A (see Figure 8) evaluates both the imaging position candidate and the imaging position candidate, but the invention is not limited to this. For example, the evaluation unit 124A may evaluate the imaging position candidate (and the corresponding imaging posture candidate) based on a comparison with the working position. Alternatively, the evaluation unit 124A may evaluate the imaging posture candidate (and the corresponding imaging position candidate) based on a comparison with the working posture.

[0109] Furthermore, in the first embodiment, the case in which the determination result of the imaging position and orientation determination device 10 is displayed on the display device 50 (see Figure 3) was described, but the display device 50 can be omitted as appropriate. Also, in the first embodiment, when the work image is a video, the case in which a representative frame included in the video is used to compare this frame with the candidate image was described, but it is not limited to this. For example, the processing unit 12 (see Figure 2) may generate moment-by-moment candidate images (videos) based on the moment-by-moment candidate imaging position and imaging orientation of the worker, and the evaluation unit 124 may compare the work image (video) and the candidate images (videos). In this case, the evaluation unit 124 (see Figure 2) calculates the similarity between the frame of the work image and the frame of the candidate image at the corresponding time, and further performs an overall evaluation by summing the similarities at multiple time points.

[0110] Furthermore, each embodiment can be combined as appropriate. For example, the first embodiment (see Figure 7) and the second embodiment (see Figure 9) can be combined to give a higher evaluation the greater the similarity between the work image and the candidate image, and also to give a higher evaluation the greater the similarity between the position and posture of the first camera 20 during imaging of the worker and the candidate imaging position and posture. Alternatively, for example, the first embodiment (see Figure 7) and the third embodiment (see Figure 11) can be combined to give a higher evaluation the greater the similarity between the work image and the candidate image, and also to give a higher evaluation the better the image of the object in the candidate image. Alternatively, for example, the third embodiment (see Figure 11) and the fourth embodiment (see Figure 12) can be combined to give a higher evaluation the better the image of the object in the candidate image, and also to give a higher evaluation the more appropriate the distance between the candidate imaging position and the object. Many other combinations are also possible.

[0111] Furthermore, the processing performed by the imaging position and orientation determination system 100 and the imaging position and orientation determination device 10 (such as the imaging position and orientation determination method) may be executed as a predetermined program on a computer. The aforementioned program can be provided via a communication line, or it can be written to a recording medium such as a CD-ROM and distributed.

[0112] Furthermore, this disclosure is not limited to the embodiments and includes various modifications. For example, the embodiments are described in detail for illustrative purposes and are not necessarily limited to having all the configurations described. Also, some of the configurations of the embodiments can be added, deleted, or replaced with other configurations.

[0113] Furthermore, each of the aforementioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the aforementioned configurations, functions, etc., may be implemented in software by having the processor interpret and execute programs that realize each function. Information such as programs, tables, and files that realize each function can be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0114] Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it can be assumed that almost all components are interconnected.

[0115] 10, 10A, 10B Imaging position and orientation determination device 11 Storage unit 12, 12A, 12B Processing unit 20 First camera 30, 200 Robot 31, 60 Second camera 40 Input device 50 Display device 70 Control device 100, 100A, 100B Imaging position and orientation determination system 121 Data acquisition unit 122 Interference condition generation unit 123 Position and orientation calculation unit 124, 124A, 124B Evaluation unit 125, 125A, 125B Determination unit K11, K12, K13, K14 Position (candidate imaging position) K21, K22, K23, K24, K25 Position (candidate imaging position) P1 Pit (first environment) S101 Step (data acquisition processing) S102 Step (interference condition generation processing) S103 Step (Position and orientation calculation processing) S104, S105, S204, S205, S206, S304, S305, S404 Step (Evaluation processing) S106, S207, S306, S405 Step (Decision processing) V1 Railway vehicle (object) Y1 Yard (second environment)

Claims

1. An imaging position and posture determination device comprising: a data acquisition unit that acquires imaging information when an operator images an object with a first camera and robot environment information including three-dimensional information of the surroundings when a robot images the object with a second camera; an interference condition generation unit that generates interference conditions between the robot and the environment when the robot images the object with the second camera based on the robot environment information; a position and posture calculation unit that calculates candidate imaging position and candidate imaging posture when the robot images the object with the second camera based on the imaging information and the interference conditions, wherein the imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera takes an image; an evaluation unit that evaluates the candidate imaging position and candidate imaging posture based on the data; and a determination unit that determines the imaging position and imaging posture when the robot images the object with the second camera based on the evaluation results of the evaluation unit.

2. The imaging position and posture determination device according to claim 1, wherein the data is an image taken during work, which is an image obtained when the worker takes an image of the object with the first camera, and the evaluation unit generates candidate images, which are images taken when the object is photographed at the candidate imaging position and the candidate imaging posture, and evaluates the candidate imaging position and the candidate imaging posture corresponding to the candidate images based on a comparison between the image taken during work and the candidate images.

3. The imaging position and orientation determination device according to claim 2, characterized in that the evaluation unit performs the comparison based on the similarity between the working image and the candidate image.

4. The imaging position and orientation determination device according to claim 1, wherein the data is the position and orientation of the first camera when the operator images the object with the first camera, and the evaluation unit evaluates the imaging position candidate based on a comparison of the position of the first camera with the imaging position candidate, and evaluates the imaging orientation candidate based on a comparison of the orientation of the first camera with the imaging orientation candidate.

5. The imaging position and posture determination device according to claim 4, characterized in that the evaluation unit evaluates the imaging position candidate based on the similarity between the position of the first camera and the imaging position candidate, and evaluates the imaging posture candidate based on the similarity between the posture of the first camera and the imaging posture candidate.

6. The imaging position and orientation determination device according to claim 1, characterized in that the position and orientation calculation unit calculates a plurality of pairs of imaging position candidates and imaging orientation candidates; the evaluation unit generates a candidate image for each pair of the plurality of imaging position candidates and imaging orientation candidates, which is an image of the object when it is captured at the said imaging position candidate and imaging orientation candidate; and further calculates an evaluation value indicating the quality of the image of the object in the candidate image; and the determination unit determines the imaging position and imaging orientation based on the evaluation value.

7. The imaging position and orientation determination device according to claim 1, characterized in that the position and orientation calculation unit calculates a plurality of pairs of imaging position candidates and imaging orientation candidates, the evaluation unit calculates an evaluation value for each pair in the plurality of pairs of imaging position candidates and imaging orientation candidates based on the positional relationship between the imaging position candidate and the imaging orientation candidate and the object, and the determination unit determines the imaging position and the imaging orientation based on the evaluation value.

8. The imaging position and orientation determination device according to claim 1, characterized in that the imaging information is one or more of the following: an image taken during work, which is an image taken by the worker with the first camera of the object; position and orientation information indicating the position and orientation of the first camera; line of sight information indicating the starting point position and direction of the worker's line of sight; and worker environment information, which is three-dimensional information of the worker's surroundings when the worker took an image of the object with the first camera.

9. The imaging position and orientation determination device according to claim 1, characterized in that the position and orientation calculation unit sets the nodes obtained by dividing the line segment connecting the operator's position and the position of the object at the time of imaging the object into multiple segments as candidate imaging positions, and sets the orientation when imaging the object along the line segment from the candidate imaging positions as a candidate imaging orientation.

10. The imaging position and orientation determination device according to claim 1, characterized in that the position and orientation calculation unit derives a path which is the solution to a path planning problem when one of the following is taken as the starting point and the other as the ending point: a position near the operator's position when imaging the object and which conforms to the interference conditions, and the position of the object; sets a plurality of points included in the path as imaging position candidates; and sets the orientation when imaging the object from the imaging position candidates as the imaging orientation candidates.

11. The imaging position and orientation determination device according to claim 1, characterized in that the position and orientation calculation unit excludes from the plurality of pairs of imaging position candidates and imaging orientation candidates those that cause interference between the robot and the environment, and also excludes those that cause the object to be obscured when the robot images the object with the second camera.

12. The imaging position and orientation determination device according to any one of claims 1 to 11, characterized in that the first environment, which is the environment in which the operator images the object with the first camera, and the second environment, which is the environment in which the robot images the object with the second camera, are different.

13. An imaging position and attitude determination system comprising a robot and an imaging position and attitude determination device according to claim 1, wherein the robot comprises a control device that performs control based on the imaging position and the imaging attitude determined by the imaging position and attitude determination device.

14. A robot comprising an imaging position and orientation determination device as described in claim 1, wherein the robot comprises a control device that performs control based on the imaging position and orientation determined by the imaging position and orientation determination device.

15. A method for determining imaging position and posture, comprising: a data acquisition process that acquires imaging information when an operator images an object with a first camera and robot environment information including three-dimensional information of the surroundings when a robot images the object with a second camera; an interference condition generation process that generates interference conditions between the robot and the environment when the robot images the object with the second camera based on the robot environment information; a position and posture calculation process that calculates candidate imaging position and candidate imaging posture when the robot images the object with the second camera based on the imaging information and the interference conditions, wherein the imaging information includes data that identifies an object included in the field of view of the first camera or an object included in the field of view of the operator when the first camera takes an image; an evaluation process that evaluates the candidate imaging position and candidate imaging posture based on the data; and a determination process that determines the imaging position and imaging posture when the robot images the object with the second camera based on the evaluation result of the evaluation process.

Citation Information

Patent Citations

  • Appearance inspection apparatus

    JP1996313225A

  • Image processor and image processing method

    JP2013117795A

  • Imaging device and autonomous traveling device including the same, and imaging method

    JP2019161444A

  • Device and method for detection and localization of vehicles

    US20200090366A1

  • Image-capturing plan generation device, image-capturing plan generation method, and program

    WO2018070354A1