Image analysis device, image analysis method, and computer program
The image analysis device and method efficiently identify the position of a target region on a three-dimensional model by generating a three-dimensional model from two-dimensional images and calculating the intersection point of a straight line, addressing the challenge of manual search and reducing computational complexity.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
It is difficult to search for and specify the position of a target area on a three-dimensional model from a three-dimensional model.
An image analysis device and method that generates a three-dimensional model from multiple two-dimensional images, places the target two-dimensional image in a virtual space based on imaging coordinates and posture information, and calculates the position where a straight line through the imaging coordinate intersects the three-dimensional model to identify the target region.
Enables the simple and accurate identification of the target region on the three-dimensional model by processing a single two-dimensional image, reducing manual work and computational burden, and ensuring high-speed and precise target detection.
Smart Images

Figure JP2024034909_02042026_PF_FP_ABST
Abstract
Description
Image analysis device, image analysis method, and computer program
[0001] The present disclosure relates to an image analysis device, an image analysis method, and a computer program.
[0002] The diagnostic system described in Patent Document 1 has a function of automatically determining and detecting deteriorated portions by performing diagnostic processing by a computer based on continuous images obtained by aerial photography of an object in the past and at present. The object is a building, infrastructure facilities, etc. A predetermined area on the surface of the object is a diagnostic target area. Based on a user's input operation, the diagnostic target area of the object is set as a basic setting.
[0003] The diagnostic system described in Patent Document 1 has an SFM processing unit. The SFM processing unit performs SFM processing on a plurality of input images to restore a three-dimensional structure and outputs a three-dimensional model.
[0004] Japanese Unexamined Patent Application Publication No. 2019-070631
[0005] Generally, there are cases where a diagnostic target area is searched from a three-dimensional model of an object and the diagnostic target area is diagnosed.
[0006] However, it is often difficult to search for a target area such as a diagnostic target area from a three-dimensional model and specify the position of the target area on the three-dimensional model.
[0007] Therefore, the present disclosure has been made in view of the above problems, and an object thereof is to provide an image analysis device, an image analysis method, and a computer program capable of specifying the position of a target area on a three-dimensional model by simple processing.
[0008] According to this disclosure, an image analysis device is provided, comprising: a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, which are indicated by three-dimensional coordinates in a virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging; and a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate strikes the three-dimensional model, wherein the first target three-dimensional coordinate indicates the position in the virtual space of a target image included in the target two-dimensional image.
[0009] Furthermore, the present disclosure provides an image analysis method that includes the steps of: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquiring imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; imaging pose information indicating the pose of the camera that performed the imaging; placing the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculating a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
[0010] Furthermore, the present disclosure provides a computer program that causes a computer to perform the following steps: generate a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquire imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; and imaging pose information indicating the pose of the camera that performed the imaging; place the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculate a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculate a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
[0011] This disclosure provides an image analysis device, an image analysis method, and a computer program that can identify the position of a target region on a three-dimensional model through a simple process.
[0012] This is a block diagram showing an example configuration of an image analysis system according to one embodiment of the present invention. This is a block diagram showing an example configuration of a server according to the same embodiment. This is a schematic perspective view showing a camera, a two-dimensional image, and a subject in the real space according to the same embodiment. This is a schematic perspective view showing an imaging position and a three-dimensional model of a subject in the virtual space according to the same embodiment. This is a schematic perspective view showing an imaging position, a target two-dimensional image, and a three-dimensional model in the virtual space according to the same embodiment. (a) is a schematic side view showing an imaging position represented in the virtual space according to the same embodiment, (b) is a schematic side view showing a target two-dimensional image placed in the virtual space, (c) is a schematic side view showing a target image and bounding box included in the target two-dimensional image placed in the virtual space, and a target region included in the three-dimensional model, and (d) is a schematic side view showing a state in the virtual space where a straight line passing through the imaging position and the target image hits the three-dimensional model. This is a front view showing a target two-dimensional image placed in the virtual space according to the same embodiment. This is a flowchart showing an image analysis method according to the same embodiment. This is a flowchart showing a three-dimensional model generation process according to the same embodiment. This is a flowchart showing an image analysis method according to a modified example of the same embodiment.
[0013] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.
[0014] Figure 1 is a diagram showing an example configuration of an image analysis system SYS according to one embodiment of the present invention. As shown in Figure 1, the image analysis system SYS includes a server 1. Server 1 corresponds to an example of the "image analysis device" in this disclosure.
[0015] Server 1 generates a three-dimensional model of the subject, which is placed in a virtual space, based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions. Server 1 calculates the position in the virtual space where a straight line passing through the imaging position and the target image contained in one of the two-dimensional images intersects the three-dimensional model (i.e., the position of the target region on the three-dimensional model). As a result, the position of the target region on the three-dimensional model is identified. Thus, in this embodiment, instead of performing a process to search for the target region across the entire area of the three-dimensional model, the position of the target region on the three-dimensional model is calculated from the camera position and one of the two-dimensional images. Therefore, the position of the target region on the three-dimensional model can be identified by a simple process. Details of this point will be described later.
[0016] Furthermore, the image analysis system SYS comprises at least one terminal 2, at least one mobile device 3, or at least one imaging device 5.
[0017] Server 1, terminal 2, mobile device 3, and imaging device 5 are connected to a network NW. The network NW includes, for example, the Internet, a private network, a public telephone network, a LAN (Local Area Network), and a short-range wireless network.
[0018] Terminal 2 is, for example, a personal computer (e.g., a laptop computer, a desktop computer, or a tablet).
[0019] The mobile device 3 is, for example, an unmanned mobile device or a manned mobile device. An unmanned mobile device is, for example, an unmanned aerial vehicle such as a drone, an unmanned ground vehicle, an unmanned submersible, or an unmanned surface vessel. An unmanned ground vehicle is, for example, an unmanned ground vehicle or an unmanned ground vehicle that mimics a living organism (for example, a snake-shaped unmanned ground vehicle). A manned mobile device is, for example, an aircraft, an automobile, a ship, or a submarine. The mobile device 3 includes a camera 4.
[0020] The imaging device 5 includes a camera 6. The imaging device 5 is, for example, a mobile device such as a smartphone. The imaging device 5 may also be, for example, the camera 6 itself.
[0021] Cameras 4 and 6 capture images of the subject and generate video data showing a video containing images of the subject. The video is a collection of consecutive two-dimensional images. Cameras 4 and 6 may also generate multiple still image data, each showing multiple still images containing images of the subject. The still images are two-dimensional images.
[0022] Hereafter, video data will be referred to as "video data 511" (Figure 2). Similarly, multiple still image data will be referred to as "still image data set 512" (Figure 2).
[0023] In the following, unless it is necessary to distinguish between cameras 4 and 6, cameras 4 and 6 will be collectively referred to as "camera CM".
[0024] Furthermore, the subject is not particularly limited as long as it can be captured by the camera CM, and the size, shape, pattern, and color of the subject are not particularly limited. For example, the subject is one or more movable or immovable property. The subject is, for example, one or more objects. Typically, the subject is one or more stationary objects. Stationary objects are, for example, man-made or natural objects. Man-made objects are, for example, structures, machinery, electronic equipment, or copyrighted works. Structures are, for example, buildings or infrastructure facilities. Buildings are, for example, office buildings or houses. Infrastructure facilities are facilities for developing social infrastructure. For example, infrastructure facilities are roads, bridges, road traffic facilities, power generation facilities, power transmission facilities, water treatment facilities, or gas distribution facilities. Machinery are, for example, automobiles, work vehicles, trains, aircraft, ships, submarines, or robots. Natural objects are, for example, trees, forests, the ground surface, cliffs, coastlines, or rivers.
[0025] Furthermore, the subject includes one or more targets whose positions on the three-dimensional model are to be identified. The targets are not particularly limited and can be arbitrarily set, such as all of the area included in the subject, part of the area, all of the objects included in the subject, or part of the objects.
[0026] Terminal 2 acquires video data 511 or still image data set 512 generated by camera CM. For example, terminal 2 receives video data 511 or still image data set 512 transmitted from camera CM via network NW. For example, terminal 2 acquires video data 511 or still image data set 512 from removable media. Terminal 2 transmits video data 511 or still image data set 512 to server 1 via network NW. Alternatively, the mobile device 3 and imaging device 5 may transmit video data 511 or still image data set 512 to server 1 via network NW.
[0027] Figure 2 is a block diagram showing an example configuration of Server 1 in Figure 1. As shown in Figure 2, Server 1 includes an arithmetic unit 10, a communication unit 40, and a storage unit 50. Server 1 may also include an input unit 20 and a display unit 30.
[0028] The input unit 20 is an input device for inputting various types of information to the calculation unit 10. For example, the input unit 20 may be a keyboard and pointing device, or a touch panel.
[0029] The display unit 30 displays various information. The display unit 30 is, for example, a liquid crystal display or an organic electroluminescent display.
[0030] The communication unit 40 is connected to a network NW. The communication unit 40 communicates with external devices connected to the network NW. The external devices are, for example, a terminal 2, a mobile device 3, and an imaging device 5. The communication unit 40 is a communication device that performs communication according to a predetermined communication protocol, and includes, for example, a network interface controller. The predetermined communication protocol is, for example, a protocol compliant with Ethernet® and an Internet Protocol Suite.
[0031] The communication unit 40 receives video data 511 or still image data set 512 generated by imaging each subject from the terminal 2, mobile device 3, and imaging device 5 via the network NW.
[0032] The storage unit 50 includes a storage device and stores data and computer programs. The storage unit 50 includes a main storage device such as a semiconductor memory and an auxiliary storage device such as a semiconductor memory and a hard disk drive. The storage unit 50 may also include a removable medium such as an optical disc. The storage unit 50 may be, for example, a non-temporary computer-readable storage medium.
[0033] The memory unit 50 stores the first database 51, the second database 52, and the trained model 53.
[0034] The first database 51 includes video data 511 and still image data set 512 received by the communication unit 40. In the first database 51, the video data 511 and still image data set 512 are associated with attribute information (hereinafter, "attribute information AT"). Attribute information AT includes, for example, user information, imaging conditions, and subject information. User information includes, for example, identification information of the user of terminal 2, mobile device 3, or imaging device 5. Imaging conditions include, for example, the date and time of imaging and calibration information of camera CM. Calibration information includes, for example, information regarding lens distortion, focal length, and principal point position of camera CM. Information regarding lens distortion includes, for example, information regarding distortion aberration. Imaging conditions may also include information on the imaging location. The imaging location is indicated, for example, by the position coordinates of camera CM obtained by GPS (Global Positioning System) or GNSS (Global Navigation Satellite System). Subject information includes, for example, identification information of the subject. The subject information may include the coordinates of the ground control point (GCP).
[0035] The second database 52 includes multiple datasets 520. Each of the datasets 520 includes multiple two-dimensional images 150, a three-dimensional model 301 (point cloud data), multiple imaging position information 521, and multiple imaging orientation information 522. In the second database 52, the two-dimensional images 150, the three-dimensional model 301, the imaging position information 521, and the imaging orientation information 522 are related to each other and are also related to attribute information AT. Details of these will be described later.
[0036] The trained model 53 is a computer program. The trained model 53 includes, for example, a neural network. Details of the trained model 53 will be described later.
[0037] The arithmetic unit 10 performs various calculations. The arithmetic unit 10 includes processors such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).
[0038] Specifically, the calculation unit 10 includes a model generation unit 11, a distortion correction unit 12, an image placement unit 13, a target detection unit 14, a target coordinate calculation unit 15, and a display control unit 16. For example, the calculation unit 10 functions as the model generation unit 11, the distortion correction unit 12, the image placement unit 13, the target detection unit 14, the target coordinate calculation unit 15, and the display control unit 16 by executing a computer program stored in the storage unit 50.
[0039] Next, the processing performed by the arithmetic unit 10 will be explained with reference to Figures 2 to 6. Figure 3 is a schematic perspective view showing the camera CM, a two-dimensional image 150, and the subject 300 in real space RS. As shown in Figure 3, the subject 300 includes the target 310. The two-dimensional image 150 includes a subject image 200 that shows the subject 300. The subject image 200 includes a target image 210 that shows the target 310, depending on the imaging position 100. In this embodiment, the subject 300 is imaged while moving one camera CM, and multiple two-dimensional images 150 are generated. Figure 4 is a schematic perspective view showing the imaging position 101 and a three-dimensional model 301 of the subject 300 in virtual space VS. In Figure 4, the imaging position 101 is indicated by a black circle. For ease of understanding, region A of the three-dimensional model 301 is shown enlarged.
[0040] As shown in Figures 2 to 4, the model generation unit 11 of the calculation unit 10 generates a three-dimensional model 301 of the subject 300 based on a plurality of two-dimensional images 150 generated by imaging the subject 300 from a plurality of different imaging positions 100. The three-dimensional model 301 includes a target region 311 that indicates the target 310 of the subject 300. The three-dimensional model 301 is placed in the virtual space VS. The three-dimensional model 301 shows the three-dimensional shape of the subject 300. The three-dimensional model 301 is composed of point cloud data. The point cloud data is data that represents a point cloud 321. The point cloud 321 is a collection of a plurality of points 322. The point cloud data includes the three-dimensional coordinates of each point 322. The point cloud data may further include one or more of the following information for each point 322: color information (e.g., RGB values), normal vector information for each point, and reflection intensity information.
[0041] As an example, the model generation unit 11 generates a three-dimensional model 301 by performing SfM (Structure from Motion) processing. SfM processing is a process that generates a three-dimensional model 301 by using the principle of triangulation based on multiple two-dimensional images 150 generated by capturing a subject 300 having multiple feature points from multiple imaging positions 100. SfM processing preferably includes bundle adjustment. Bundle adjustment is a process that minimizes reprojection errors.
[0042] Specifically, first, the model generation unit 11 acquires a plurality of two-dimensional images 150 including the subject image 200 from the video data 511 or the still image data set 512 in the first database 51. Next, the model generation unit 11 extracts the feature points of the subject 300 from each two-dimensional image 150 and associates the feature points between the two-dimensional images 150 (feature point matching). The associated feature points are called tie points. In feature point matching, the model generation unit 11 uses, for example, SIFT (Scale Invariant Feature Transformation), SURF (Speeded-Up Robust Features), or FAST (Features from Accelerated Segment Test).
[0043] Next, the model generation unit 11 estimates the imaging position 100 (the position of the camera CM) and the posture (orientation) of the camera CM when each two-dimensional image 150 was generated based on the associated feature points, and calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimation result. In this case, the model generation unit 11 optimizes the estimation results of the imaging position 100 and the posture of the camera CM, as well as the three-dimensional coordinates of the three-dimensional points, by performing bundle adjustment. The model generation unit 11 calculates the three-dimensional coordinates of a plurality of three-dimensional points corresponding to a plurality of feature points. The three-dimensional coordinates in this case indicate the coordinates in the three-dimensional coordinate system CS set in the virtual space VS. The three-dimensional coordinate system CS is defined by the X-axis, Y-axis, and Z-axis that are orthogonal to each other. The three-dimensional coordinate system CS may be a coordinate system with a predetermined position in the virtual space VS as the origin, or may be a coordinate system to which geospatial coordinates (coordinates on the ground) or actual size information is assigned.
[0044] The plurality of three-dimensional points constitute a point cloud 321. That is, each three-dimensional point is each point 322 constituting the point cloud 321. As described above, the model generation unit 11 generates point cloud data. In this case, for example, the point cloud data indicates a sparse point cloud 321. Since the three-dimensional model 301 is constituted by the point cloud data, the generation of the point cloud data is synonymous with the generation of the three-dimensional model 301.
[0045] The model generation unit 11 may add one or more types of information, such as color information, luminance information, normal vector information, and reflection intensity information, to each point 322 that constitutes the point cloud 321. Further, in addition to the SfM process, the model generation unit 11 may execute an MVS (Multi View Stereo) process. The MVS process is a process of calculating depth and normal for each pixel of each two-dimensional image 150 by multi-view stereo measurement, integrating them, and generating a dense point cloud 321 of the subject 300. Therefore, the model generation unit 11 can generate point cloud data indicating the dense point cloud 321 by executing the MVS process. In this case, the three-dimensional model 301 is constituted by the dense point cloud 321.
[0046] As described above, in the process of generating the three-dimensional model 301, the model generation unit 11 estimates the imaging position 100 and the posture of the camera CM for each two-dimensional image 150.
[0047] The imaging position 100 indicates the position of the camera CM when the subject 300 is imaged to generate the two-dimensional image 150. The estimation result of the imaging position 100 (hereinafter, "imaging position information 521") is indicated by three-dimensional coordinates (hereinafter, "imaging three-dimensional coordinates") in the three-dimensional coordinate system CS. As shown in FIG. 4, in the present embodiment, the imaging position 100 in the real space RS is indicated as "imaging position 101" in the virtual space VS. Note that the imaging position information 521 is estimated for each two-dimensional image 150.
[0048] The posture of the camera CM indicates the posture (orientation) of the camera CM when the subject 300 is imaged to generate the two-dimensional image 150. The estimation result of the posture of the camera CM (hereinafter, "imaging posture information 522") is indicated by the rotation angle of the camera CM around the X axis, the rotation angle of the camera CM around the Y axis, and the rotation angle of the camera CM around the Z axis in the three-dimensional coordinate system CS. In FIG. 4, for simplicity of the drawing, the imaging posture information 522 is indicated by an arrow. Note that the imaging posture information 522 is estimated for each two-dimensional image 150. Further, when estimating the posture of the camera CM, the model generation unit 11 may execute a process of reducing the influence of gimbal lock. Gimbal lock is a phenomenon in which the three degrees of freedom of rotation become two degrees of freedom when expressing a three-dimensional posture by Euler angles.
[0049] Furthermore, it is preferable for the model generation unit 11 to assign faces to the three-dimensional model 301 based on the point cloud data. The process of assigning faces is, for example, a process of converting the point cloud data constituting the three-dimensional model 301 into mesh data. In this case, for example, the model generation unit 11 generates a plurality of polygonal faces (for example, triangles) by connecting points 322 of the point cloud data constituting the three-dimensional model 301, and represents the three-dimensional model 301 with the plurality of polygonal faces. In this case, for example, the model generation unit 11 generates TIN (Triangulated Irregular Network) data based on the point cloud data, and represents the three-dimensional model 301 with the TIN data. As a result, the three-dimensional model 301 is represented by a set of triangular faces.
[0050] The process of assigning faces to the three-dimensional model 301 is not particularly limited and may be performed in the following ways. For example, the model generation unit 11 assigns faces to each point 322 that constitutes the point cloud data. In this case, the faces are, for example, circular faces centered on each point 322 that constitutes the point cloud data. Alternatively, for example, the model generation unit 11 assigns faces to the three-dimensional model 301 composed of point cloud data by executing an octree algorithm. In this case, specifically, the model generation unit 11 represents the three-dimensional model 301 by multiple cubes by recursively dividing the three-dimensional model 301 into eight cubes (octants). Alternatively, for example, the model generation unit 11 assigns faces to the three-dimensional model 301 by converting the point cloud data into surface data.
[0051] Furthermore, the model generation unit 11 may add material information to the surfaces applied to the three-dimensional model 301 based on the two-dimensional image 150. The material information includes information on the color and pattern of the subject 300. For example, the model generation unit 11 may perform a process of mapping a texture to each surface (e.g., each polygon) that constitutes the mesh data of the three-dimensional model 301. In this case, the mesh data is, for example, TIN data.
[0052] The storage unit 50 stores the three-dimensional model 301 (point cloud data constituting the three-dimensional model 301) in the second database 52. Furthermore, the storage unit 50 stores multiple two-dimensional images 150, multiple imaging position information 521, and multiple imaging orientation information 522 in the second database 52 in association with the three-dimensional model 301. In the second database 52, the two-dimensional images 150 are stored in association with the imaging position information 521 and the imaging orientation information 522.
[0053] Furthermore, the method for generating the three-dimensional model 301 is not particularly limited, as long as the imaging position information 521 and imaging orientation information 522 can be obtained. For example, the model generation unit 11 may generate the three-dimensional model 301 by 3D Gaussian splatting.
[0054] Continuing with the reference to Figure 2, the distortion correction unit 12 of the calculation unit 10 will be explained. The distortion correction unit 12 obtains the two-dimensional image 150 to be processed (hereinafter referred to as "target two-dimensional image 150A") from the second database 52, which is one of the multiple two-dimensional images 150 used when generating the three-dimensional model 301.
[0055] Furthermore, the distortion correction unit 12 acquires calibration information for the camera CM from the second database 52. Based on the calibration information, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 150A. In this case, the distortion is, for example, lens distortion. Lens distortion is, for example, distortion aberration.
[0056] Next, the image placement unit 13 of the calculation unit 10 will be described with reference to Figures 2, 5, and 6(a). Figures 5 and 6(a) to 6(d) show a camera CM for ease of understanding.
[0057] Figure 5 is a schematic perspective view showing the imaging position 101, the target two-dimensional image 151, and the three-dimensional model 301 in the virtual space VS. As shown in Figure 5, in this embodiment, the target two-dimensional image 150A placed in the virtual space VS is referred to as "target two-dimensional image 151". Furthermore, the subject image 200 and target image 210 of the target two-dimensional image 150A placed in the virtual space VS are referred to as "subject image 201" and "target image 211", respectively.
[0058] Figure 6(a) is a schematic side view showing the imaging position 101 represented in the virtual space VS. Figure 6(b) is a schematic side view showing the target two-dimensional image 151 placed in the virtual space VS.
[0059] First, as shown in Figure 2, the image placement unit 13 acquires imaging position information 521 and imaging orientation information 522 associated with the target two-dimensional image 150A from the second database 52. The imaging position information 521 indicates the imaging three-dimensional coordinates (aX, aY, aZ). As shown in Figures 5 and 6(a), the imaging position 101 in the virtual space VS is indicated by the imaging three-dimensional coordinates (aX, aY, aZ).
[0060] Next, as shown in Figures 5 and 6(b), the image placement unit 13 places the distortion-corrected target two-dimensional image 151 in the virtual space VS based on the imaging position information 521 and the imaging orientation information 522.
[0061] Specifically, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS at a predetermined distance L away from the imaging three-dimensional coordinates (aX, aY, aZ) along the optical axis 523. In this case, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS such that the target two-dimensional image 151 is orthogonal to the optical axis 523 of the camera CM, and the size of the target two-dimensional image 151 corresponds to the field of view θ of the camera CM.
[0062] In this case, the image placement unit 13 calculates the three-dimensional coordinates (bX, bY, bZ), (cX, cY, cZ), and (dX, dY, dZ) in the virtual space VS where at least three of the four corners 161, 162, 163, and 164 of the target two-dimensional image 151 are placed. The image placement unit 13 then places the target two-dimensional image 151 in the virtual space VS by arranging at least three of the four corners 161 to 163 of the target two-dimensional image 151 at the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ).
[0063] Next, the target detection unit 14 of the calculation unit 10 will be described with reference to Figures 2, 5, 6(c), and 7. Figure 6(c) is a schematic side view showing the target image 211 and bounding box 221 included in the target two-dimensional image 151 placed in the virtual space VS, and the target region 311 included in the three-dimensional model 301. Figure 7 is a front view showing the target two-dimensional image 151 placed in the virtual space VS. In Figure 7, the target two-dimensional image 151 is viewed from the imaging position 101.
[0064] As shown in Figure 2, preferably, the target detection unit 14 utilizes a trained model 53 constructed by training on training data. The trained model 53 is constructed by training on training data. In this case, the training is supervised learning. As an example, the trained model 53 uses YOLO (You Only Look Once), an object detection algorithm that utilizes a convolutional neural network. The object detection algorithm is not particularly limited and may be, for example, R-CNN (Regions with Convolutional Neural Network) or SSD (Single Shot Multi-Box Detector). The training data includes a two-dimensional image containing the target image and an information tag. The information tag is attached to the two-dimensional image. The information tag includes a bounding box surrounding the target image contained in the subject image in the two-dimensional image. The information tag may further include the name of the target image. The bounding box corresponds to an example of "information identifying the target image" in this disclosure.
[0065] The object detection algorithm may also be a segmentation algorithm. In this case, the object detection algorithm may be, for example, "U-NET" which performs semantic segmentation, "Mask R-CNN" which performs instant segmentation, or "Panoptic Feature Pyramid Network" which performs panoptic segmentation.
[0066] As shown in Figures 2, 5, 6(c), and 7, the target detection unit 14, as an example, causes the trained model 53 to detect the target image 211 contained in the target two-dimensional image 151 and sets a bounding box 221 surrounding the target image 211. In this way, the target detection unit 14 causes the trained model 53 to detect the target image 211 and identifies the target image 211. In other words, when the target detection unit 14 inputs the target two-dimensional image 151 to the trained model 53, the trained model 53 outputs the target two-dimensional image 151 including the target image 211 surrounded by the bounding box 221. Note that in Figure 6(c), the bounding box 221 is exaggerated for easier viewing.
[0067] In this case, the target detection unit 14 uses a pre-trained model 53, but for example, a rule-based algorithm may be used to detect the target image 211.
[0068] Next, the target coordinate calculation unit 15 will be described with reference to Figures 2, 5, 6(c), 6(d), and 7. Figure 6(d) is a schematic side view showing a state in the virtual space VS where a straight line 400 passing through the imaging position 101 and the target image 211 intersects with the three-dimensional model 301.
[0069] As shown in Figures 2, 5, 6(c), and 7, the target coordinate calculation unit 15 calculates the three-dimensional coordinates (eX, eY, eZ) of the target image 211 contained in the target two-dimensional image 151 in the virtual space VS. Hereinafter, the three-dimensional coordinates (eX, eY, eZ) will be referred to as "first target three-dimensional coordinates (eX, eY, eZ)".
[0070] Specifically, the target coordinate calculation unit 15 calculates the first target three-dimensional coordinates (eX, eY, eZ) based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.
[0071] Preferably, as shown in Figure 7, the first target three-dimensional coordinates (eX, eY, eZ) are indicated by the three-dimensional coordinates in the virtual space VS of a specific point 224 in the bounding box 221. For example, the specific point 224 is the intersection of diagonals 222 and 223 of the bounding box 221. In this way, the specific point 224 is determined based on the bounding box 221. The specific point 224 represents a point inside the target image 211. The bounding box 221 is a rectangle surrounding the target image 211.
[0072] In this case, specifically, the target coordinate calculation unit 15 calculates the two-dimensional coordinates (u, v) of a specific point 224 in the bounding box 221 of the target two-dimensional image 151. The two-dimensional coordinates (u, v) are coordinates in the screen coordinate system SC. The screen coordinate system SC is defined by mutually orthogonal U-axis and V-axis. The origin of the screen coordinate system SC is not particularly limited, but for example, it is set at a corner 161 of the target two-dimensional image 151. The origin may also be set at the center of the target two-dimensional image 151.
[0073] Then, the target coordinate calculation unit 15 converts the two-dimensional coordinates (u, v) of a specific point 224 in the screen coordinate system SC into three-dimensional coordinates, which are the first target three-dimensional coordinates (eX, eY, eZ), based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.
[0074] The specific point 224 was determined based on the bounding box 221, but is not particularly limited as long as it is a point inside the target image 211.
[0075] Next, as shown in Figures 5 and 6(d), the target coordinate calculation unit 15 calculates the three-dimensional coordinates (fX, fY, fZ) of the position where the straight line 400, passing through the imaging three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ), intersects with the three-dimensional model 301. In this case, the three-dimensional coordinates (fX, fY, fZ) are referred to as the "second target three-dimensional coordinates (fX, fY, fZ)". In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point when the straight line 400 intersects with the three-dimensional model 301. The straight line 400 is, for example, a virtual straight line.
[0076] The second target three-dimensional coordinates (fX, fY, fZ) indicate the three-dimensional coordinates of the target region 311 on the three-dimensional model 301. The target region 311 corresponds to the target image 211 and is the region on the three-dimensional model 301 that reproduces the target 310 in the subject 300.
[0077] Specifically, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ), which are the positions where the straight line 400 intersects with the surface of the three-dimensional model 301. In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point where the straight line 400 intersects with the surface of the three-dimensional model 301.
[0078] For example, if a three-dimensional model 301 is assigned multiple polygonal faces (e.g., triangles) with point 322 of the point cloud 321 as its vertex, the target coordinate calculation unit 15 sets the average of the three-dimensional coordinates of the multiple vertices that define the face to which the line 400 meets as the second target three-dimensional coordinate. Alternatively, for example, the target coordinate calculation unit 15 sets the three-dimensional coordinate of the vertex closest to the point to which the line 400 meets as the second target three-dimensional coordinate.
[0079] For example, if a circular surface is assigned to the three-dimensional model 301, with each point 322 constituting the point cloud 321 as its center, the target coordinate calculation unit 15 sets the three-dimensional coordinate of point 322 at the center of the circular surface where the straight line 400 intersects as the second target three-dimensional coordinate.
[0080] For example, if multiple cubes are assigned to a three-dimensional model 301 using an octree algorithm, the target coordinate calculation unit 15 sets the three-dimensional coordinates of point 322, which is contained within the cube having the face that the line 400 intersects, as the second target three-dimensional coordinates. If multiple points 322 are contained within the cube, the average of the three-dimensional coordinates of the multiple points 322 is set as the second target three-dimensional coordinates.
[0081] As described above with reference to Figures 5 and 6, according to this embodiment, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ), which are the three-dimensional coordinates of the target region 311 on the three-dimensional model 301, by determining the position where a straight line 400 passing through the imaging three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ) of the target two-dimensional image 151 intersects with the three-dimensional model 301 in the virtual space VS. Thus, in this embodiment, instead of performing a process to search for the target region over the entire area of the three-dimensional model, the position of the target region 311 on the three-dimensional model 301 is determined based on the imaging three-dimensional coordinates (aX, aY, aZ) and a single target two-dimensional image 151. In other words, the position of the target region 311 on the three-dimensional model 301 is determined by processing a two-dimensional image (target two-dimensional image 151). Therefore, the position of the target region 311 on the three-dimensional model 301 can be determined by a simple process.
[0082] In addition, according to this embodiment, the calculation unit 10 automatically identifies the position of the target region 311 on the three-dimensional model 301. Therefore, the manual work of searching for the target region 311 on the three-dimensional model 301 can be omitted. Thus, the workload on humans can be reduced.
[0083] In particular, the imaging position information 521 (imaging three-dimensional coordinates (aX, aY, aZ)) and the imaging orientation information 522 are information that is always calculated as a by-product when generating the three-dimensional model 301. Therefore, it is possible to suppress the occurrence of additional processing for the process of identifying the target region 311 on the three-dimensional model 301.
[0084] Furthermore, according to this embodiment, the model generation unit 11 assigns surfaces to the three-dimensional model 301. Therefore, the straight line 400 always intersects with the three-dimensional model 301. Thus, the second target three-dimensional coordinates (fX, fY, fZ) indicating the position of the target region 311 can be reliably calculated without depending on the density of the point cloud 321.
[0085] Furthermore, according to this embodiment, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 151. Therefore, even if the target image 211 is located in a region of the target two-dimensional image 151 where distortion occurs, the two-dimensional coordinates (u, v) indicating the position of the target image 211 and the first target three-dimensional coordinates (eX, eY, eZ) can be calculated with high accuracy.
[0086] Furthermore, according to this embodiment, the target detection unit 14 inputs the target two-dimensional image 151 into the trained model 53 to detect the target image 211. In other words, the trained model 53 is constructed by learning the two-dimensional image 150 as training data. Therefore, compared to the case where a trained model is constructed to detect the target image by training a three-dimensional model, the burden on the worker when creating the training data, the capacity of the training data, and the amount of computation required for training can be reduced. As a result, the cost and training time when constructing the trained model 53 can be reduced. In particular, instead of detecting the target region by analyzing a three-dimensional model, the target image 211 is detected on the two-dimensional image (target two-dimensional image 151). Therefore, according to this embodiment, compared to the case where the target region is detected by analyzing a three-dimensional model, the target region 311 on the three-dimensional model 301 can be identified with high speed and simple processing.
[0087] Furthermore, according to this embodiment, the first target three-dimensional coordinates (eX, eY, eZ) are the three-dimensional coordinates of a specific point 224 determined based on the bounding box 221. In other words, the specific point 224 is determined as the intersection of the diagonals 222 and 223 of the bounding box 221. Therefore, compared to the case where the specific point 224 is determined as any point in the target image 211, the specific point 224 can be easily determined with a simpler process.
[0088] Next, an image analysis method according to this embodiment will be described with reference to Figures 2 and 8. Figure 8 is a flowchart of the image analysis method. The image analysis method is executed by the server 1. As shown in Figure 8, the image analysis method includes steps S1 to S9. The computer program stored in the storage unit 50 causes the arithmetic unit 10 to execute steps S1 to S9. In other words, the computer program product realizes steps S1 to S9 when the computer program is executed by the arithmetic unit 10. The arithmetic unit 10 corresponds to an example of a "computer" in this disclosure.
[0089] As shown in Figures 2 and 8, first, in step S1, the model generation unit 11 generates a three-dimensional model 301 of the subject 300 based on a plurality of two-dimensional images 150 generated by capturing the subject 300 from a plurality of different imaging positions 100.
[0090] Next, in step S2, the distortion correction unit 12 acquires the target two-dimensional image 150A from the second database 52 among the multiple two-dimensional images 150.
[0091] Next, in step S3, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 150A based on the calibration information. The distortion-corrected target two-dimensional image 150A is stored in the storage unit 50.
[0092] Next, in step S4, the image placement unit 13 acquires imaging position information 521 and imaging orientation information 522 from the second database 52. The imaging position information 521 indicates the position of the camera CM at the time of imaging using the imaging three-dimensional coordinates (aX, aY, aZ) of the three-dimensional coordinate system CS. The imaging orientation information 522 indicates the orientation of the camera CM at the time of imaging using the rotation angle in the three-dimensional coordinate system CS.
[0093] Next, in step S5, the image placement unit 13 places the distortion-corrected target two-dimensional image 150A in the virtual space VS based on the imaging position information 521 and the imaging orientation information 522. The target two-dimensional image 150A placed in the virtual space VS will be referred to as "target two-dimensional image 151".
[0094] Next, in step S6, the target detection unit 14 causes the trained model 53 to detect the target image 211 included in the target two-dimensional image 151 and sets a bounding box 221 surrounding the target image 211.
[0095] Next, in step S7, the target coordinate calculation unit 15 calculates the first target three-dimensional coordinates (eX, eY, eZ) which represent the three-dimensional coordinates of a specific point 224 in the bounding box 221.
[0096] Next, in step S8, the target coordinate calculation unit 15 calculates a second target three-dimensional coordinate (fX, fY, fZ) which indicates the position where a straight line 400 passing through the imaging three-dimensional coordinate (aX, aY, aZ) indicated by the imaging position information 521 and the first target three-dimensional coordinate (eX, eY, eZ) intersects with the three-dimensional model 301. The second target three-dimensional coordinate (fX, fY, fZ) indicates the three-dimensional coordinate of the target region 311 on the three-dimensional model 301. In this way, the position of the target 310 included in the subject 300 is identified as the target region 311 on the three-dimensional model 301.
[0097] Next, in step S9, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 via the communication unit 40 in response to a request from the user's terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in an operable manner. For example, the display control unit 16 overlays the target two-dimensional image 151 onto the three-dimensional model 301 and displays it on the terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in a position where the target area 311 on the three-dimensional model 301 is visible to the terminal 2. In this case, for example, the display control unit 16 enlarges the target area 311 on the three-dimensional model 301 and displays it on the terminal 2. Therefore, according to this embodiment, the user of the terminal 2 can omit the task of visually searching for the target area 311 on the three-dimensional model 301. In other words, the workload of the user can be reduced. When step S9 is completed, the image analysis method is finished.
[0098] As described above with reference to Figure 8, according to the image analysis method of this embodiment, the position of the target region 311 on the three-dimensional model 301 is identified by calculating the position where the straight line 400 intersects with the three-dimensional model 301 in the virtual space VS. Therefore, the position of the target region 311 can be identified by a simple process.
[0099] Figure 9 is a flowchart showing the three-dimensional model generation process in step S1 of Figure 8. As shown in Figure 9, the three-dimensional model generation process includes steps S11 to S15.
[0100] First, in step S11, the model generation unit 11 acquires a plurality of two-dimensional images 150, including the subject image 200, from the first database 51.
[0101] Next, in step S12, the model generation unit 11 extracts feature points of the subject 300 from each two-dimensional image 150 and associates the feature points with each other in the two-dimensional images 150.
[0102] Next, in step S13, the model generation unit 11 estimates the imaging position 100 and the orientation of the camera CM at the time each two-dimensional image 150 was generated, based on the associated feature points. Then, the model generation unit 11 calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimated result of the imaging position 100 (imaging position information 521) and the estimated result of the orientation of the camera CM (imaging orientation information 522). As a result, point cloud data representing a point cloud 321 consisting of multiple three-dimensional points (multiple points 322) is generated.
[0103] Next, in step S14, the model generation unit 11 adds surfaces to the three-dimensional model 301 based on the point cloud data.
[0104] Next, in step S15, the model generation unit 11 adds material information to the surfaces applied to the three-dimensional model 301 based on the two-dimensional image 150. Then, the process returns to step S2 in Figure 8.
[0105] (Modified Version) A modified version of this embodiment will be described with reference to Figures 2 and 10. In the modified version, the target 310 included in the subject 300 is the part of the subject 300 that shows a specific state. The following will mainly describe the differences between the modified version and the above embodiment.
[0106] In the modified example, the target image 211 shown in Figure 7 is an image of a portion of the subject 300 that represents a specific state (hereinafter referred to as the "specific state portion"). Therefore, according to the modified example, the specific state portion of the subject 300 is identified as the target region 311 on the three-dimensional model 301. In other words, the position of the target region 311 representing the specific state portion of the subject 300 is calculated as the second target three-dimensional coordinate on the three-dimensional model 301. As a result, the manual work of visually searching for the target region 311 representing the specific state portion on the three-dimensional model 301 can be omitted. Furthermore, the possibility of missing the detection of the target region 311 representing the specific state portion can be suppressed.
[0107] The specific state portion of the subject 300 is, for example, the entire subject 300 or a portion of the area of interest that shows an abnormal state that is not in a normal state (hereinafter referred to as the "abnormal state portion"). Therefore, in this case, the target image 211 is an image of the abnormal state portion. As a result, the human task of visually searching for the target region 311 representing the abnormal state portion on the three-dimensional model 301 can be omitted. In addition, the failure to detect the target region 311 representing the abnormal state portion can be suppressed.
[0108] For example, the target image 211 shows a region of the target two-dimensional image 151 that has changed relative to the reference two-dimensional image. The region that has changed relative to the reference two-dimensional image is, for example, a region that has been deformed, mutated, displaced, increased, decreased, or disappeared relative to the reference two-dimensional image. The reference two-dimensional image is, for example, a two-dimensional image of the subject 300 that was captured earlier than the target two-dimensional image 151. Alternatively, the reference two-dimensional image is, for example, a two-dimensional image generated by capturing the subject 300 in a reference state. The reference state is, for example, a state in which the entire subject 300 or the region of interest is in a normal state. In this example, the target image 211 is an image of the entire subject 300 or a portion of the region of interest that shows an abnormal state.
[0109] Next, an image analysis method relating to a modified example will be described with reference to Figures 2 and 10. Figure 10 is a flowchart of the image analysis method relating to a modified example. The image analysis method is executed by Server 1. As shown in Figure 10, the image analysis method includes steps S21 to S23. The computer program stored in the storage unit 50 causes the arithmetic unit 10 to execute steps S21 to S23. In other words, the computer program product realizes steps S21 to S23 when the computer program is executed by the arithmetic unit 10.
[0110] First, in step S21, the target detection unit 14 performs a target image detection process. The target image detection process involves detecting a target image 210 from the video data 511 of the subject 300, the two-dimensional image 150 extracted from the video data 511, the still image dataset 512, or the two-dimensional image 150 extracted from the still image dataset 512. Hereinafter, the video data 511, the two-dimensional image 150 extracted from the video data 511, the still image dataset 512, and the still image dataset 512 will be collectively referred to as "detection target image data".
[0111] For example, the target detection unit 14 inputs the image data to be detected into the trained model. If the trained model detects the target image 210 from the image data to be detected, it outputs information indicating that the image data to be detected contains the target image 210. If the trained model does not detect the target image 210 from the image data to be detected, it outputs information indicating that the image data to be detected does not contain the target image 210.
[0112] A trained model is a computer program built by learning from training data. Examples of trained models include neural networks. Training can be, for example, supervised learning. Training data consists of image data to be detected and information indicating the presence of the target image.
[0113] Next, in step S22, the target detection unit 14 determines whether or not the target image 210 has been detected in the target image detection process.
[0114] If a negative result is obtained in step S22 (NO), the image analysis method is terminated.
[0115] On the other hand, if a positive determination is made in step S22 (YES), the process proceeds to step S23.
[0116] Next, in step S23, the calculation unit 10 identifies a target region 311 on the three-dimensional model 301 corresponding to the target image 210. The processing in step S23 is the same as the processing in steps S1 to S9 shown in Figure 8. When step S23 is completed, the image analysis method is finished.
[0117] As explained above with reference to Figure 10, in this modified example, the process in step S23 is executed only when the target image 210 is detected. Therefore, the processing load on the calculation unit 10 can be reduced. In particular, if the target image 210 is not detected, the three-dimensional model 301 is not generated, thus reducing the processing load on the calculation unit 10. In this case, the increase in the amount of data stored in the storage unit 50 can also be suppressed.
[0118] Furthermore, the image analysis method shown in Figure 10 can also be applied to the above-described embodiment with reference to Figures 1 to 9.
[0119] Preferred embodiments and modifications of the present disclosure have been described in detail above with reference to the attached drawings, but the technical scope of the present disclosure is not limited to such examples. It is clear to any person with ordinary skill in the art of the present disclosure that various modifications or alterations can be conceived within the scope of the technical idea set forth in the claims, and these too are understood to fall within the technical scope of the present disclosure.
[0120] The apparatus or system described herein may be implemented as a single apparatus, or as a group of apparatuses (e.g., a cloud server) partially or entirely connected by a network. For example, some or all of the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by the same computer or server. Alternatively, for example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by a terminal 2. For example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may each be implemented by separate computers or servers. For example, the first database 51, the second database 52, and the trained model 53 may each be stored in separate storage devices or servers.
[0121] The series of processes performed by the apparatus described herein may be implemented using software, hardware, or a combination of software and hardware. Computer programs for implementing each function of the arithmetic unit 10 according to this embodiment can be created and implemented on a PC or the like. Furthermore, a computer-readable recording medium containing such a computer program can also be provided. Examples of recording media include magnetic disks, optical disks, magneto-optical disks, and flash memory. Alternatively, the computer program may be distributed without using a recording medium, for example, via a network.
[0122] Furthermore, the processes described using flowcharts in this specification do not necessarily have to be executed in the order shown. Some processing steps may be executed in parallel. Additional processing steps may be adopted, and some processing steps may be omitted.
[0123] Furthermore, the effects described herein are merely descriptive or illustrative and not limiting. In other words, the technology relating to this disclosure may produce other effects that will be apparent to those skilled in the art from the description herein, in addition to or in lieu of the effects described herein.
[0124] Furthermore, the following configurations also fall within the technical scope of this disclosure.
[0125] (Item 1) An image analysis device comprising: a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, which are indicated by three-dimensional coordinates in a virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging; and a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model, wherein the first target three-dimensional coordinate indicates the position in the virtual space of a target image included in the target two-dimensional image.
[0126] (Item 2) The image analysis apparatus according to Item 1, wherein the three-dimensional model is composed of point cloud data, the model generation unit assigns a surface to the three-dimensional model based on the point cloud data, and the target coordinate calculation unit calculates the second target three-dimensional coordinate, which is the position where the line strikes the surface.
[0127] (Item 3) The image analysis apparatus according to Item 1 or 2, further comprising a distortion correction unit for correcting distortion of the target two-dimensional image, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the distortion-corrected target two-dimensional image.
[0128] (Item 4) An image analysis device according to any one of Items 1 to 3, further comprising a target detection unit that causes a trained model constructed by training on training data to detect the target image contained in the target two-dimensional image and identify the target image, wherein the training data includes a two-dimensional image containing the target image and an information tag, and the information tag includes information that identifies the target image.
[0129] (Item 5) The image analysis device according to Item 4, wherein the target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image, and the first target three-dimensional coordinates indicate the three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.
[0130] (Item 6) The image analysis device described in any of Items 1 to 5, wherein the target image is an image of a part of the subject that shows a specific state.
[0131] (Item 7) An image analysis method comprising: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquiring imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, and imaging pose information indicating the pose of the camera that performed the imaging; placing the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculating a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
[0132] (Item 8) A computer program that causes a computer to perform the following steps: generate a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquire imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; and imaging pose information indicating the pose of the camera that performed the imaging; place the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculate a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculate a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
[0133] This disclosure provides an image analysis device, an image analysis method, and a computer program, and has industrial applicability.
[0134] 1 Server (image analysis device), 11 Model generation unit, 12 Distortion correction unit, 13 Image placement unit, 14 Target detection unit, 15 Target coordinate calculation unit, 16 Display control unit, 53 Trained model
Claims
1. An image analysis device comprising: a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, which are indicated by three-dimensional coordinates in a virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging; and a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects with the three-dimensional model, wherein the first target three-dimensional coordinate indicates the position in the virtual space of a target image included in the target two-dimensional image.
2. The image analysis apparatus according to claim 1, wherein the three-dimensional model is composed of point cloud data, the model generation unit assigns a surface to the three-dimensional model based on the point cloud data, and the target coordinate calculation unit calculates the second target three-dimensional coordinate which is the position where the line strikes the surface.
3. The image analysis apparatus according to claim 1 or claim 2, further comprising a distortion correction unit for correcting distortion of the target two-dimensional image, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the distortion-corrected target two-dimensional image.
4. The image analysis device according to claim 1 or 2, further comprising a target detection unit that causes a trained model constructed by training on training data to detect the target image contained in the target two-dimensional image and identify the target image, wherein the training data includes a two-dimensional image containing the target image and an information tag, and the information tag includes information that identifies the target image.
5. The image analysis apparatus according to claim 4, wherein the target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image, and the first target three-dimensional coordinates indicate the three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.
6. The image analysis apparatus according to claim 1 or claim 2, wherein the target image is an image of a portion of the subject that shows a specific state.
7. An image analysis method comprising: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquiring imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, and imaging pose information indicating the pose of the camera that performed the imaging; placing the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculating a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
8. A computer program that causes a computer to perform the following steps: generate a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquire imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; and imaging pose information indicating the pose of the camera that performed the imaging; place the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculate a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculate a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.
Citation Information
Patent Citations
Deterioration diagnosis system using flight vehicle
JP2019070631A
Inspection system and inspection method
JP2020021466A