Image analysis device, image analysis method, and computer program

The image analysis device and method efficiently identify the position of a target region on a three-dimensional model by generating the model from two-dimensional images and calculating intersection points, simplifying the process and reducing manual effort.

JP2026062603APending Publication Date: 2026-04-09CALTA INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

It is difficult to search for and specify the position of a diagnostic target area on a three-dimensional model.

Method used

An image analysis device and method that generates a three-dimensional model from multiple two-dimensional images, places the target image in a virtual space based on imaging coordinates and pose information, and calculates the position where a straight line through the imaging coordinates intersects the model to identify the target region.

Benefits of technology

Enables simple and accurate identification of the target region on the three-dimensional model, reducing manual effort and computational burden, and improving processing speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062603000001_ABST
    Figure 2026062603000001_ABST
Patent Text Reader

Abstract

This invention provides an image analysis device, an image analysis method, and a computer program that can easily identify the location of a target region on a three-dimensional model. [Solution] The server, which is an image analysis device, comprises: a model generation unit that generates a three-dimensional model of a subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions; an image placement unit that places the target two-dimensional image in a virtual space based on the imaging position when the target two-dimensional image was generated from the multiple two-dimensional images, the imaging three-dimensional coordinates indicated by the three-dimensional coordinates in a virtual space where the three-dimensional model is placed, and imaging pose information indicating the pose of the camera that performed the imaging; and a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinates and the first target three-dimensional coordinates intersects with the three-dimensional model. The first target three-dimensional coordinate indicates the position in the virtual space of the target image included in the target two-dimensional image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image analysis apparatus, an image analysis method, and a computer program.

Background Art

[0002] The diagnostic system described in Patent Document 1 has a function of automatically determining and detecting deteriorated portions by performing diagnostic processing by a computer based on continuous images obtained by aerial photography of an object in the past and at present. The object is a building, infrastructure equipment, or the like. A predetermined area on the surface of the object is a diagnostic target area. Based on a user's input operation, the diagnostic target area of the object is set as a basic setting.

[0003] The diagnostic system described in Patent Document 1 has an SFM processing unit. The SFM processing unit performs SFM processing on a plurality of input images to restore a three-dimensional structure and outputs a three-dimensional model.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Generally, there are cases where a diagnostic target area is searched from a three-dimensional model of an object and the diagnostic target area is diagnosed.

[0006] However, it is often difficult to search for a target area such as a diagnostic target area from a three-dimensional model and specify the position of the target area on the three-dimensional model.

[0007] Therefore, this disclosure has been made in view of the above-mentioned problems, and its purpose is to provide an image analysis device, an image analysis method, and a computer program that can identify the position of a target region on a three-dimensional model through simple processing. [Means for solving the problem]

[0008] According to this disclosure, an image analysis device is provided, comprising: a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images, which are indicated by three-dimensional coordinates in a virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging; and a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate strikes the three-dimensional model, wherein the first target three-dimensional coordinate indicates the position in the virtual space of a target image included in the target two-dimensional image.

[0009] Furthermore, the present disclosure provides an image analysis method that includes the steps of: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquiring imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; imaging pose information indicating the pose of the camera that performed the imaging; placing the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculating a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.

[0010] Furthermore, the present disclosure provides a computer program that causes a computer to perform the following steps: generate a three-dimensional model of a subject based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions; acquire imaging three-dimensional coordinates, which are three-dimensional coordinates in a virtual space where the three-dimensional model is placed, indicating the imaging position when a target two-dimensional image was generated from the plurality of two-dimensional images; and imaging pose information indicating the pose of the camera that performed the imaging; place the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging pose information; calculate a first target three-dimensional coordinate, which is the three-dimensional coordinate in the virtual space of a target image included in the target two-dimensional image; and calculate a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model. [Effects of the Invention]

[0011] This disclosure provides an image analysis device, an image analysis method, and a computer program that can identify the position of a target region on a three-dimensional model through a simple process. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram showing an example configuration of an image analysis system according to one embodiment of the present invention. [Figure 2] This block diagram shows an example of a server configuration according to the same embodiment. [Figure 3] This is a schematic perspective view showing the camera, a two-dimensional image, and the subject in real space according to the embodiment. [Figure 4] This is a schematic perspective view showing the imaging position and a three-dimensional model of the subject in the virtual space according to the same embodiment. [Figure 5]This is a schematic perspective view showing the imaging position, the target two-dimensional image, and the three-dimensional model in the virtual space according to the same embodiment. [Figure 6] (a) is a schematic side view showing the imaging position represented in the virtual space according to the embodiment; (b) is a schematic side view showing the target two-dimensional image placed in the virtual space; (c) is a schematic side view showing the target image and bounding box included in the target two-dimensional image placed in the virtual space, as well as the target region included in the three-dimensional model; and (d) is a schematic side view showing the state in the virtual space where a straight line passing through the imaging position and the target image intersects with the three-dimensional model. [Figure 7] This is a front view showing a target two-dimensional image placed in a virtual space according to the same embodiment. [Figure 8] This is a flowchart showing the image analysis method according to the present invention. [Figure 9] This is a flowchart showing the three-dimensional model generation process according to the same embodiment. [Figure 10] This flowchart shows an image analysis method according to a modified example of the same embodiment. [Modes for carrying out the invention]

[0013] Preferred embodiments of this disclosure will be described in detail below with reference to the attached drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions will be omitted.

[0014] Figure 1 is a diagram showing an example configuration of an image analysis system SYS according to one embodiment of the present invention. As shown in Figure 1, the image analysis system SYS includes a server 1. Server 1 corresponds to an example of the "image analysis device" in this disclosure.

[0015] Server 1 generates a three-dimensional model of a subject to be placed in a virtual space based on a plurality of two-dimensional images generated by imaging the subject from a plurality of different imaging positions. Server 1 calculates, in the virtual space, the position where the straight line passing through the imaging position and the target image included in one two-dimensional image hits the three-dimensional model (that is, the position of the target area on the three-dimensional model). As a result, the position of the target area on the three-dimensional model is specified. Thus, in the present embodiment, instead of executing a process of searching for the target area over the entire area of the three-dimensional model, the position of the target area on the three-dimensional model is calculated from the camera position and one two-dimensional image. Therefore, the position of the target area on the three-dimensional model can be specified by a simple process. Details of this point will be described later.

[0016] The image analysis system SYS further includes at least one terminal 2, at least one mobile device 3, or at least one imaging device 5.

[0017] Server 1, terminal 2, mobile device 3, and imaging device 5 are connected to the network NW. The network NW includes, for example, the Internet, a closed network, a public telephone network, a LAN (Local Area Network), and a short-range wireless network.

[0018] The terminal 2 is, for example, a personal computer (for example, a notebook personal computer, a desktop personal computer, or a tablet).

[0019] The mobile device 3 is, for example, an unmanned mobile device or a manned mobile device. The unmanned mobile device is, for example, an unmanned aircraft such as a drone, an unmanned ground vehicle, an unmanned submersible, or an unmanned watercraft. The unmanned ground vehicle is, for example, an unmanned ground vehicle or an unmanned ground vehicle modeled after a living thing (for example, a snake-shaped unmanned ground vehicle). The manned mobile device is, for example, an aircraft, an automobile, a ship, or a submarine. The mobile device 3 includes a camera 4.

[0020] The imaging device 5 includes a camera 6. The imaging device 5 is, for example, a mobile device such as a smartphone. The imaging device 5 may also be, for example, the camera 6 itself.

[0021] Cameras 4 and 6 capture images of the subject and generate video data showing a video containing images of the subject. The video is a collection of consecutive two-dimensional images. Cameras 4 and 6 may also generate multiple still image data, each showing multiple still images containing images of the subject. The still images are two-dimensional images.

[0022] Hereafter, video data will be referred to as "Video Data 511" (Figure 2). Similarly, multiple still image data will be referred to as "Still Image Dataset 512" (Figure 2).

[0023] In the following, unless it is necessary to distinguish between cameras 4 and 6, cameras 4 and 6 will be collectively referred to as "camera CM".

[0024] Furthermore, the subject is not particularly limited as long as it can be captured by the camera CM, and the size, shape, pattern, and color of the subject are not particularly limited. For example, the subject is one or more movable or immovable property. The subject is, for example, one or more objects. Typically, the subject is one or more stationary objects. Stationary objects are, for example, man-made or natural objects. Man-made objects are, for example, structures, machinery, electronic equipment, or copyrighted works. Structures are, for example, buildings or infrastructure facilities. Buildings are, for example, office buildings or houses. Infrastructure facilities are facilities for developing social infrastructure. For example, infrastructure facilities are roads, bridges, road traffic facilities, power generation facilities, power transmission facilities, water treatment facilities, or gas distribution facilities. Machinery is, for example, automobiles, work vehicles, trains, aircraft, ships, submarines, or robots. Natural objects are, for example, trees, forests, the ground surface, cliffs, coastlines, or rivers.

[0025] Furthermore, the subject includes one or more targets whose positions on the three-dimensional model are to be identified. The targets are not particularly limited and can be arbitrarily set, such as all of the region included in the subject, part of the region, all of the objects included in the subject, or parts of the objects.

[0026] Terminal 2 acquires video data 511 or still image data set 512 generated by camera CM. For example, terminal 2 receives video data 511 or still image data set 512 transmitted from camera CM via network NW. For example, terminal 2 acquires video data 511 or still image data set 512 from removable media. Terminal 2 transmits video data 511 or still image data set 512 to server 1 via network NW. Alternatively, mobile device 3 and imaging device 5 may transmit video data 511 or still image data set 512 to server 1 via network NW.

[0027] Figure 2 is a block diagram showing an example configuration of Server 1 in Figure 1. As shown in Figure 2, Server 1 includes an arithmetic unit 10, a communication unit 40, and a storage unit 50. Server 1 may also include an input unit 20 and a display unit 30.

[0028] The input unit 20 is an input device for inputting various types of information to the calculation unit 10. For example, the input unit 20 may be a keyboard and pointing device, or a touch panel.

[0029] The display unit 30 displays various information. The display unit 30 is, for example, a liquid crystal display or an organic electroluminescent display.

[0030] The communication unit 40 is connected to a network NW. The communication unit 40 communicates with external devices connected to the network NW. The external devices are, for example, a terminal 2, a mobile device 3, and an imaging device 5. The communication unit 40 is a communication device that performs communication according to a predetermined communication protocol, and includes, for example, a network interface controller. The predetermined communication protocol is, for example, a protocol compliant with Ethernet® and an Internet Protocol Suite.

[0031] The communication unit 40 receives video data 511 or still image data set 512 generated by imaging each subject from the terminal 2, mobile device 3, and imaging device 5 via the network NW.

[0032] The storage unit 50 includes a storage device and stores data and computer programs. The storage unit 50 includes a main storage device such as a semiconductor memory and an auxiliary storage device such as a semiconductor memory and a hard disk drive. The storage unit 50 may also include a removable medium such as an optical disc. The storage unit 50 may be, for example, a non-temporary computer-readable storage medium.

[0033] The memory unit 50 stores the first database 51, the second database 52, and the trained model 53.

[0034] The first database 51 includes video data 511 and still image data set 512 received by the communication unit 40. In the first database 51, the video data 511 and still image data set 512 are associated with attribute information (hereinafter, "attribute information AT"). Attribute information AT includes, for example, user information, imaging conditions, and subject information. User information includes, for example, identification information of the user of terminal 2, mobile device 3, or imaging device 5. Imaging conditions include, for example, the date and time of imaging and calibration information of camera CM. Calibration information includes, for example, information regarding lens distortion, focal length, and principal point position of camera CM. Information regarding lens distortion includes, for example, information regarding distortion aberration. Imaging conditions may also include information on the imaging location. The imaging location is indicated, for example, by the position coordinates of camera CM obtained by GPS (Global Positioning System) or GNSS (Global Navigation Satellite System). Subject information includes, for example, identification information of the subject. The subject information may include the coordinates of the ground control point (GCP).

[0035] The second database 52 contains multiple datasets 520. Each of the datasets 520 contains multiple two-dimensional images 150, a three-dimensional model 301 (point cloud data), multiple imaging location information 521, and multiple imaging orientation information 522. In the second database 52, the two-dimensional images 150, the three-dimensional model 301, the imaging location information 521, and the imaging orientation information 522 are related to each other and are also related to attribute information AT. Details of these will be described later.

[0036] The trained model 53 is a computer program. The trained model 53 includes, for example, a neural network. Details of the trained model 53 will be described later.

[0037] The arithmetic unit 10 performs various calculations. The arithmetic unit 10 includes processors such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).

[0038] Specifically, the calculation unit 10 includes a model generation unit 11, a distortion correction unit 12, an image placement unit 13, a target detection unit 14, a target coordinate calculation unit 15, and a display control unit 16. For example, the calculation unit 10 functions as the model generation unit 11, the distortion correction unit 12, the image placement unit 13, the target detection unit 14, the target coordinate calculation unit 15, and the display control unit 16 by executing a computer program stored in the storage unit 50.

[0039] Next, the processing performed by the arithmetic unit 10 will be explained with reference to Figures 2 to 6. Figure 3 is a schematic perspective view showing the camera CM, a two-dimensional image 150, and the subject 300 in real space RS. As shown in Figure 3, the subject 300 includes the target 310. The two-dimensional image 150 includes a subject image 200 that shows the subject 300. The subject image 200 includes a target image 210 that shows the target 310, depending on the imaging position 100. In this embodiment, the subject 300 is imaged while moving one camera CM, and multiple two-dimensional images 150 are generated. Figure 4 is a schematic perspective view showing the imaging position 101 and a three-dimensional model 301 of the subject 300 in virtual space VS. In Figure 4, the imaging position 101 is indicated by a black circle. Also, for ease of understanding, region A of the three-dimensional model 301 is shown enlarged.

[0040] As shown in Figures 2 to 4, the model generation unit 11 of the calculation unit 10 generates a three-dimensional model 301 of the subject 300 based on multiple two-dimensional images 150 generated by imaging the subject 300 from multiple different imaging positions 100. The three-dimensional model 301 includes a target region 311 that indicates the target 310 of the subject 300. The three-dimensional model 301 is placed in the virtual space VS. The three-dimensional model 301 shows the three-dimensional shape of the subject 300. The three-dimensional model 301 is composed of point cloud data. The point cloud data is data that represents a point cloud 321. The point cloud 321 is a collection of multiple points 322. The point cloud data includes the three-dimensional coordinates of each point 322. The point cloud data may further include one or more of the following information for each point 322: color information (e.g., RGB values), normal vector information for each point, and reflectance information.

[0041] As an example, the model generation unit 11 generates a three-dimensional model 301 by performing SfM (Structure from Motion) processing. SfM processing is a process that generates a three-dimensional model 301 by using the principle of triangulation based on multiple two-dimensional images 150 generated by capturing a subject 300 having multiple feature points from multiple imaging positions 100. SfM processing preferably includes bundle adjustment. Bundle adjustment is a process that minimizes reprojection errors.

[0042] Specifically, first, the model generation unit 11 obtains multiple two-dimensional images 150, including the subject image 200, from the video data 511 or still image dataset 512 of the first database 51. Next, the model generation unit 11 extracts feature points of the subject 300 from each two-dimensional image 150 and associates the feature points with each other (feature point matching). The associated feature points are called tie points. For feature point matching, the model generation unit 11 uses, for example, SIFT (Scale Invariant Feature Transformation), SURF (Speeded-Up Robust Features), or FAST (Features from Accelerated Segment Test).

[0043] Next, the model generation unit 11 estimates the imaging position 100 (position of camera CM) and the orientation (direction) of the camera CM when each two-dimensional image 150 is generated, based on the associated feature points, and calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimation results. In this case, the model generation unit 11 optimizes the estimation results of the imaging position 100 and the orientation of the camera CM, as well as the three-dimensional coordinates of the three-dimensional points, by performing bundle adjustment. The model generation unit 11 calculates the three-dimensional coordinates of multiple three-dimensional points corresponding to multiple feature points. In this case, the three-dimensional coordinates represent the coordinates in the three-dimensional coordinate system CS set in the virtual space VS. The three-dimensional coordinate system CS is defined by mutually orthogonal X, Y, and Z axes. The three-dimensional coordinate system CS may be a coordinate system with a predetermined position in the virtual space VS as its origin, or it may be a coordinate system to which geospatial coordinates (ground coordinates) or actual size information is assigned.

[0044] Multiple three-dimensional points constitute the point cloud 321. That is, each three-dimensional point is one of the points 322 that make up the point cloud 321. In this manner, the model generation unit 11 generates point cloud data. In this case, for example, the point cloud data represents a sparse point cloud 321. Since the three-dimensional model 301 is constructed from the point cloud data, the generation of point cloud data is synonymous with the generation of the three-dimensional model 301.

[0045] The model generation unit 11 may add one or more pieces of information from color information, brightness information, normal vector information, and reflection intensity information to each point 322 that constitutes the point cloud 321. In addition to SfM processing, the model generation unit 11 may also perform MVS (Multi View Stereo) processing. MVS processing involves calculating the depth and normal for each pixel of each two-dimensional image 150 using multi-view stereo measurement, integrating these, and generating a dense point cloud 321 of the subject 300. Therefore, by performing MVS processing, the model generation unit 11 can generate point cloud data that represents a dense point cloud 321. In this case, the three-dimensional model 301 is constructed from the dense point cloud 321.

[0046] As described above, in the process of generating the three-dimensional model 301, the model generation unit 11 estimates the imaging position 100 and the orientation of the camera CM for each two-dimensional image 150.

[0047] The imaging position 100 indicates the position of the camera CM when the subject 300 is imaged and a two-dimensional image 150 is generated. The estimated result of the imaging position 100 (hereinafter referred to as "imaging position information 521") is shown by three-dimensional coordinates in the three-dimensional coordinate system CS (hereinafter referred to as "imaging three-dimensional coordinates"). As shown in Figure 4, in this embodiment, the imaging position 100 in the real space RS is shown as "imaging position 101" in the virtual space VS. Note that imaging position information 521 is estimated for each two-dimensional image 150.

[0048] The camera CM's orientation indicates the camera CM's orientation (direction) when the subject 300 is imaged and a two-dimensional image 150 is generated. The estimated camera CM orientation (hereinafter referred to as "imaging orientation information 522") is shown in the three-dimensional coordinate system CS by the rotation angle of the camera CM around the X axis, the rotation angle of the camera CM around the Y axis, and the rotation angle of the camera CM around the Z axis. In Figure 4, for the sake of simplification, the imaging orientation information 522 is shown by an arrow. Note that the imaging orientation information 522 is estimated for each two-dimensional image 150. The model generation unit 11 may also perform processing to reduce the effect of gimbal lock when estimating the camera CM's orientation. Gimbal lock is a phenomenon in which the three degrees of freedom of rotation become two degrees of freedom when representing a three-dimensional orientation using Euler angles.

[0049] Furthermore, it is preferable for the model generation unit 11 to assign faces to the three-dimensional model 301 based on the point cloud data. One example of the process of assigning faces is the process of converting the point cloud data constituting the three-dimensional model 301 into mesh data. In this case, for example, the model generation unit 11 generates multiple polygonal faces (e.g., triangles) by connecting points 322 of the point cloud data constituting the three-dimensional model 301, and represents the three-dimensional model 301 with these multiple polygonal faces. In this case, for example, the model generation unit 11 generates TIN (Triangulated Irregular Network) data based on the point cloud data, and represents the three-dimensional model 301 with the TIN data. As a result, the three-dimensional model 301 is represented by a set of triangular faces.

[0050] The process of assigning faces to the three-dimensional model 301 is not particularly limited and may be performed in the following ways. For example, the model generation unit 11 assigns faces to each point 322 that constitutes the point cloud data. In this case, the faces are, for example, circular faces centered on each point 322 that constitutes the point cloud data. Alternatively, for example, the model generation unit 11 assigns faces to the three-dimensional model 301 composed of point cloud data by executing an octree algorithm. In this case, specifically, the model generation unit 11 represents the three-dimensional model 301 by multiple cubes by recursively dividing the three-dimensional model 301 into eight cubes (octants). Alternatively, for example, the model generation unit 11 assigns faces to the three-dimensional model 301 by converting the point cloud data into surface data.

[0051] Furthermore, the model generation unit 11 may add material information to the faces assigned to the three-dimensional model 301 based on the two-dimensional image 150. The material information includes information on the color and pattern of the subject 300. For example, the model generation unit 11 may perform a process of mapping a texture to each face (e.g., each polygon) that constitutes the mesh data of the three-dimensional model 301. In this case, the mesh data is, for example, TIN data.

[0052] The storage unit 50 stores the three-dimensional model 301 (point cloud data constituting the three-dimensional model 301) in the second database 52. Furthermore, the storage unit 50 stores multiple two-dimensional images 150, multiple imaging position information 521, and multiple imaging orientation information 522 in the second database 52 in association with the three-dimensional model 301. In the second database 52, the two-dimensional images 150 are stored in association with the imaging position information 521 and the imaging orientation information 522.

[0053] Furthermore, the method for generating the three-dimensional model 301 is not particularly limited, as long as the imaging position information 521 and imaging orientation information 522 can be obtained. For example, the model generation unit 11 may generate the three-dimensional model 301 by 3D Gaussian splatting.

[0054] Continuing with Figure 2, the distortion correction unit 12 of the calculation unit 10 will be explained. The distortion correction unit 12 obtains the two-dimensional image 150 to be processed (hereinafter referred to as "target two-dimensional image 150A") from the second database 52, which is one of the multiple two-dimensional images 150 used when generating the three-dimensional model 301.

[0055] Furthermore, the distortion correction unit 12 acquires calibration information for the camera CM from the second database 52. Based on the calibration information, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 150A. In this case, the distortion is, for example, lens distortion. Lens distortion is, for example, aberration distortion.

[0056] Next, the image placement unit 13 of the calculation unit 10 will be described with reference to Figures 2, 5, and 6(a). Figures 5 and 6(a) to 6(d) show the camera CM for ease of understanding.

[0057] Figure 5 is a schematic perspective view showing the imaging position 101, the target two-dimensional image 151, and the three-dimensional model 301 in the virtual space VS. As shown in Figure 5, in this embodiment, the target two-dimensional image 150A placed in the virtual space VS is referred to as "target two-dimensional image 151". Furthermore, the subject image 200 and target image 210 of the target two-dimensional image 150A placed in the virtual space VS are referred to as "subject image 201" and "target image 211", respectively.

[0058] Figure 6(a) is a schematic side view showing the imaging position 101 represented in the virtual space VS. Figure 6(b) is a schematic side view showing the target two-dimensional image 151 placed in the virtual space VS.

[0059] First, as shown in Figure 2, the image placement unit 13 acquires imaging position information 521 and imaging orientation information 522 associated with the target two-dimensional image 150A from the second database 52. The imaging position information 521 indicates the imaging three-dimensional coordinates (aX, aY, aZ). As shown in Figures 5 and 6(a), the imaging position 101 in the virtual space VS is indicated by the imaging three-dimensional coordinates (aX, aY, aZ).

[0060] Next, as shown in Figures 5 and 6(b), the image placement unit 13 places the distortion-corrected target two-dimensional image 151 in the virtual space VS based on the imaging position information 521 and the imaging orientation information 522.

[0061] Specifically, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS at a predetermined distance L away from the imaging three-dimensional coordinates (aX, aY, aZ) along the optical axis 523. In this case, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS such that the target two-dimensional image 151 is orthogonal to the optical axis 523 of the camera CM, and the size of the target two-dimensional image 151 corresponds to the field of view θ of the camera CM.

[0062] In this case, the image placement unit 13 calculates the three-dimensional coordinates (bX, bY, bZ), (cX, cY, cZ), and (dX, dY, dZ) in the virtual space VS where at least three of the four corners 161, 162, 163, and 164 of the target two-dimensional image 151 are placed. The image placement unit 13 then places the target two-dimensional image 151 in the virtual space VS by arranging at least three of the four corners 161 to 163 of the target two-dimensional image 151 at the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ).

[0063] Next, the target detection unit 14 of the calculation unit 10 will be described with reference to Figures 2, 5, 6(c), and 7. Figure 6(c) is a schematic side view showing the target image 211 and bounding box 221 included in the target two-dimensional image 151 placed in the virtual space VS, as well as the target region 311 included in the three-dimensional model 301. Figure 7 is a front view showing the target two-dimensional image 151 placed in the virtual space VS. In Figure 7, the target two-dimensional image 151 is viewed from the imaging position 101.

[0064] As shown in Figure 2, preferably, the target detection unit 14 utilizes a trained model 53 constructed by training on training data. The trained model 53 is constructed by training on training data. In this case, training is supervised learning. As an example, the trained model 53 uses YOLO (You Only Look Once), an object detection algorithm that utilizes a convolutional neural network. The object detection algorithm is not particularly limited and may be, for example, R-CNN (Regions with Convolutional Neural Network) or SSD (Single Shot Multi-Box Detector). The training data includes a two-dimensional image containing the target image and an information tag. The information tag is attached to the two-dimensional image. The information tag includes a bounding box surrounding the target image contained in the subject image in the two-dimensional image. The information tag may further include the name of the target image. The bounding box corresponds to an example of "information identifying the target image" in this disclosure.

[0065] The object detection algorithm may also be a segmentation algorithm. In this case, the object detection algorithm may be, for example, "U-NET" which performs semantic segmentation, "Mask R-CNN" which performs instant segmentation, or "Panoptic Feature Pyramid Network" which performs panoptic segmentation.

[0066] As shown in Figures 2, 5, 6(c), and 7, the target detection unit 14, as an example, causes the trained model 53 to detect the target image 211 contained in the target two-dimensional image 151 and sets a bounding box 221 surrounding the target image 211. In this way, the target detection unit 14 causes the trained model 53 to detect the target image 211 and identifies the target image 211. In other words, when the target detection unit 14 inputs the target two-dimensional image 151 to the trained model 53, the trained model 53 outputs the target two-dimensional image 151 including the target image 211 surrounded by the bounding box 221. Note that in Figure 6(c), the bounding box 221 is exaggerated for easier viewing.

[0067] In this case, the target detection unit 14 uses a pre-trained model 53, but the target image 211 may also be detected using, for example, a rule-based algorithm.

[0068] Next, the target coordinate calculation unit 15 will be described with reference to Figures 2, 5, 6(c), 6(d), and 7. Figure 6(d) is a schematic side view showing the state in the virtual space VS where a straight line 400 passing through the imaging position 101 and the target image 211 intersects with the three-dimensional model 301.

[0069] As shown in Figures 2, 5, 6(c), and 7, the target coordinate calculation unit 15 calculates the three-dimensional coordinates (eX, eY, eZ) of the target image 211 contained in the target two-dimensional image 151 in the virtual space VS. Hereinafter, the three-dimensional coordinates (eX, eY, eZ) will be referred to as the "first target three-dimensional coordinates (eX, eY, eZ)".

[0070] Specifically, the target coordinate calculation unit 15 calculates the first target three-dimensional coordinate (eX, eY, eZ) based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.

[0071] Preferably, as shown in Figure 7, the first target three-dimensional coordinates (eX, eY, eZ) are indicated by the three-dimensional coordinates in the virtual space VS of a specific point 224 in the bounding box 221. For example, the specific point 224 is the intersection of diagonals 222 and 223 of the bounding box 221. Thus, the specific point 224 is determined based on the bounding box 221. The specific point 224 represents a point inside the target image 211. The bounding box 221 is a rectangle surrounding the target image 211.

[0072] In this case, specifically, the target coordinate calculation unit 15 calculates the two-dimensional coordinates (u, v) of a specific point 224 in the bounding box 221 of the target two-dimensional image 151. The two-dimensional coordinates (u, v) are coordinates in the screen coordinate system SC. The screen coordinate system SC is defined by mutually orthogonal U and V axes. The origin of the screen coordinate system SC is not particularly limited, but for example, it is set at a corner 161 of the target two-dimensional image 151. The origin may also be set at the center of the target two-dimensional image 151.

[0073] Then, the target coordinate calculation unit 15 converts the two-dimensional coordinates (u, v) of a specific point 224 in the screen coordinate system SC into the three-dimensional coordinates (eX, eY, eZ) of the first target three-dimensional coordinates (eX, eY, eZ) based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.

[0074] Although the specific point 224 was determined based on the bounding box 221, it is not particularly limited as long as it is a point inside the target image 211.

[0075] Next, as shown in Figures 5 and 6(d), the target coordinate calculation unit 15 calculates the three-dimensional coordinates (fX, fY, fZ) of the position where the straight line 400, passing through the imaging three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ), intersects with the three-dimensional model 301. In this case, the three-dimensional coordinates (fX, fY, fZ) are referred to as the "second target three-dimensional coordinates (fX, fY, fZ)". In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point when the straight line 400 intersects with the three-dimensional model 301. The straight line 400 is, for example, a virtual straight line.

[0076] The second target three-dimensional coordinates (fX, fY, fZ) indicate the three-dimensional coordinates of the target region 311 on the three-dimensional model 301. The target region 311 corresponds to the target image 211 and is the region on the three-dimensional model 301 that reproduces the target 310 in the subject 300.

[0077] Specifically, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ), which are the positions where the line 400 intersects with the surface of the three-dimensional model 301. In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point where the line 400 intersects with the surface of the three-dimensional model 301.

[0078] For example, if a three-dimensional model 301 is assigned multiple polygonal faces (e.g., triangles) with point 322 of the point cloud 321 as its vertex, the target coordinate calculation unit 15 sets the average of the three-dimensional coordinates of the multiple vertices that define the face to which the line 400 intersects as the second target three-dimensional coordinate. Alternatively, for example, the target coordinate calculation unit 15 sets the three-dimensional coordinate of the vertex closest to the point to which the line 400 intersects as the second target three-dimensional coordinate.

[0079] For example, if a circular surface is assigned to the three-dimensional model 301, with each point 322 constituting the point cloud 321 as its center, the target coordinate calculation unit 15 sets the three-dimensional coordinate of point 322 at the center of the circular surface where the straight line 400 intersects as the second target three-dimensional coordinate.

[0080] For example, if multiple cubes are assigned to a three-dimensional model 301 using an octree algorithm, the target coordinate calculation unit 15 sets the three-dimensional coordinates of point 322, which is contained within the cube having the face that the line 400 intersects, as the second target three-dimensional coordinates. If multiple points 322 are contained within the cube, the average of the three-dimensional coordinates of the multiple points 322 is set as the second target three-dimensional coordinates.

[0081] As described above with reference to Figures 5 and 6, according to this embodiment, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ), which are the three-dimensional coordinates of the target region 311 on the three-dimensional model 301, by determining the position where a straight line 400 passing through the imaging three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ) of the target two-dimensional image 151 intersects with the three-dimensional model 301 in the virtual space VS. Thus, in this embodiment, instead of performing a process to search for the target region over the entire area of ​​the three-dimensional model, the position of the target region 311 on the three-dimensional model 301 is determined based on the imaging three-dimensional coordinates (aX, aY, aZ) and one target two-dimensional image 151. In other words, the position of the target region 311 on the three-dimensional model 301 is determined by processing a two-dimensional image (target two-dimensional image 151). Therefore, the position of the target region 311 on the three-dimensional model 301 can be determined by a simple process.

[0082] In addition, according to this embodiment, the calculation unit 10 automatically identifies the position of the target region 311 on the three-dimensional model 301. Therefore, the manual task of visually searching for the target region 311 on the three-dimensional model 301 can be omitted. Thus, the workload on humans can be reduced.

[0083] In particular, the imaging position information 521 (imaging three-dimensional coordinates (aX, aY, aZ)) and imaging orientation information 522 are information that is always calculated as a by-product when generating the three-dimensional model 301. Therefore, it is possible to suppress the occurrence of additional processing required to identify the target region 311 on the three-dimensional model 301.

[0084] Furthermore, according to this embodiment, the model generation unit 11 assigns surfaces to the three-dimensional model 301. Therefore, the straight line 400 always intersects with the three-dimensional model 301. Thus, the second target three-dimensional coordinates (fX, fY, fZ) indicating the position of the target region 311 can be reliably calculated without depending on the density of the point cloud 321.

[0085] Furthermore, according to this embodiment, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 151. Therefore, even if the target image 211 is located in a region of the target two-dimensional image 151 where distortion occurs, the two-dimensional coordinates (u, v) indicating the position of the target image 211 and the first target three-dimensional coordinates (eX, eY, eZ) can be calculated with high accuracy.

[0086] Furthermore, according to this embodiment, the target detection unit 14 inputs the target two-dimensional image 151 into the trained model 53 to detect the target image 211. In other words, the trained model 53 is constructed by training the two-dimensional image 150 as training data. Therefore, compared to the case where a trained model is constructed to detect the target image by training a three-dimensional model, the burden on the worker when creating the training data, the volume of training data, and the amount of computation required for training can be reduced. As a result, the cost and training time when constructing the trained model 53 can be reduced. In particular, instead of detecting the target region by analyzing a three-dimensional model, the target image 211 is detected on the two-dimensional image (target two-dimensional image 151). Therefore, according to this embodiment, compared to the case where the target region is detected by analyzing a three-dimensional model, the target region 311 on the three-dimensional model 301 can be identified with faster and simpler processing.

[0087] Furthermore, according to this embodiment, the first target three-dimensional coordinates (eX, eY, eZ) are the three-dimensional coordinates of a specific point 224 determined based on the bounding box 221. In other words, the specific point 224 is determined as the intersection of the diagonals 222 and 223 of the bounding box 221. Therefore, compared to the case where the specific point 224 is determined as any point in the target image 211, the specific point 224 can be easily determined with a simpler process.

[0088] Next, an image analysis method according to this embodiment will be described with reference to Figures 2 and 8. Figure 8 is a flowchart of the image analysis method. The image analysis method is executed by Server 1. As shown in Figure 8, the image analysis method includes steps S1 to S9. The computer program stored in the storage unit 50 causes the arithmetic unit 10 to execute steps S1 to S9. In other words, the computer program product realizes steps S1 to S9 when the computer program is executed by the arithmetic unit 10. The arithmetic unit 10 corresponds to an example of a "computer" in this disclosure.

[0089] As shown in Figures 2 and 8, first, in step S1, the model generation unit 11 generates a three-dimensional model 301 of the subject 300 based on a plurality of two-dimensional images 150 generated by capturing the subject 300 from a plurality of different imaging positions 100.

[0090] Next, in step S2, the distortion correction unit 12 acquires the target two-dimensional image 150A from the second database 52 among the multiple two-dimensional images 150.

[0091] Next, in step S3, the distortion correction unit 12 corrects the distortion of the target two-dimensional image 150A based on the calibration information. The distortion-corrected target two-dimensional image 150A is stored in the storage unit 50.

[0092] Next, in step S4, the image placement unit 13 obtains imaging position information 521 and imaging orientation information 522 from the second database 52. The imaging position information 521 indicates the position of the camera CM at the time of imaging using the imaging three-dimensional coordinates (aX, aY, aZ) in the three-dimensional coordinate system CS. The imaging orientation information 522 indicates the orientation of the camera CM at the time of imaging using the rotation angle in the three-dimensional coordinate system CS.

[0093] Next, in step S5, the image placement unit 13 places the distortion-corrected target two-dimensional image 150A in the virtual space VS based on the imaging position information 521 and the imaging orientation information 522. The target two-dimensional image 150A placed in the virtual space VS will be referred to as "target two-dimensional image 151".

[0094] Next, in step S6, the target detection unit 14 instructs the trained model 53 to detect the target image 211 included in the target two-dimensional image 151 and sets a bounding box 221 surrounding the target image 211.

[0095] Next, in step S7, the target coordinate calculation unit 15 calculates the first target three-dimensional coordinates (eX, eY, eZ) which represent the three-dimensional coordinates of a specific point 224 in the bounding box 221.

[0096] Next, in step S8, the target coordinate calculation unit 15 calculates a second target three-dimensional coordinate (fX, fY, fZ) which indicates the position where a straight line 400 passing through the imaging three-dimensional coordinate (aX, aY, aZ) indicated by the imaging position information 521 and the first target three-dimensional coordinate (eX, eY, eZ) intersects with the three-dimensional model 301. The second target three-dimensional coordinate (fX, fY, fZ) indicates the three-dimensional coordinate of the target region 311 on the three-dimensional model 301. In this way, the position of the target 310 included in the subject 300 is identified as the target region 311 on the three-dimensional model 301.

[0097] Next, in step S9, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 via the communication unit 40 in response to a request from the user's terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in an operable manner. For example, the display control unit 16 overlays the target two-dimensional image 151 onto the three-dimensional model 301 and displays it on the terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in a position where the target region 311 on the three-dimensional model 301 is displayed on the terminal 2. In this case, for example, the display control unit 16 enlarges the target region 311 on the three-dimensional model 301 and displays it on the terminal 2. Therefore, according to this embodiment, the user of terminal 2 can omit the task of visually searching for the target region 311 on the three-dimensional model 301. In other words, the workload of the user can be reduced. When step S9 is completed, the image analysis method is finished.

[0098] As explained above with reference to Figure 8, according to the image analysis method of this embodiment, the position of the target region 311 on the three-dimensional model 301 is identified by calculating the position where the line 400 intersects with the three-dimensional model 301 in the virtual space VS. Therefore, the position of the target region 311 can be identified by a simple process.

[0099] Figure 9 is a flowchart showing the three-dimensional model generation process in step S1 of Figure 8. As shown in Figure 9, the three-dimensional model generation process includes steps S11 to S15.

[0100] First, in step S11, the model generation unit 11 acquires multiple two-dimensional images 150, including the subject image 200, from the first database 51.

[0101] Next, in step S12, the model generation unit 11 extracts feature points of the subject 300 from each two-dimensional image 150 and associates the feature points with each other in the two-dimensional images 150.

[0102] Next, in step S13, the model generation unit 11 estimates the imaging position 100 and the orientation of the camera CM when each two-dimensional image 150 was generated, based on the associated feature points. Then, the model generation unit 11 calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimated result of the imaging position 100 (imaging position information 521) and the estimated result of the orientation of the camera CM (imaging orientation information 522). As a result, point cloud data representing a point cloud 321 consisting of multiple three-dimensional points (multiple points 322) is generated.

[0103] Next, in step S14, the model generation unit 11 adds surfaces to the three-dimensional model 301 based on the point cloud data.

[0104] Next, in step S15, the model generation unit 11 adds material information to the surfaces applied to the three-dimensional model 301 based on the two-dimensional image 150. Then, the process returns to step S2 in Figure 8.

[0105] (modified version) A modified version of this embodiment will be described with reference to Figures 2 and 10. In the modified version, the target 310 included in the subject 300 is the part of the subject 300 that indicates a specific state. The following will mainly describe the differences between the modified version and the above embodiment.

[0106] In the modified example, the target image 211 shown in Figure 7 is an image of a portion of the subject 300 that represents a specific state (hereinafter referred to as the "specific state portion"). Therefore, according to the modified example, the specific state portion of the subject 300 is identified as the target region 311 on the three-dimensional model 301. In other words, the position of the target region 311 representing the specific state portion of the subject 300 is calculated as the second target three-dimensional coordinate on the three-dimensional model 301. As a result, the manual work of visually searching for the target region 311 representing the specific state portion on the three-dimensional model 301 can be omitted. Furthermore, the possibility of missing the detection of the target region 311 representing the specific state portion can be suppressed.

[0107] The specific state portion of the subject 300 is, for example, the portion of the subject 300, either entirely or within the area of ​​interest, that exhibits an abnormal state (hereinafter referred to as the "abnormal state portion"). Therefore, in this case, the target image 211 is an image of the abnormal state portion. As a result, the manual task of visually searching for the target region 311 representing the abnormal state portion on the three-dimensional model 301 can be omitted. Furthermore, the failure to detect the target region 311 representing the abnormal state portion can be suppressed.

[0108] For example, the target image 211 shows a region of the target two-dimensional image 151 that has changed relative to the reference two-dimensional image. The region that has changed relative to the reference two-dimensional image is, for example, a region that has been deformed, mutated, displaced, increased, decreased, or disappeared relative to the reference two-dimensional image. The reference two-dimensional image is, for example, a two-dimensional image of the subject 300 that was captured earlier than the target two-dimensional image 151. Alternatively, the reference two-dimensional image is, for example, a two-dimensional image generated by capturing the subject 300 in a reference state. The reference state is, for example, a state in which the entire subject 300 or the region of interest is in a normal state. In this example, the target image 211 is an image of the entire subject 300 or a portion of the region of interest that shows an abnormal state.

[0109] Next, an image analysis method relating to a modified example will be described with reference to Figures 2 and 10. Figure 10 is a flowchart of the image analysis method relating to a modified example. The image analysis method is executed by Server 1. As shown in Figure 10, the image analysis method includes steps S21 to S23. The computer program stored in the storage unit 50 causes the arithmetic unit 10 to execute steps S21 to S23. In other words, the computer program product realizes steps S21 to S23 when the computer program is executed by the arithmetic unit 10.

[0110] First, in step S21, the target detection unit 14 performs a target image detection process. The target image detection process involves detecting a target image 210 from the video data 511 of the subject 300, the two-dimensional image 150 extracted from the video data 511, the still image dataset 512, or the two-dimensional image 150 extracted from the still image dataset 512. Hereinafter, the video data 511, the two-dimensional image 150 extracted from the video data 511, the still image dataset 512, and the still image dataset 512 will be collectively referred to as "detection target image data".

[0111] For example, the target detection unit 14 inputs the image data to be detected into the trained model. If the trained model detects the target image 210 from the image data to be detected, it outputs information indicating that the image data to be detected contains the target image 210. If the trained model does not detect the target image 210 from the image data to be detected, it outputs information indicating that the image data to be detected does not contain the target image 210.

[0112] A trained model is a computer program built by learning from training data. Examples of trained models include neural networks. Training can be, for example, supervised learning. Training data consists of image data to be detected and information indicating that the image contains the target image.

[0113] Next, in step S22, the target detection unit 14 determines whether or not the target image 210 has been detected in the target image detection process.

[0114] If a negative result is obtained in step S22 (NO), the image analysis method is terminated.

[0115] On the other hand, if a positive result is obtained in step S22 (YES), the process proceeds to step S23.

[0116] Next, in step S23, the calculation unit 10 identifies the target region 311 on the three-dimensional model 301 corresponding to the target image 210. The processing in step S23 is the same as the processing in steps S1 to S9 shown in Figure 8. When step S23 is completed, the image analysis method is finished.

[0117] As explained above with reference to Figure 10, in this modified version, the process in step S23 is executed only when the target image 210 is detected. Therefore, the processing load on the calculation unit 10 can be reduced. In particular, if the target image 210 is not detected, the three-dimensional model 301 is not generated, thus reducing the processing load on the calculation unit 10. In this case, the increase in the amount of data stored in the storage unit 50 can also be suppressed.

[0118] Furthermore, the image analysis method shown in Figure 10 can also be applied to the above-described embodiment with reference to Figures 1 to 9.

[0119] Preferred embodiments and modifications of the present disclosure have been described in detail above with reference to the attached drawings, but the technical scope of the present disclosure is not limited to such examples. It is clear to any person with ordinary skill in the art of the present disclosure that various modifications or alterations can be conceived within the scope of the technical idea set forth in the claims, and these too are understood to fall within the technical scope of the present disclosure.

[0120] The apparatus or system described herein may be implemented as a single apparatus, or as a group of apparatuses (e.g., a cloud server) partially or entirely connected by a network. For example, some or all of the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by the same computer or server. Alternatively, for example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by terminal 2. For example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may each be implemented by separate computers or servers. For example, the first database 51, the second database 52, and the trained model 53 may each be stored in separate storage devices or servers.

[0121] The series of processes performed by the apparatus described herein may be implemented using software, hardware, or a combination of software and hardware. Computer programs for implementing each function of the arithmetic unit 10 according to this embodiment can be created and implemented on a PC or the like. Furthermore, a computer-readable recording medium containing such a computer program can also be provided. Examples of recording media include magnetic disks, optical disks, magneto-optical disks, and flash memory. Alternatively, the computer program may be distributed without using a recording medium, for example, via a network.

[0122] Furthermore, the processes described using flowcharts in this specification do not necessarily have to be executed in the order shown. Some processing steps may be executed in parallel. Additional processing steps may be adopted, and some processing steps may be omitted.

[0123] Furthermore, the effects described herein are merely descriptive or illustrative and not limiting. In other words, the technology relating to this disclosure may produce other effects that will be apparent to those skilled in the art from the description herein, in addition to or in lieu of the effects described herein.

[0124] Furthermore, the following configurations also fall within the technical scope of this disclosure.

[0125] (Item 1) A model generation unit generates a three-dimensional model of the subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, An image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging. The system includes a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model, The first target three-dimensional coordinates indicate the position of the target image contained in the target two-dimensional image in the virtual space, according to the image analysis device.

[0126] (Item 2) The aforementioned three-dimensional model is constructed from point cloud data. The model generation unit assigns surfaces to the three-dimensional model based on the point cloud data. The image analysis apparatus according to item 1, wherein the target coordinate calculation unit calculates the second target three-dimensional coordinate, which is the position where the straight line abuts the surface.

[0127] (Item 3) The system further includes a distortion correction unit for correcting the distortion of the aforementioned two-dimensional image. The image analysis apparatus according to item 1 or item 2, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the target two-dimensional image from which the distortion has been corrected.

[0128] (Item 4) The system further includes a target detection unit that uses a trained model constructed by training on training data to detect the target image contained in the target two-dimensional image and identify the target image, The aforementioned training data includes a two-dimensional image containing a target image and information tags, The aforementioned information tag includes information that identifies the target image, and is an image analysis device as described in any of items 1 to 3.

[0129] (Item 5) The target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image. The image analysis apparatus according to item 4, wherein the first target three-dimensional coordinates indicate the three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.

[0130] (Item 6) The image analysis device described in any of items 1 to 5, wherein the target image is an image of a part of the subject that shows a specific state.

[0131] (Item 7) A step of generating a three-dimensional model of a subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, The steps include obtaining three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging pose information indicating the pose of the camera that performed the imaging, The steps include: arranging the target two-dimensional image in the virtual space based on the three-dimensional imaging coordinates and the imaging orientation information; The steps include: calculating the first target three-dimensional coordinates, which are the three-dimensional coordinates of the target image included in the aforementioned two-dimensional image in the virtual space; An image analysis method comprising the step of calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.

[0132] (Item 8) On the computer, A step of generating a three-dimensional model of a subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, The steps include obtaining three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging pose information indicating the pose of the camera that performed the imaging, The steps include: arranging the target two-dimensional image in the virtual space based on the three-dimensional imaging coordinates and the imaging orientation information; The steps include: calculating the first target three-dimensional coordinates, which are the three-dimensional coordinates of the target image included in the aforementioned two-dimensional image in the virtual space; A computer program that performs the step of calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model. [Industrial applicability]

[0133] This disclosure provides an image analysis device, an image analysis method, and a computer program, and has industrial applicability. [Explanation of Symbols]

[0134] 1 Server (image analysis device), 11 Model generation unit, 12 Distortion correction unit, 13 Image placement unit, 14 Target detection unit, 15 Target coordinate calculation unit, 16 Display control unit, 53 Trained model

Claims

1. A model generation unit generates a three-dimensional model of the subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, An image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging posture information indicating the posture of the camera that performed the imaging. The system includes a target coordinate calculation unit that calculates a second target three-dimensional coordinate indicating the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model. The first target three-dimensional coordinates indicate the position of the target image contained in the target two-dimensional image in the virtual space, according to the image analysis device.

2. The aforementioned three-dimensional model is constructed from point cloud data. The model generation unit assigns surfaces to the three-dimensional model based on the point cloud data. The image analysis apparatus according to claim 1, wherein the target coordinate calculation unit calculates the second target three-dimensional coordinate, which is the position where the straight line abuts the surface.

3. The system further includes a distortion correction unit for correcting the distortion of the aforementioned two-dimensional image. The image analysis apparatus according to claim 1 or claim 2, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the target two-dimensional image from which the distortion has been corrected.

4. The system further includes a target detection unit that uses a trained model constructed by training on training data to detect the target image contained in the target two-dimensional image and identify the target image, The aforementioned training data includes a two-dimensional image containing a target image and information tags, The image analysis apparatus according to claim 1 or claim 2, wherein the information tag includes information that identifies the target image.

5. The target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image. The image analysis apparatus according to claim 4, wherein the first target three-dimensional coordinates indicate the three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.

6. The image analysis apparatus according to claim 1 or claim 2, wherein the target image is an image of a portion of the subject that shows a specific state.

7. A step of generating a three-dimensional model of a subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, The steps include obtaining three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging pose information indicating the pose of the camera that performed the imaging, The steps include: arranging the target two-dimensional image in the virtual space based on the three-dimensional imaging coordinates and the imaging orientation information; The steps include: calculating the first target three-dimensional coordinates, which are the three-dimensional coordinates in the virtual space of the target image included in the aforementioned two-dimensional image; An image analysis method comprising the step of calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.

8. On the computer, A step of generating a three-dimensional model of a subject based on multiple two-dimensional images generated by capturing the subject from multiple different imaging positions, The steps include obtaining three-dimensional coordinates indicating the imaging position when the target two-dimensional image was generated from the plurality of two-dimensional images, which are imaging three-dimensional coordinates indicated by the three-dimensional coordinates in the virtual space where the three-dimensional model is placed, and imaging pose information indicating the pose of the camera that performed the imaging, The steps include: arranging the target two-dimensional image in the virtual space based on the three-dimensional imaging coordinates and the imaging orientation information; The steps include: calculating the first target three-dimensional coordinates, which are the three-dimensional coordinates in the virtual space of the target image included in the aforementioned two-dimensional image; A computer program that performs the step of calculating a second target three-dimensional coordinate, which is the three-dimensional coordinate of the position where a straight line passing through the imaging three-dimensional coordinate and the first target three-dimensional coordinate intersects the three-dimensional model.

Citation Information

Patent Citations

  • Deterioration diagnosis system using flight vehicle

    JP2019070631A