Image analysis device, image analysis method, and computer program
The image analysis device simplifies the identification of target regions on three-dimensional models by generating models from multiple images and calculating intersection points, reducing manual effort and enhancing efficiency.
Patent Information
- Application Number
- JP2024573080
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2044-09-30
AI Technical Summary
Identifying the position of a target region in a three-dimensional model is difficult and time-consuming, requiring extensive manual searching.
An image analysis device and method that generates a three-dimensional model from multiple two-dimensional images, places the target image in a virtual space using imaging coordinates and attitude information, and calculates the position where a straight line from the imaging coordinates intersects with the model, allowing for simple and automated target area identification.
Enables efficient and automated identification of target areas on three-dimensional models by processing two-dimensional images, reducing human workload and avoiding extensive searches over the entire model.
Smart Images

Figure 0007800961000001 
Figure 0007800961000002 
Figure 0007800961000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image analysis device, an image analysis method, and a computer program. [Background technology]
[0002] The diagnostic system described in Patent Document 1 has the function of automatically determining and detecting deteriorated areas by performing diagnostic processing using a computer based on continuous images of past and present objects taken from the air by an aircraft. The objects are buildings, infrastructure facilities, etc. A predetermined area on the surface of the object is the diagnostic target area. The diagnostic target area of the object is set as a basic setting based on input operations by the user.
[0003] The diagnostic system described in Patent Document 1 includes an SFM processing unit that performs SFM processing on a plurality of input images to reconstruct a three-dimensional structure and output a three-dimensional model. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-070631 Summary of the Invention [Problem to be solved by the invention]
[0005] Generally, a diagnostic target region may be searched for from a three-dimensional model of an object and diagnosed.
[0006] However, it is often difficult to search for a target region, such as a diagnostic region, in a three-dimensional model and identify the position of the target region on the three-dimensional model.
[0007] Therefore, the present disclosure has been made in consideration of the above-mentioned problems, and its purpose is to provide an image analysis device, an image analysis method, and a computer program that can identify the position of a target area on a three-dimensional model through simple processing. [Means for solving the problem]
[0008] According to the present disclosure, there is provided an image analysis device comprising: a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in a virtual space based on imaging three-dimensional coordinates that indicate the imaging positions when a target two-dimensional image of the plurality of two-dimensional images was generated, the imaging three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information that indicates the attitude of a camera that performed the imaging; and a target coordinate calculation unit that calculates second target three-dimensional coordinates that indicate the position at which a straight line passing through the imaging three-dimensional coordinates and first target three-dimensional coordinates hits the three-dimensional model, wherein the first target three-dimensional coordinates indicate the position in the virtual space of a target image included in the target two-dimensional image.
[0009] Furthermore, according to the present disclosure, there is provided an image analysis method including the steps of: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions; acquiring imaging three-dimensional coordinates indicating the imaging positions when a target two-dimensional image of the plurality of two-dimensional images was generated, the imaging three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of a camera that performed the imaging; arranging the target two-dimensional image in the virtual space based on the imaging three-dimensional coordinates and the imaging attitude information; calculating first target three-dimensional coordinates, which are the three-dimensional coordinates in the virtual space of a target image included in the target two-dimensional image; and calculating second target three-dimensional coordinates, which are the three-dimensional coordinates of a position where a straight line passing through the imaging three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model.
[0010] Furthermore, according to the present disclosure, there is provided a computer program that causes a computer to execute the steps of: generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different capturing positions; acquiring capturing three-dimensional coordinates that indicate the capturing positions when a two-dimensional image of the subject from the plurality of two-dimensional images was generated, the capturing three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and capturing attitude information that indicates the attitude of a camera that performed the capturing; arranging the two-dimensional image of the subject in the virtual space based on the capturing three-dimensional coordinates and the capturing attitude information; calculating first target three-dimensional coordinates that are the three-dimensional coordinates in the virtual space of a target image included in the two-dimensional image of the subject; and calculating second target three-dimensional coordinates that are the three-dimensional coordinates of a position where a straight line passing through the capturing three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model. [Effects of the Invention]
[0011] According to the present disclosure, it is possible to provide an image analysis device, an image analysis method, and a computer program that can identify the position of a target area on a three-dimensional model through simple processing. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a block diagram showing an example of the configuration of an image analysis system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of a server according to the embodiment. [Figure 3] FIG. 2 is a perspective view schematically showing a camera, a two-dimensional image, and a subject in a real space according to the embodiment. [Figure 4] 10 is a perspective view schematically showing an imaging position and a three-dimensional model of a subject in a virtual space according to the embodiment. FIG. [Figure 5]10 is a perspective view schematically showing an imaging position, a target two-dimensional image, and a three-dimensional model in a virtual space according to the embodiment. FIG. [Figure 6] (a) is a side view schematically showing an imaging position represented in a virtual space according to the embodiment; (b) is a side view schematically showing a two-dimensional image of a target placed in a virtual space; (c) is a side view schematically showing a target image and a bounding box included in the two-dimensional image of a target placed in a virtual space, and a target area included in a three-dimensional model; and (d) is a side view schematically showing a state in which a straight line passing through the imaging position and the target image hits a three-dimensional model in the virtual space. [Figure 7] FIG. 2 is a front view showing a target two-dimensional image arranged in a virtual space according to the embodiment. [Figure 8] 10 is a flowchart showing an image analysis method according to the embodiment. [Figure 9] 10 is a flowchart showing a three-dimensional model generation process according to the embodiment. [Figure 10] 10 is a flowchart showing an image analysis method according to a modified example of the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0014] Fig. 1 is a diagram showing an example of the configuration of an image analysis system SYS according to an embodiment of the present invention. As shown in Fig. 1, the image analysis system SYS includes a server 1. The server 1 corresponds to an example of the "image analysis device" of the present disclosure.
[0015] The server 1 generates a three-dimensional model of the subject to be placed in a virtual space based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions. The server 1 calculates the position in the virtual space where a line passing through the imaging position and the target image included in one two-dimensional image strikes the three-dimensional model (i.e., the position of the target area on the three-dimensional model). As a result, the position of the target area is identified on the three-dimensional model. In this way, in this embodiment, instead of performing a process to search for the target area over the entire area of the three-dimensional model, the server 1 calculates the position of the target area on the three-dimensional model from the camera position and one two-dimensional image. Therefore, the position of the target area on the three-dimensional model can be identified by a simple process. This point will be described in detail later.
[0016] The image analysis system SYS also includes at least one terminal 2, at least one mobile device 3, or at least one imaging device 5.
[0017] The server 1, the terminal 2, the mobile device 3, and the imaging device 5 are connected to a network NW. The network NW includes, for example, the Internet, a closed network, a public telephone network, a LAN (Local Area Network), and a short-distance wireless network.
[0018] The terminal 2 is, for example, a personal computer (for example, a notebook computer, a desktop computer, or a tablet).
[0019] The mobile device 3 is, for example, an unmanned mobile device or a manned mobile device. The unmanned mobile device is, for example, an unmanned aerial vehicle such as a drone, an unmanned ground vehicle, an unmanned underwater vehicle, or an unmanned surface vessel. The unmanned ground vehicle is, for example, an unmanned ground vehicle modeled after a living organism (for example, a snake-shaped unmanned ground vehicle). The manned mobile device is, for example, an aircraft, an automobile, a ship, or a submarine. The mobile device 3 includes a camera 4.
[0020] The imaging device 5 includes a camera 6. The imaging device 5 is, for example, a mobile terminal such as a smartphone. The imaging device 5 may be, for example, the camera 6 itself.
[0021] The cameras 4 and 6 capture images of the subject and generate video data representing a video including the image of the subject. A video is a collection of successive two-dimensional images. The cameras 4 and 6 may also generate multiple still image data representing multiple still images including the image of the subject. The still images are two-dimensional images.
[0022] Hereinafter, the video data will be referred to as "video data 511" (FIG. 2), and the plurality of still image data will be referred to as "still image data set 512" (FIG. 2).
[0023] Hereinafter, when there is no need to distinguish between cameras 4 and 6, cameras 4 and 6 will be collectively referred to as "camera CM."
[0024] Furthermore, the subject is not particularly limited as long as it can be captured by the camera CM, and the size, shape, pattern, and color of the subject are also not particularly limited. For example, the subject is one or more movable or immovable property. The subject is, for example, one or more objects. Typically, the subject is one or more stationary objects. The stationary object is, for example, an artificial object or a natural object. The artificial object is, for example, a structure, a machine, an electronic device, or a copyrighted work. The structure is, for example, a building or an infrastructure facility. The building is, for example, a building or a house. The infrastructure facility is a facility for establishing a social infrastructure. For example, the infrastructure facility is a road, a bridge, a road traffic facility, a power generation facility, a power distribution facility, a water treatment facility, or a gas distribution facility. The machine is, for example, an automobile, a work vehicle, a train, an aircraft, a ship, a submarine, or a robot. The natural object is, for example, a tree, a forest, the earth's surface, a cliff, a coast, or a river.
[0025] Furthermore, the subject includes one or more targets whose positions on the three-dimensional model are to be identified. The targets can be any targets, such as the entire area included in the subject, a part of an area, the entire object included in the subject, or a part of an object, without being particularly limited thereto.
[0026] The terminal 2 acquires the video data 511 or the still image data set 512 generated by the camera CM. For example, the terminal 2 receives the video data 511 or the still image data set 512 transmitted from the camera CM via the network NW. For example, the terminal 2 acquires the video data 511 or the still image data set 512 from a removable medium. The terminal 2 transmits the video data 511 or the still image data set 512 to the server 1 via the network NW. Note that the mobile device 3 and the imaging device 5 may also transmit the video data 511 or the still image data set 512 to the server 1 via the network NW.
[0027] Fig. 2 is a block diagram showing an example configuration of the server 1 in Fig. 1. As shown in Fig. 2, the server 1 includes a calculation unit 10, a communication unit 40, and a storage unit 50. The server 1 may also include an input unit 20 and a display unit 30.
[0028] The input unit 20 is an input device for inputting various pieces of information to the calculation unit 10. For example, the input unit 20 is a keyboard and pointing device, or a touch panel.
[0029] The display unit 30 displays various types of information and is, for example, a liquid crystal display or an organic electroluminescence display.
[0030] The communication unit 40 is connected to the network NW. The communication unit 40 communicates with external devices connected to the network NW. The external devices are, for example, the terminal 2, the mobile device 3, and the imaging device 5. The communication unit 40 is a communication device that communicates according to a predetermined communication protocol, and includes, for example, a network interface controller. The predetermined communication protocol is, for example, a protocol compliant with Ethernet (registered trademark) and the Internet Protocol Suite.
[0031] The communication unit 40 receives, for each subject, video data 511 or still image data set 512 generated by capturing an image of the subject from the terminal 2, the mobile device 3, and the imaging device 5 via the network NW.
[0032] The storage unit 50 includes a storage device and stores data and computer programs. The storage unit 50 includes a main storage device such as a semiconductor memory, and an auxiliary storage device such as a semiconductor memory and a hard disk drive. The storage unit 50 may also include removable media such as an optical disk. The storage unit 50 may be, for example, a non-transitory computer-readable storage medium.
[0033] The storage unit 50 stores a first database 51, a second database 52, and a trained model 53.
[0034] The first database 51 includes video data 511 and a still image data set 512 received by the communication unit 40. In the first database 51, the video data 511 and the still image data set 512 are associated with attribute information (hereinafter referred to as "attribute information AT"). The attribute information AT includes, for example, user information, imaging conditions, and subject information. The user information includes, for example, identification information of the user of the terminal 2, the mobile device 3, or the imaging device 5. The imaging conditions include, for example, the imaging date and time, and calibration information of the camera CM. The calibration information includes, for example, information on the lens distortion of the camera CM, the focal length, and information on the principal point position. The information on the lens distortion includes, for example, information on distortion aberration. The imaging conditions may include information on the imaging location. The imaging location is indicated by, for example, the position coordinates of the camera CM acquired by a GPS (Global Positioning System) or a GNSS (Global Navigation Satellite System). The subject information includes, for example, identification information of the subject. The subject information may include coordinates of ground control points (GCPs).
[0035] The second database 52 includes a plurality of data sets 520. Each of the plurality of data sets 520 includes a plurality of two-dimensional images 150, a three-dimensional model 301 (point cloud data), a plurality of pieces of imaging position information 521, and a plurality of pieces of imaging attitude information 522. In the second database 52, the two-dimensional images 150, the three-dimensional model 301, the imaging position information 521, and the imaging attitude information 522 are associated with each other and with attribute information AT. These will be described in detail later.
[0036] The trained model 53 is a computer program. The trained model 53 includes, for example, a neural network. Details of the trained model 53 will be described later.
[0037] The calculation unit 10 executes various calculations and includes processors such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).
[0038] Specifically, the calculation unit 10 includes a model generation unit 11, a distortion correction unit 12, an image placement unit 13, a target detection unit 14, a target coordinate calculation unit 15, and a display control unit 16. For example, the calculation unit 10 functions as the model generation unit 11, the distortion correction unit 12, the image placement unit 13, the target detection unit 14, the target coordinate calculation unit 15, and the display control unit 16 by executing a computer program stored in the storage unit 50.
[0039] Next, the processing executed by the calculation unit 10 will be described with reference to FIGS. 2 to 6. FIG. 3 is a perspective view schematically showing a camera CM, a two-dimensional image 150, and a subject 300 in a real space RS. As shown in FIG. 3, the subject 300 includes a target 310. The two-dimensional image 150 includes a subject image 200 representing the subject 300. The subject image 200 includes a target image 210 representing the target 310 according to the imaging position 100. In this embodiment, the subject 300 is imaged while a single camera CM is moved, and multiple two-dimensional images 150 are generated. FIG. 4 is a perspective view schematically showing an imaging position 101 and a three-dimensional model 301 of the subject 300 in a virtual space VS. In FIG. 4, the imaging position 101 is indicated by a black circle. For ease of understanding, an area A of the three-dimensional model 301 is shown enlarged.
[0040] As shown in FIGS. 2 to 4, the model generation unit 11 of the calculation unit 10 generates a three-dimensional model 301 of the subject 300 based on a plurality of two-dimensional images 150 generated by capturing images of the subject 300 from a plurality of different imaging positions 100. The three-dimensional model 301 includes a target area 311 indicating a target 310 of the subject 300. The three-dimensional model 301 is placed in a virtual space VS. The three-dimensional model 301 indicates the three-dimensional shape of the subject 300. The three-dimensional model 301 is configured by point cloud data. The point cloud data is data indicating a point cloud 321. The point cloud 321 is a collection of a plurality of points 322. The point cloud data includes three-dimensional coordinates of each point 322. The point cloud data may further include one or more of color information (e.g., RGB values) of each point 322, normal vector information of each point, and reflection intensity information.
[0041] As an example, the model generation unit 11 generates a three-dimensional model 301 by executing SfM (Structure from Motion) processing. The SfM processing refers to processing that generates the three-dimensional model 301 by utilizing the principle of triangulation based on a plurality of two-dimensional images 150 generated by capturing an object 300 having a plurality of feature points from a plurality of imaging positions 100. The SfM processing preferably includes bundle adjustment. The bundle adjustment refers to processing that minimizes reprojection errors.
[0042] Specifically, first, the model generation unit 11 acquires a plurality of two-dimensional images 150 including a subject image 200 from the video data 511 or the still image data set 512 of the first database 51. Next, the model generation unit 11 extracts feature points of the subject 300 from each of the two-dimensional images 150 and associates the feature points between the two-dimensional images 150 (feature point matching). The associated feature points are called tie points. In feature point matching, the model generation unit 11 uses, for example, SIFT (Scale Invariant Feature Transformation), SURF (Speeded-Up Robust Features), or FAST (Features from Accelerated Segment Test).
[0043] Next, the model generation unit 11 estimates the imaging position 100 (position of the camera CM) and the attitude (orientation) of the camera CM when each 2D image 150 was generated based on the associated feature points, and calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimation results. In this case, the model generation unit 11 performs bundle adjustment to optimize the estimation results of the imaging position 100 and the attitude of the camera CM, as well as the three-dimensional coordinates of the three-dimensional points. The model generation unit 11 calculates the three-dimensional coordinates of multiple three-dimensional points corresponding to the multiple feature points. In this case, the three-dimensional coordinates represent coordinates in a three-dimensional coordinate system CS set in the virtual space VS. The three-dimensional coordinate system CS is defined by mutually orthogonal X-, Y-, and Z-axes. The three-dimensional coordinate system CS may be a coordinate system with its origin at a predetermined position in the virtual space VS, or may be a coordinate system assigned geospatial coordinates (ground coordinates) or actual size information.
[0044] A plurality of three-dimensional points constitute a point cloud 321. That is, each three-dimensional point is a point 322 constituting the point cloud 321. In this manner, the model generation unit 11 generates point cloud data. In this case, for example, the point cloud data indicates a sparse point cloud 321. Because the three-dimensional model 301 is constituted by the point cloud data, generating the point cloud data is synonymous with generating the three-dimensional model 301.
[0045] The model generation unit 11 may add one or more pieces of information selected from color information, brightness information, normal vector information, and reflection intensity information to each point 322 constituting the point cloud 321. Furthermore, the model generation unit 11 may perform MVS (Multi View Stereo) processing in addition to the SfM processing. The MVS processing refers to a process of calculating depth and normals for each pixel of each two-dimensional image 150 by multi-view stereo measurement, integrating these, and generating a dense point cloud 321 of the subject 300. Therefore, the model generation unit 11 can generate point cloud data indicating the dense point cloud 321 by performing the MVS processing. In this case, the three-dimensional model 301 is configured by the dense point cloud 321.
[0046] As described above, the model generation unit 11 estimates the image capturing position 100 and the posture of the camera CM for each two-dimensional image 150 in the process of generating the three-dimensional model 301.
[0047] The imaging position 100 indicates the position of the camera CM when the subject 300 is imaged and the two-dimensional image 150 is generated. The estimation result of the imaging position 100 (hereinafter referred to as "imaging position information 521") is indicated by three-dimensional coordinates in the three-dimensional coordinate system CS (hereinafter referred to as "imaging three-dimensional coordinates"). As shown in FIG. 4, in this embodiment, the imaging position 100 in the real space RS is indicated as "imaging position 101" in the virtual space VS. Note that the imaging position information 521 is estimated for each two-dimensional image 150.
[0048] The attitude of the camera CM indicates the attitude (orientation) of the camera CM when the subject 300 is imaged and the two-dimensional image 150 is generated. The estimated result of the attitude of the camera CM (hereinafter referred to as "image capture attitude information 522") is indicated by the rotation angle of the camera CM around the X axis, the rotation angle of the camera CM around the Y axis, and the rotation angle of the camera CM around the Z axis in the three-dimensional coordinate system CS. In FIG. 4, the image capture attitude information 522 is indicated by arrows for the sake of simplicity. Note that the image capture attitude information 522 is estimated for each two-dimensional image 150. Furthermore, when estimating the attitude of the camera CM, the model generation unit 11 may perform processing to reduce the influence of gimbal lock. Gimbal lock is a phenomenon in which three degrees of freedom of rotation are reduced to two degrees of freedom when a three-dimensional attitude is expressed using Euler angles.
[0049] Furthermore, it is preferable that the model generation unit 11 assigns surfaces to the three-dimensional model 301 based on the point cloud data. One example of the surface assignment process is a process of converting the point cloud data that constitutes the three-dimensional model 301 into mesh data. In this case, for example, the model generation unit 11 generates a plurality of polygonal (e.g., triangular) surfaces (e.g., polygons) that are generated by connecting points 322 of the point cloud data that constitutes the three-dimensional model 301, and represents the three-dimensional model 301 using the plurality of polygonal surfaces. In this case, for example, the model generation unit 11 generates TIN (Triangulated Irregular Network) data based on the point cloud data, and represents the three-dimensional model 301 using the TIN data. As a result, the three-dimensional model 301 is represented by a collection of triangular surfaces.
[0050] The process of adding surfaces to the three-dimensional model 301 is not particularly limited, and may be the following process. For example, the model generation unit 11 adds surfaces to each of the points 322 constituting the point cloud data. In this case, the surfaces are, for example, circular surfaces centered on each of the points 322 constituting the point cloud data. Alternatively, for example, the model generation unit 11 adds surfaces to the three-dimensional model 301 configured from the point cloud data by executing an octree algorithm. In this case, specifically, the model generation unit 11 represents the three-dimensional model 301 using a plurality of cubes by recursively dividing the three-dimensional model 301 into eight cubes (octants). Alternatively, for example, the model generation unit 11 adds surfaces to the three-dimensional model 301 by converting the point cloud data into surface data.
[0051] Furthermore, the model generation unit 11 may add material information to the surfaces assigned to the three-dimensional model 301 based on the two-dimensional image 150. The material information is information including the color and pattern of the subject 300. For example, the model generation unit 11 may execute a process of mapping texture to each surface (e.g., each polygon) constituting the mesh data of the three-dimensional model 301. The mesh data in this case is, for example, TIN data.
[0052] The storage unit 50 stores the three-dimensional model 301 (point cloud data constituting the three-dimensional model 301) in the second database 52. Furthermore, the storage unit 50 stores a plurality of two-dimensional images 150, a plurality of pieces of imaging position information 521, and a plurality of pieces of imaging orientation information 522 in the second database 52, in association with the three-dimensional model 301. In the second database 52, the two-dimensional images 150 are stored in association with the imaging position information 521 and the imaging orientation information 522.
[0053] Note that there are no particular limitations on the method for generating the three-dimensional model 301, as long as it is possible to obtain the image capturing position information 521 and the image capturing attitude information 522. For example, the model generation unit 11 may generate the three-dimensional model 301 by 3D Gaussian splatting.
[0054] 2, the distortion correction unit 12 of the calculation unit 10 will be described. The distortion correction unit 12 acquires, from the second database 52, one of the multiple two-dimensional images 150 used to generate the three-dimensional model 301, a two-dimensional image 150 to be processed (hereinafter, referred to as "target two-dimensional image 150A").
[0055] Furthermore, the distortion correction unit 12 acquires calibration information for the camera CM from the second database 52. The distortion correction unit 12 corrects distortion in the target two-dimensional image 150A based on the calibration information. The distortion in this case is, for example, lens distortion. The lens distortion is, for example, distortion aberration.
[0056] Next, the image placement unit 13 of the calculation unit 10 will be described with reference to Figures 2, 5, and 6(a). For ease of understanding, a camera CM is shown in Figures 5 and 6(a) to 6(d).
[0057] 5 is a perspective view schematically showing the imaging position 101, the target two-dimensional image 151, and the three-dimensional model 301 in the virtual space VS. As shown in FIG. 5, in this embodiment, the target two-dimensional image 150A placed in the virtual space VS will be referred to as the "target two-dimensional image 151." Furthermore, the subject image 200 and the target image 210 of the target two-dimensional image 150A placed in the virtual space VS will be referred to as the "subject image 201" and the "target image 211," respectively.
[0058] Fig. 6(a) is a side view schematically showing an imaging position 101 displayed in a virtual space VS. Fig. 6(b) is a side view schematically showing a target two-dimensional image 151 placed in the virtual space VS.
[0059] 2, the image arrangement unit 13 acquires imaging position information 521 and imaging posture information 522 associated with the target two-dimensional image 150A from the second database 52. The imaging position information 521 indicates the imaging three-dimensional coordinates (aX, aY, aZ). As shown in FIGS. 5 and 6(a), the imaging position 101 in the virtual space VS is indicated by the imaging three-dimensional coordinates (aX, aY, aZ).
[0060] Next, as shown in FIGS. 5 and 6(b), the image placement unit 13 places the distortion-corrected target two-dimensional image 151 in the virtual space VS based on the imaging position information 521 and the imaging attitude information 522.
[0061] Specifically, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS at a position a predetermined distance L away from the captured three-dimensional coordinates (aX, aY, aZ) along the optical axis 523. In this case, the image placement unit 13 places the target two-dimensional image 151 in the virtual space VS so that the target two-dimensional image 151 is perpendicular to the optical axis 523 of the camera CM and the size of the target two-dimensional image 151 corresponds to the angle of view θ of the camera CM.
[0062] In this case, the image arrangement unit 13 calculates three-dimensional coordinates (bX, bY, bZ), (cX, cY, cZ), and (dX, dY, dZ) in the virtual space VS for arranging at least three corners 161, 162, and 163 of the four corners 161, 162, 163, and 164 of the target two-dimensional image 151. Then, the image arrangement unit 13 arranges at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151 at the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ), thereby arranging the target two-dimensional image 151 in the virtual space VS.
[0063] Next, the target detection unit 14 of the calculation unit 10 will be described with reference to Fig. 2, Fig. 5, Fig. 6(c), and Fig. 7. Fig. 6(c) is a side view schematically showing the target image 211 and bounding box 221 included in the target two-dimensional image 151 placed in the virtual space VS, and the target area 311 included in the three-dimensional model 301. Fig. 7 is a front view showing the target two-dimensional image 151 placed in the virtual space VS. In Fig. 7, the target two-dimensional image 151 is viewed from the imaging position 101.
[0064] As shown in FIG. 2, the target detection unit 14 preferably uses a trained model 53 constructed by training training data. The trained model 53 is constructed by training training data. In this case, the training is supervised learning. As an example, the trained model 53 uses YOLO (You Only Look Once), an object detection algorithm that uses a convolutional neural network. The object detection algorithm is not particularly limited and may be, for example, R-CNN (Regions with Convolutional Neural Network) or SSD (Single Shot Multi-Box Detector). The training data includes a two-dimensional image including a target image and an information tag. The information tag is added to the two-dimensional image. The information tag includes a bounding box that surrounds the target image included in the subject image in the two-dimensional image. The information tag may further include the name of the target image. The bounding box corresponds to an example of "information identifying the target image" in the present disclosure.
[0065] The object detection algorithm may also be an algorithm that performs segmentation, such as U-NET, which performs semantic segmentation, Mask R-CNN, which performs instant segmentation, or Panoptic Feature Pyramid Network, which performs panoptic segmentation.
[0066] As shown in Figures 2, 5, 6(c), and 7, the target detection unit 14, for example, causes the trained model 53 to detect the target image 211 included in the target two-dimensional image 151 and sets a bounding box 221 surrounding the target image 211. In this way, the target detection unit 14 causes the trained model 53 to detect the target image 211 and identify the target image 211. In other words, when the target detection unit 14 inputs the target two-dimensional image 151 to the trained model 53, the trained model 53 outputs the target two-dimensional image 151 including the target image 211 surrounded by the bounding box 221. Note that in Figure 6(c), the bounding box 221 is exaggerated to make the drawing easier to see.
[0067] Although the target detection unit 14 uses the trained model 53, it may also detect the target image 211 using, for example, a rule-based algorithm.
[0068] Next, the target coordinate calculation unit 15 will be described with reference to Fig. 2, Fig. 5, Fig. 6(c), Fig. 6(d), and Fig. 7. Fig. 6(d) is a side view that schematically shows a state in which a straight line 400 passing through the imaging position 101 and the target image 211 hits a three-dimensional model 301 in the virtual space VS.
[0069] 2, 5, 6(c), and 7, the target coordinate calculation unit 15 calculates the three-dimensional coordinates (eX, eY, eZ) in the virtual space VS of the target image 211 included in the target two-dimensional image 151. Hereinafter, the three-dimensional coordinates (eX, eY, eZ) will be referred to as "first target three-dimensional coordinates (eX, eY, eZ)."
[0070] Specifically, the target coordinate calculation unit 15 calculates the first target three-dimensional coordinates (eX, eY, eZ) based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.
[0071] 7, the first target three-dimensional coordinates (eX, eY, eZ) are preferably indicated by the three-dimensional coordinates in the virtual space VS of a specific point 224 in a bounding box 221. As an example, the specific point 224 is the intersection of a diagonal line 222 and a diagonal line 223 of the bounding box 221. In this manner, the specific point 224 is determined based on the bounding box 221. The specific point 224 indicates a point inside the target image 211. The bounding box 221 is a rectangle that surrounds the target image 211.
[0072] In this case, specifically, the target coordinate calculation unit 15 calculates the two-dimensional coordinates (u, v) of a specific point 224 in a bounding box 221 in the target two-dimensional image 151. The two-dimensional coordinates (u, v) are coordinates in a screen coordinate system SC. The screen coordinate system SC is defined by a U axis and a V axis that are perpendicular to each other. The origin of the screen coordinate system SC is not particularly limited, but is set to a corner 161 of the target two-dimensional image 151, for example. Note that the origin may also be set to the center of the target two-dimensional image 151, for example.
[0073] Then, the target coordinate calculation unit 15 converts the two-dimensional coordinates (u, v) of the specific point 224 in the screen coordinate system SC into three-dimensional coordinates, i.e., first target three-dimensional coordinates (eX, eY, eZ), based on the three-dimensional coordinates (bX, bY, bZ) to (dX, dY, dZ) of at least three corners 161 to 163 of the four corners 161 to 164 of the target two-dimensional image 151.
[0074] Although the specific point 224 is determined based on the bounding box 221, it is not particularly limited as long as it is a point inside the target image 211.
[0075] Next, as shown in FIGS. 5 and 6(d), the target coordinate calculation unit 15 calculates the three-dimensional coordinates (fX, fY, fZ) of the position where a straight line 400 passing through the captured three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ) hits the three-dimensional model 301. The three-dimensional coordinates (fX, fY, fZ) in this case are referred to as "second target three-dimensional coordinates (fX, fY, fZ)." In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point when the straight line 400 intersects with the three-dimensional model 301. The straight line 400 is, for example, a virtual straight line.
[0076] The second target three-dimensional coordinates (fX, fY, fZ) indicate the three-dimensional coordinates of a target area 311 on the three-dimensional model 301. The target area 311 corresponds to the target image 211, and is an area in which the target 310 in the subject 300 is reproduced on the three-dimensional model 301.
[0077] Specifically, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ) which are the position where the straight line 400 hits the surface of the three-dimensional model 301. In other words, the second target three-dimensional coordinates (fX, fY, fZ) are the three-dimensional coordinates of the intersection point when the straight line 400 intersects with the surface of the three-dimensional model 301.
[0078] For example, when a plurality of polygonal (e.g., triangular) faces having point 322 of point cloud 321 as a vertex are assigned to three-dimensional model 301, target coordinate calculation unit 15 sets the average value of the three-dimensional coordinates of a plurality of vertices that define the face on which straight line 400 strikes, among the plurality of polygonal faces, as the second target three-dimensional coordinates. Alternatively, for example, target coordinate calculation unit 15 sets the three-dimensional coordinates of the vertex that is closest to the position on which straight line 400 strikes, among the plurality of vertices that define the face on which straight line 400 strikes, as the second target three-dimensional coordinates.
[0079] For example, if the three-dimensional model 301 is assigned a circular surface centered on each point 322 that constitutes the point cloud 321, the target coordinate calculation unit 15 sets the three-dimensional coordinates of the point 322 at the center of the circular surface where the straight line 400 hits as the second target three-dimensional coordinates.
[0080] For example, if multiple cubes are assigned to the three-dimensional model 301 by the octree algorithm, the target coordinate calculation unit 15 sets the three-dimensional coordinates of the point 322 contained in the cube having the face on which the line 400 strikes as the second target three-dimensional coordinates. If multiple points 322 are contained in the cube, the average value of the three-dimensional coordinates of the multiple points 322 is set as the second target three-dimensional coordinates.
[0081] As described above with reference to FIGS. 5 and 6 , according to this embodiment, the target coordinate calculation unit 15 calculates the second target three-dimensional coordinates (fX, fY, fZ), which are the three-dimensional coordinates of the target area 311 on the three-dimensional model 301, by determining the position in the virtual space VS where a line 400 passing through the captured three-dimensional coordinates (aX, aY, aZ) and the first target three-dimensional coordinates (eX, eY, eZ) of the target two-dimensional image 151 strikes the three-dimensional model 301. As described above, in this embodiment, instead of performing a process of searching for the target area throughout the entire area of the three-dimensional model, the position of the target area 311 on the three-dimensional model 301 is identified based on the captured three-dimensional coordinates (aX, aY, aZ) and one target two-dimensional image 151. In other words, the position of the target area 311 on the three-dimensional model 301 is identified by processing the two-dimensional image (target two-dimensional image 151). Therefore, the position of the target area 311 on the three-dimensional model 301 can be identified by simple processing.
[0082] In addition, according to this embodiment, the position of the target region 311 on the three-dimensional model 301 is automatically identified by the calculation unit 10. This eliminates the need for a human to visually search for the target region 311 on the three-dimensional model 301. This reduces the workload on the human worker.
[0083] In particular, the imaging position information 521 (imaging three-dimensional coordinates (aX, aY, aZ)) and the imaging attitude information 522 are information that are necessarily calculated as by-products when generating the three-dimensional model 301. Therefore, it is possible to prevent additional processing from occurring for processing to identify the target region 311 on the three-dimensional model 301.
[0084] Furthermore, according to this embodiment, the model generation unit 11 adds a surface to the three-dimensional model 301. Therefore, the straight line 400 always hits the three-dimensional model 301. Therefore, the second target three-dimensional coordinates (fX, fY, fZ) indicating the position of the target region 311 can be reliably calculated without depending on the density of the point cloud 321.
[0085] Furthermore, according to this embodiment, the distortion correction unit 12 corrects distortion of the target two-dimensional image 151. Therefore, even if the target image 211 is located in an area of the target two-dimensional image 151 where distortion occurs, the two-dimensional coordinates (u, v) and the first target three-dimensional coordinates (eX, eY, eZ) indicating the position of the target image 211 can be calculated with high accuracy.
[0086] Furthermore, according to this embodiment, the target detection unit 14 inputs the target two-dimensional image 151 into the trained model 53 to detect the target image 211. That is, the trained model 53 is constructed by learning using the two-dimensional image 150 as training data. Therefore, compared to constructing a trained model that detects a target image by training a three-dimensional model, the burden on the operator when creating training data, the size of the training data, and the amount of calculation required for training can be reduced. As a result, the cost and training time required for constructing the trained model 53 can be reduced. In particular, instead of detecting the target area by analyzing a three-dimensional model, the target image 211 is detected on the two-dimensional image (the target two-dimensional image 151). Therefore, according to this embodiment, the target area 311 on the three-dimensional model 301 can be identified by fast and simple processing compared to detecting the target area by analyzing a three-dimensional model.
[0087] Furthermore, according to this embodiment, the first target three-dimensional coordinates (eX, eY, eZ) are the three-dimensional coordinates of a specific point 224 that is determined based on the bounding box 221. That is, the specific point 224 is determined as the intersection of the diagonal line 222 and the diagonal line 223 of the bounding box 221. Therefore, compared to the case where the specific point 224 is determined at an arbitrary point on the target image 211, the specific point 224 can be determined easily by simple processing.
[0088] Next, an image analysis method according to this embodiment will be described with reference to FIGS. 2 and 8. FIG. 8 is a flowchart showing the image analysis method. The image analysis method is executed by the server 1. As shown in FIG. 8, the image analysis method includes steps S1 to S9. A computer program stored in the storage unit 50 causes the calculation unit 10 to execute steps S1 to S9. In other words, the computer program product realizes steps S1 to S9 when the computer program is executed by the calculation unit 10. The calculation unit 10 corresponds to an example of a "computer" in the present disclosure.
[0089] As shown in Figures 2 and 8, first, in step S1, the model generation unit 11 generates a three-dimensional model 301 of the subject 300 based on multiple two-dimensional images 150 generated by capturing images of the subject 300 from multiple different imaging positions 100.
[0090] Next, in step S2, the distortion corrector 12 acquires a target two-dimensional image 150A from the second database 52, out of the plurality of two-dimensional images 150.
[0091] Next, in step S3, the distortion corrector 12 corrects the distortion of the target two-dimensional image 150A based on the calibration information. The target two-dimensional image 150A after the distortion correction is stored in the storage unit 50.
[0092] Next, in step S4, the image arrangement unit 13 acquires imaging position information 521 and imaging attitude information 522 from the second database 52. The imaging position information 521 indicates the position of the camera CM at the time of imaging by imaging three-dimensional coordinates (aX, aY, aZ) in the three-dimensional coordinate system CS. The imaging attitude information 522 indicates the attitude of the camera CM at the time of imaging by a rotation angle in the three-dimensional coordinate system CS.
[0093] Next, in step S5, the image placement unit 13 places the distortion-corrected target two-dimensional image 150A in the virtual space VS based on the imaging position information 521 and the imaging attitude information 522. The target two-dimensional image 150A placed in the virtual space VS is referred to as a "target two-dimensional image 151."
[0094] Next, in step S6, the target detection unit 14 causes the trained model 53 to detect the target image 211 included in the target two-dimensional image 151, and sets a bounding box 221 surrounding the target image 211.
[0095] Next, in step S7, the target coordinate calculation unit 15 calculates first target three-dimensional coordinates (eX, eY, eZ) that indicate the three-dimensional coordinates of the specific point 224 in the bounding box 221.
[0096] Next, in step S8, target coordinate calculation unit 15 calculates second target three-dimensional coordinates (fX, fY, fZ) indicating the position where line 400 passing through the imaging three-dimensional coordinates (aX, aY, aZ) indicated by imaging position information 521 and the first target three-dimensional coordinates (eX, eY, eZ) hits three-dimensional model 301. The second target three-dimensional coordinates (fX, fY, fZ) indicate the three-dimensional coordinates of target area 311 on three-dimensional model 301. In this way, the position of target 310 included in subject 300 is identified as target area 311 on three-dimensional model 301.
[0097] Next, in step S9, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 via the communication unit 40 in response to a request from the user's terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in an operable manner. For example, the display control unit 16 displays the target two-dimensional image 151 on the three-dimensional model 301 superimposed on the three-dimensional model 301 on the terminal 2. For example, the display control unit 16 displays the three-dimensional model 301 on the terminal 2 in such a manner that the target area 311 on the three-dimensional model 301 is displayed on the terminal 2. In this case, for example, the display control unit 16 enlarges the target area 311 on the three-dimensional model 301 and displays it on the terminal 2. Therefore, according to this embodiment, the user of the terminal 2 can omit the task of visually searching for the target area 311 on the three-dimensional model 301. In other words, the user's workload can be reduced. When step S9 is completed, the image analysis method ends.
[0098] 8, according to the image analysis method of this embodiment, the position of the target area 311 on the three-dimensional model 301 is identified by calculating the position where the straight line 400 strikes the three-dimensional model 301 in the virtual space VS. Therefore, the position of the target area 311 can be identified by simple processing.
[0099] Fig. 9 is a flowchart showing the three-dimensional model generation process in step S1 of Fig. 8. As shown in Fig. 9, the three-dimensional model generation process includes steps S11 to S15.
[0100] First, in step S11, the model generating unit 11 acquires from the first database 51 a plurality of two-dimensional images 150 including the subject image 200.
[0101] Next, in step S12, the model generation unit 11 extracts feature points of the subject 300 from each of the two-dimensional images 150, and associates the feature points between the two-dimensional images 150.
[0102] Next, in step S13, the model generation unit 11 estimates, based on the associated feature points, the image capture position 100 and the posture of the camera CM when each two-dimensional image 150 was generated. Then, the model generation unit 11 calculates the three-dimensional coordinates of the three-dimensional points corresponding to the feature points based on the estimation result of the image capture position 100 (image capture position information 521) and the estimation result of the posture of the camera CM (image capture posture information 522). As a result, point cloud data indicating a point cloud 321 consisting of a plurality of three-dimensional points (a plurality of points 322) is generated.
[0103] Next, in step S14, the model generation unit 11 adds surfaces to the three-dimensional model 301 based on the point cloud data.
[0104] Next, in step S15, the model generation unit 11 adds material information to the surfaces added to the three-dimensional model 301 based on the two-dimensional image 150. Then, the process returns to step S2 in FIG.
[0105] (Variation) A modified example of this embodiment will be described with reference to Figures 2 and 10. In this modified example, target 310 included in subject 300 is a portion of subject 300 that indicates a specific state. Below, differences between this modified example and the above embodiment will be mainly described.
[0106] In the modified example, the target image 211 shown in FIG. 7 is an image of a portion of the subject 300 that indicates a specific state (hereinafter, "specific state portion"). Therefore, according to the modified example, the specific state portion of the subject 300 is specified as a target region 311 on the three-dimensional model 301. That is, the position of the target region 311 that indicates the specific state portion of the subject 300 is calculated as second target three-dimensional coordinates on the three-dimensional model 301. As a result, it is possible to omit the human task of visually searching for the target region 311 that indicates the specific state portion on the three-dimensional model 301. It is also possible to prevent the target region 311 that indicates the specific state portion from being overlooked when it is detected.
[0107] The specific state portion of the subject 300 is, for example, a portion of the entire subject 300 or a region of interest that indicates an abnormal state that is not a normal state (hereinafter referred to as an "abnormal state portion"). Therefore, in this case, the target image 211 is an image of the abnormal state portion. As a result, it is possible to omit the human task of visually searching for the target region 311 that indicates the abnormal state portion on the three-dimensional model 301. It is also possible to prevent the target region 311 that indicates the abnormal state portion from being overlooked.
[0108] As an example, the target image 211 indicates an area of the object 2D image 151 that has changed relative to the reference 2D image. An area that has changed relative to the reference 2D image is, for example, an area that has been deformed, mutated, displaced, increased, decreased, or disappeared relative to the reference 2D image. The reference 2D image is, for example, a 2D image of the object 300 captured earlier than the object 2D image 151. Alternatively, the reference 2D image is, for example, a 2D image generated by capturing an image of the object 300 in a reference state. The reference state is, for example, a state in which the entire object 300 or a region of interest is normal. In this example, the target image 211 is an image of a portion of the entire object 300 or a region of interest that shows an abnormal state.
[0109] Next, an image analysis method according to a modified example will be described with reference to Fig. 2 and Fig. 10. Fig. 10 is a flowchart showing the image analysis method according to the modified example. The image analysis method is executed by the server 1. As shown in Fig. 10, the image analysis method includes steps S21 to S23. A computer program stored in the storage unit 50 causes the calculation unit 10 to execute steps S21 to S23. In other words, the computer program product realizes steps S21 to S23 when the computer program is executed by the calculation unit 10.
[0110] First, in step S21, the target detection unit 14 executes a target image detection process. The target image detection process refers to a process of detecting a target image 210 from video data 511 of the subject 300, a two-dimensional image 150 extracted from the video data 511, a still image data set 512, or a two-dimensional image 150 extracted from the still image data set 512. Hereinafter, the video data 511, the two-dimensional image 150 extracted from the video data 511, the still image data set 512, and the still image data set 512 will be collectively referred to as "detection target image data."
[0111] For example, the target detection unit 14 inputs detection target image data to a trained model. If the trained model detects a target image 210 from the detection target image data, it outputs information indicating that the detection target image data includes the target image 210. If the trained model does not detect a target image 210 from the detection target image data, it outputs information indicating that the detection target image data does not include the target image 210.
[0112] A trained model is a computer program constructed by learning training data. The trained model includes, for example, a neural network. The learning is, for example, supervised learning. The training data is detection target image data and "information indicating that the target image is included."
[0113] Next, in step S22, the target detection unit 14 determines whether or not the target image 210 has been detected in the target image detection process.
[0114] If the determination in step S22 is negative (NO), the image analysis method ends.
[0115] On the other hand, if the determination in step S22 is affirmative (YES), the process proceeds to step S23.
[0116] Next, in step S23, the calculation unit 10 identifies a target region 311 on the three-dimensional model 301 that corresponds to the target image 210. The processing in step S23 is similar to the processing in steps S1 to S9 shown in Fig. 8. When step S23 ends, the image analysis method ends.
[0117] As described above with reference to Fig. 10, according to the modified example, the processing of step S23 is executed only when the target image 210 is detected. This reduces the processing load on the calculation unit 10. In particular, when the target image 210 is not detected, the three-dimensional model 301 is not generated, thereby reducing the processing load on the calculation unit 10. Furthermore, in this case, an increase in the amount of data stored in the storage unit 50 can be suppressed.
[0118] The image analysis method shown in FIG. 10 can also be applied to the above-described embodiment described with reference to FIGS.
[0119] Although the preferred embodiments and modifications of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0120] The devices or systems described herein may be implemented as a single device, or may be implemented by multiple devices (e.g., cloud servers) partially or entirely connected via a network. For example, some or all of the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by the same computer or server. Furthermore, for example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may be implemented by the terminal 2. For example, the model generation unit 11, distortion correction unit 12, image placement unit 13, target detection unit 14, target coordinate calculation unit 15, and display control unit 16 may each be implemented by a separate computer or server. For example, the first database 51, the second database 52, and the trained model 53 may each be stored in separate storage devices or servers.
[0121] The series of processes performed by the device described herein may be implemented using software, hardware, or a combination of software and hardware. A computer program for implementing each function of the calculation unit 10 according to this embodiment may be created and installed on a PC or the like. A computer-readable recording medium storing such a computer program may also be provided. Examples of the recording medium include a magnetic disk, an optical disk, a magneto-optical disk, and a flash memory. The computer program may also be distributed, for example, via a network, without using a recording medium.
[0122] Furthermore, the processes described herein using flowchart diagrams do not necessarily have to be performed in the order shown. Some process steps may be performed in parallel. Additional process steps may be employed, and some process steps may be omitted.
[0123] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0124] The following configurations also fall within the technical scope of the present disclosure.
[0125] (Item 1) a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating an imaging position when a target two-dimensional image of the plurality of two-dimensional images was generated, the three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of a camera that performed imaging; a target coordinate calculation unit that calculates second target three-dimensional coordinates that indicate a position where a straight line passing through the imaged three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model, The first target three-dimensional coordinates indicate the position in the virtual space of a target image included in the object two-dimensional image.
[0126] (Item 2) the three-dimensional model is configured by point cloud data, the model generation unit adds a surface to the three-dimensional model based on the point cloud data; Item 2. The image analysis device according to item 1, wherein the target coordinate calculation unit calculates the second target three-dimensional coordinates that are the position where the straight line hits the surface.
[0127] (Item 3) Further, a distortion correction unit corrects distortion of the target two-dimensional image, 3. The image analyzing device according to claim 1, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the object two-dimensional image in which the distortion has been corrected.
[0128] (Item 4) The method further includes a target detection unit that causes a trained model constructed by training training data to detect the target image included in the target two-dimensional image and identifies the target image, the training data includes a two-dimensional image including a target image and an information tag; 4. The image analyzing device according to any one of items 1 to 3, wherein the information tag includes information for identifying the target image.
[0129] (Item 5) the target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image; 5. The image analysis device according to item 4, wherein the first target three-dimensional coordinates indicate three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.
[0130] (Item 6) 6. The image analyzing device according to any one of items 1 to 5, wherein the target image is an image of a portion of the subject that shows a specific state.
[0131] (Item 7) generating a three-dimensional model of the object based on a plurality of two-dimensional images generated by capturing images of the object from a plurality of different imaging positions; acquiring three-dimensional coordinates indicating an imaging position when a target two-dimensional image of the plurality of two-dimensional images was generated, the three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of a camera that performed imaging; placing the target two-dimensional image in the virtual space based on the image capturing three-dimensional coordinates and the image capturing attitude information; calculating first target three-dimensional coordinates, which are three-dimensional coordinates in the virtual space of a target image included in the target two-dimensional image; and calculating second target three-dimensional coordinates, which are the three-dimensional coordinates of the position where a straight line passing through the imaging three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model.
[0132] (Item 8) On the computer, generating a three-dimensional model of the object based on a plurality of two-dimensional images generated by capturing images of the object from a plurality of different imaging positions; acquiring three-dimensional coordinates indicating an imaging position when a target two-dimensional image of the plurality of two-dimensional images was generated, the three-dimensional coordinates being indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of a camera that performed imaging; placing the target two-dimensional image in the virtual space based on the image capturing three-dimensional coordinates and the image capturing attitude information; calculating first target three-dimensional coordinates, which are three-dimensional coordinates in the virtual space of a target image included in the target two-dimensional image; and calculating second target three-dimensional coordinates, which are the three-dimensional coordinates of a position where a straight line passing through the imaged three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model. [Industrial Applicability]
[0133] The present disclosure provides an image analysis device, an image analysis method, and a computer program, and has industrial applicability. [Explanation of symbols]
[0134] 1 Server (Image Analysis Device), 11 Model Generation Unit, 12 Distortion Correction Unit, 13 Image Placement Unit, 14 Target Detection Unit, 15 Target Coordinate Calculation Unit, 16 Display Control Unit, 53 Trained Model
Claims
1. a model generation unit that generates a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions; an image placement unit that places the target two-dimensional image in the virtual space based on three-dimensional coordinates indicating an imaging position, which is the position of a camera when a target two-dimensional image of the plurality of two-dimensional images is generated, and which is indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of the camera that performed imaging; a target coordinate calculation unit that calculates second target three-dimensional coordinates that indicate a position where a line passing through the imaged three-dimensional coordinates and the first target three-dimensional coordinates strikes the three-dimensional model; a display control unit, the first target three-dimensional coordinates indicate a position in the virtual space of a target image included in the target two-dimensional image; the image capture three-dimensional coordinates and the image capture posture information are calculated as by-products when generating the three-dimensional model, the image placement unit places the two-dimensional image of the target at a position in the virtual space that is a predetermined distance away from the imaging three-dimensional coordinates along the optical axis so that the two-dimensional image of the target is perpendicular to the optical axis of the camera and the size of the two-dimensional image of the target corresponds to the angle of view of the camera; The display control unit overlays the target two-dimensional image on the three-dimensional model and displays it on the terminal.
2. the three-dimensional model is configured by point cloud data, the model generation unit adds a surface to the three-dimensional model based on the point cloud data; The image analysis device according to claim 1 , wherein the target coordinate calculation unit calculates the second target three-dimensional coordinates that are positions where the straight line strikes the surface.
3. Further, a distortion correction unit corrects distortion of the target two-dimensional image, 3. The image analysis device according to claim 1, wherein the target coordinate calculation unit calculates the first target three-dimensional coordinates based on the object two-dimensional image in which the distortion has been corrected.
4. The method further includes a target detection unit that causes a trained model constructed by training training data to detect the target image included in the target two-dimensional image and identifies the target image, the training data includes a two-dimensional image including a target image and an information tag; The image analysis device according to claim 1 , wherein the information tag includes information that identifies the target image.
5. the target detection unit causes the trained model to detect the target image and sets a bounding box surrounding the target image; The image analysis device according to claim 4 , wherein the first target three-dimensional coordinates indicate three-dimensional coordinates in the virtual space of a specific point determined based on the bounding box.
6. 3. The image analysis device according to claim 1, wherein the target image is an image of a portion of the subject that shows a specific state.
7. The point cloud data represents a point cloud which is a set of a plurality of points, the face is a polygonal face having three or more of the points as vertices, a circular face having the point as its center, or a cubic face including the point, The target coordinate calculation unit If the surface on which the straight line strikes is a surface of the polygon, calculate the second target three-dimensional coordinates based on the three-dimensional coordinates of the point of the vertex; The image analysis device according to claim 2 , wherein when the surface that the straight line strikes is the circular surface, the three-dimensional coordinates of the point at the center are set as the second target three-dimensional coordinates.
8. The model generation unit recursively divides the three-dimensional model to represent the three-dimensional model using a plurality of cubes; The image analysis device according to claim 2 , wherein the faces are faces of the cube.
9. An image analysis device as described in claim 1 or claim 2, further comprising a target detection unit that detects the target image contained in the target two-dimensional image after the target two-dimensional image is placed in the virtual space.
10. A method of generating a three-dimensional model of a subject based on a plurality of two-dimensional images generated by capturing images of the subject from a plurality of different imaging positions by a computer; a step in which the computer acquires three-dimensional coordinates indicating an imaging position, which is a position of a camera when a target two-dimensional image of the plurality of two-dimensional images is generated, and which are indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of the camera that performed the imaging; a step of the computer placing the target two-dimensional image in the virtual space based on the image capturing three-dimensional coordinates and the image capturing attitude information; a step of calculating, by the computer, first target three-dimensional coordinates which are three-dimensional coordinates in the virtual space of a target image included in the object two-dimensional image; and calculating, by the computer, second target three-dimensional coordinates which are three-dimensional coordinates of a position where a straight line passing through the imaging three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model, the image capture three-dimensional coordinates and the image capture posture information are calculated as by-products when generating the three-dimensional model, In the step of arranging the two-dimensional image of the target, the computer arranges the two-dimensional image of the target in the virtual space at a position a predetermined distance away from the imaging three-dimensional coordinate along the optical axis so that the two-dimensional image of the target is perpendicular to the optical axis of the camera and the size of the two-dimensional image of the target corresponds to the angle of view of the camera; An image analysis method in which the two-dimensional image of the object is displayed on a terminal while being superimposed on the three-dimensional model.
11. On the computer, generating a three-dimensional model of the object based on a plurality of two-dimensional images generated by capturing images of the object from a plurality of different imaging positions; acquiring three-dimensional coordinates indicating an imaging position, which is the position of a camera when a target two-dimensional image of the plurality of two-dimensional images is generated, and which are indicated by three-dimensional coordinates in a virtual space in which the three-dimensional model is placed, and imaging attitude information indicating the attitude of the camera that performed the imaging; placing the target two-dimensional image in the virtual space based on the image capturing three-dimensional coordinates and the image capturing attitude information; calculating first target three-dimensional coordinates, which are three-dimensional coordinates in the virtual space of a target image included in the target two-dimensional image; calculating second target three-dimensional coordinates, which are three-dimensional coordinates of a position where a straight line passing through the imaged three-dimensional coordinates and the first target three-dimensional coordinates hits the three-dimensional model; the image capture three-dimensional coordinates and the image capture posture information are calculated as by-products when generating the three-dimensional model, In the step of arranging the two-dimensional image of the target, the two-dimensional image of the target is arranged in the virtual space at a position that is a predetermined distance away from the imaging three-dimensional coordinate along the optical axis so that the two-dimensional image of the target is perpendicular to the optical axis of the camera and the size of the two-dimensional image of the target corresponds to the angle of view of the camera; The two-dimensional image of the object is displayed on a terminal so as to be superimposed on the three-dimensional model.
Citation Information
Patent Citations
Electronic device and three-dimensional model generation support method
JP2013114498A
Inspection system and inspection method
JP2020021466A
Program, method, system, road map, and road map creation method
JP2024013788A
Deterioration diagnosis system using flight vehicle
JP2019070631A