Information processing method and information processing system
The method addresses the challenge of efficiently tagging the same object across multiple images by calculating correspondence information and assigning tags using algorithms like LoFTR, achieving precise and cost-effective tagging without extensive 3D reconstruction.
Patent Information
- Application Number
- JP2024069064
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-04
AI Technical Summary
Existing AI-based object detection systems struggle to efficiently assign the same tag information to the same object across multiple images, failing to correctly identify and tag unknown objects.
An information processing method that calculates correspondence information between images using algorithms like LoFTR, identifies corresponding points, and assigns the same tag to these points across multiple images, allowing for efficient tagging even with a small number of images or unknown objects.
Enables precise and efficient assignment of tag information to the same object in multiple images, reducing processing costs and improving accuracy without the need for 3D reconstruction from numerous images.
Smart Images

Figure 2025165140000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing method and an information processing system. [Background technology]
[0002] There are technologies that support the assignment of tag information to objects and their accessories in images captured from multiple viewpoints. For example, Patent Document 1 discloses a technology that identifies an object to which annotation tag information is to be assigned from objects included in images captured from multiple viewpoints, and changes the display position of the tag information on the screen when the position of the object to which tag information is to be assigned on the screen changes in response to a change in viewpoint. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2023-85999 [Non-patent literature]
[0004] [Non-Patent Document 1] Jiaming Sun, et al., “LoFTR: Detector-Free Local Feature Matching with Transformers,” [online], June 20-25, 2021, Conference on Computer Vision and Pattern Recognition (CVPR), [Retrieved March 29, 2024], Internet, <url>https: / / openaccess.thecvf.com / content / CVPR2021 / papers / Sun_LoFTR_Detector-Free_Local_Feature_Matching_With_Transformers_CVPR_2021_paper.pdf. [Non-patent document 2] Johannes L. Schonberger, et al., "Structure-from-Motion Revisited," [online], June 27-30, 2016, Conference on Computer Vision and Pattern Recognition (CVPR), [Retrieved March 29, 2024], Internet, <url>https: / / demuc.de / papers / schoenberger2016sfm.pdf. Summary of the Invention [Problem to be solved by the invention]
[0005] In recent years, there has been a growing demand for systems that use object detection, which uses AI (Artificial Intelligence) to detect objects from images, to improve the efficiency of object confirmation tasks that have previously been performed visually in various situations such as facility maintenance, safety management, and situation assessment.
[0006] However, AI-based object detection cannot correctly identify the same object that appears in multiple images, and cannot identify unknown objects. Therefore, the conventional technology disclosed in the above-mentioned Patent Document 1 has a problem in that even if the results of AI-based object detection are used, it is not possible to efficiently assign the same tag information to the same object in multiple images.
[0007] The present invention has been made in consideration of the above-mentioned problems, and has as its object to efficiently assign the same tag information to the same object in multiple images. [Means for solving the problem]
[0008] In order to solve the above-mentioned problems, the present invention provides an information processing method executed by an information processing system that assigns tags to images, the information processing system having a processor and a memory, wherein the processor accepts input of a plurality of images of an object photographed from a plurality of viewpoints, calculates correspondence information relating to the correspondence of pixels between the images for each pair of images among the plurality of images according to the resolution required for assigning the tag, assigns the tag to a specified position of the image specified by a user, identifies a corresponding point that is a pixel of another of the images that corresponds to a pixel included in the specified position based on the correspondence information, and assigns to the identified corresponding point a tag that is the same as the tag assigned to the pixel corresponding to the specified position. [Effects of the Invention]
[0009] According to the present invention, the same tag information can be efficiently assigned to the same object in multiple images. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing the configuration of an information processing system according to a first embodiment. [Figure 2A] FIG. 3 is a diagram showing the arrangement of a correspondence information table according to the first embodiment. [Figure 2B] FIG. 3 is a diagram showing the arrangement of a tag information table according to the first embodiment. [Figure 3] 10 is a flowchart showing tagging processing according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing the configuration of an input / output screen according to the first embodiment. [Figure 5] 10 is a flowchart showing a corresponding point display process according to the second embodiment. [Figure 6] FIG. 10 is a diagram showing the configuration of an input / output screen according to the second embodiment. [Figure 7] 11 is a flowchart showing tag search processing according to the third embodiment. [Figure 8] FIG. 11 is a diagram showing the configuration of an input / output screen according to the third embodiment. [Figure 9] FIG. 10 is a diagram showing the configuration of an information processing device according to a fourth embodiment. [Figure 10] FIG. 10 is a diagram showing the arrangement of a 2D-3D correspondence information table according to the fourth embodiment. [Figure 11] 10 is a flowchart showing a three-dimensional display process according to the fourth embodiment. [Figure 12A] FIG. 10 is a diagram showing the configuration of an input / output screen according to the fourth embodiment. [Figure 12B] FIG. 10 is a diagram showing the configuration of an input / output screen according to the fourth embodiment. [Figure 13] FIG. 1 is a diagram showing the hardware configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0012] [Embodiment 1] (Configuration of Information Processing System 1 According to Embodiment 1) 1 is a diagram showing the configuration of an information processing system 1 according to a first embodiment. The information processing system 1 is a computer constructed in an on-premise environment or a cloud environment. The information processing system 1 has an image input unit 11, a correspondence information calculation unit 12, a tag assignment unit 13, and an input / output unit 14, which are realized by a CPU (Central Processing Unit) that executes a program. The information processing system 1 also has an image data storage unit 15, a correspondence information storage unit 16, and a tag information storage unit 17, which are provided in a predetermined storage area.
[0013] The image input unit 11 accepts input of multiple images of the same structure. The image input unit 11 stores the multiple images that it has accepted in the image data storage unit 15 for each structure that is the subject. The multiple images stored in the image data storage unit 15 for each structure are still images of the same structure taken from multiple viewpoints. In this embodiment, the subject of the images is a structure, but is not limited to a structure. Also, in this embodiment, it is assumed that the number of multiple images stored in the image data storage unit 15 is small enough not to form a precise mesh, but is not particularly limited.
[0014] The correspondence information calculation unit 12 calculates correspondence information T11 (FIG. 2A) relating to pixel correspondence between images according to the resolution required for tagging for each pair of images (hereinafter referred to as "image pair") from the multiple images for each structure stored in the image data storage unit 15. The "resolution required for tagging" refers to an image resolution that allows tags to be assigned to pixels within the image.
[0015] The correspondence information calculation unit 12 calculates correspondence information T11 for all image pairs created from multiple images of each structure using an algorithm for determining pixel correspondence between images, such as LoFTR (see Non-Patent Document 1). The correspondence information T11 is a set of combinations of pixel coordinates of each pixel in correspondence between images and a reliability indicating the likelihood of the correspondence. In the case of an image pair of image 1 and image 2 that are in correspondence, the correspondence information T11 is a set of (coordinates of image 1, coordinates of image 1, reliability). The correspondence information calculation unit 12 stores the calculated correspondence information T11 in the correspondence information storage unit 16 in the form of a correspondence information table T1 (FIG. 2A).
[0016] Since the correspondence information obtained by LoFTR generally includes outliers, it is possible to thin out the outliers using a robust estimation algorithm such as RANSAC (RANdom Sample Consensus). Thinning outliers allows more accurate correspondence information to remain, preventing the need to redo work due to correspondence errors. Furthermore, although correspondence information is obtained at the pixel level, it can be thinned to the resolution level required for tagging, depending on the size of the subject. Thinning reduces the amount of data in the database.
[0017] The tagging unit 13 assigns a tag to a specified position in an image specified by the user. The "specified position" is a portion of the image to be tagged that the user focuses on, such as a damaged portion of a structure. The tagging unit 13 also identifies corresponding points, which are pixels in other images that correspond to the pixels included in the specified position, based on the correspondence information T11 stored in the correspondence information table T1. The tagging unit 13 then assigns the same tag to the identified corresponding points as the tag assigned to the pixel corresponding to the specified position. Here, the "other images" refer to, for example, all images stored in the image data storage unit 15 that form an image pair with the image whose specified position is specified by the user.
[0018] The input / output unit 14 includes a display device such as a display that outputs images and the like to the user, and an input device such as a touch panel or keyboard that receives input from the user. The input / output unit 14 may be provided in the information processing system 1 or in a terminal (not shown) connected via a network (not shown). The display device of the input / output unit 14 displays input / output screens D1 (FIG. 4), D2 (FIG. 6), and D3 (FIG. 8), which will be described later.
[0019] The image data storage unit 15 stores images received by the image input unit 11. The correspondence information storage unit 16 stores correspondence information T11 in the form of a correspondence information table T1. As shown in FIG. 2A, the correspondence information table T1 manages correspondence information for each image pair. In FIG. 2A, for example, an image pair identified as image pair #1 includes image #1 and image #1', and correspondence information 11' is associated with and managed.
[0020] The tag information storage unit 17 stores tag information in the form of a tag information table T2 (FIG. 2B). The tag information storage unit 17 is an example of a storage unit that stores tag information related to tags assigned to images. The tag information table T2 is an example of tag information.
[0021] Tag information table T2 (FIG. 2B) shows the case where the tag is a category tag, and for each tag assigned to each image by tag assigning unit 13, the tag category is associated with the identification information of the image to which the tag is assigned and the coordinates indicating the position of the tag within the image, and the table manages these in association with each other. In FIG. 2B, for example, the tag identified by tag #1 is assigned the category of "XXX", and the image identified by image #1 and coordinate 1, the image identified by image #1' and coordinate 1', ... are the images to which the tag is assigned. Note that the tag assigned to each image by tag assigning unit 13 may be a free-description tag that has free-description text information regarding the content of the tag, etc., instead of or in addition to the tag category.
[0022] (Tag Assignment Process According to the First Embodiment) 3 is a flowchart showing the tagging process according to the first embodiment. The tagging process according to the first embodiment is executed at any timing designated by the user. When the tagging process is executed, it is assumed that an image received by the image input unit 11 is stored in the image data storage unit 15, and that correspondence information T11 is generated for the image stored in the image data storage unit 15 and stored in the correspondence information table T1.
[0023] First, in step S11, the tagging unit 13 assigns a tag to a specified position in an image specified by the user on the display screen (input / output screen D1, described below). Next, in step S12, the tagging unit 13 determines one point (pixel) having corresponding information within the specified position of the image specified in step S11. The user's specification of the specified position within the image in step S11 may be by selecting a pixel, selecting a rectangular range, selecting a range of multiple pixels, or by other methods. If a range is selected, the pixel with the highest reliability among the pixels included in the range that have corresponding information is determined to be the pixel from which the pixel of another image corresponding to the selected range is obtained.
[0024] Next, in step S13, the tag assignment unit 13 refers to the correspondence information table T1 and determines whether there is a corresponding point, which is a pixel in another image, that corresponds to the point determined in step S12. If there is a corresponding point that corresponds to the point determined in step S12 (step S13 YES), the tag assignment unit 13 proceeds to step S14, and if there is no corresponding point (step S13 NO), the tag assignment process ends.
[0025] In step S14, the tag assigning unit 13 refers to the correspondence information table T1, reads coordinate information of corresponding points in other corresponding images from the correspondence information, and assigns candidate tags to the tags displayed on the display screen (input / output screen D1). The candidate tags are tags that have not yet been officially assigned, and are displayed on the display screen together with an interface that prompts the user to input whether or not to officially assign the tags.
[0026] Next, in step S15, if the user confirms the candidate tag on the display screen (input / output screen D1 described later) and agrees to officially assign the tag, the tag assigning unit 13 officially assigns the tag to the corresponding point in the other image. If the user refuses to officially assign the tag, the tag assigning unit 13 does not assign the tag to the corresponding point in the other image and ends the tag assignment process.
[0027] Here, the corresponding points determined in step S12 based on the designated position of the image designated in step S11 are not necessarily correct, so the user is prompted to confirm them in step S15. This makes it possible to correct errors in the calculation algorithm for the corresponding information including the corresponding points and reliability such as LoFTR.
[0028] When the image input unit 11 receives input of a new image of an object, the correspondence information calculation unit 12 calculates correspondence information for each image pair between the new image and one of the multiple images of the same structure stored in the image data storage unit 15. Then, based on the correspondence information, the tag assignment unit 13 identifies pixels in one of the multiple images that correspond to pixels in the new image. Then, based on the tag information, the tag assignment unit 13 assigns the same tag to the pixels of the new image as the tag assigned to a pixel in one of the multiple images that corresponds to the pixel of the new image. This allows tags to be automatically assigned as needed even when a new image is added, without the user having to perform tag assignment work.
[0029] (Input / output screen D1 according to embodiment 1) 4 is a diagram showing the configuration of the input / output screen D1 according to the first embodiment. The input / output screen D1 is displayed on the display screen by cooperation of the image input unit 11, the correspondence information calculation unit 12, the tag assignment unit 13, and the input / output unit 14. The input / output screen D1 has an image list display area D11 and an image display area D12. The image list display area D11 displays a list of image names of images stored in the image data storage unit 15. For images to which tags have been assigned, the image list display area D11 displays the tag names in a tree format along with the image names.
[0030] In addition, in the image list display area D11, for images ("image2", "image3", "image4") for which corresponding points in other images corresponding to points to which tags have been assigned exist in the correspondence information, an identification mark "!" is displayed in the image name to indicate the presence of corresponding points. In addition, in the image list display area D11, for "tag1" of "image2" for which a confirmation display D12d is displayed, which indicates to the user whether or not to assign a tag, an identification mark "!" is displayed along with a dashed frame.
[0031] The image display area D12 displays an image selected by the user from the image list display area D11. In the example of Fig. 4, "image1" and "image2" are selected in the image list display area D11 as indicated by hatching, and therefore an image D12a of "image1" and an image D12b of "image2" are displayed in the image display area D12.
[0032] Image D12a of "image1" displays tag D12c (with text "Damaged A") of "tag1" that has already been assigned. On the other hand, image D12b of "image2" displays confirmation display D12d that prompts the user to confirm whether or not to assign a tag similar to tag D12c to a corresponding point in image D12a to which tag D12c has been assigned in the correspondence information. When the user presses "Confirm" on confirmation display D12d, a tag of "tag1" (category "Damaged A") is assigned to the corresponding point in image D12b. When the user presses "Cancel," the assignment of "tag1" to the corresponding point is canceled.
[0033] (Effects of the First Embodiment) In the first embodiment described above, for each pair of images among a plurality of images, correspondence information relating to pixel correspondence between the images is calculated according to the resolution required for tagging. Then, a tag is assigned to a specified position of the image designated by the user, and corresponding points, which are pixels in other images that correspond to the pixels included in the specified position, are identified based on the correspondence information. Then, the same tag as that assigned to the pixel corresponding to the specified position is assigned to the identified corresponding point.
[0034] Therefore, according to embodiment 1, even in situations where there are only a small number of images or unknown objects are being identified, it is possible to obtain precise correspondence information between multiple images and efficiently assign information to the images.
[0035] Furthermore, according to the first embodiment, an algorithm is used to determine the pixel correspondence between images. Therefore, unlike the prior art typified by Patent Document 1, it is not necessary to prepare 3D correspondence information at enormous processing costs using a depth camera or 3D reconstruction from a large number of images. Therefore, it is possible to correctly identify the same object captured in multiple images at low processing cost and with high accuracy.
[0036] Furthermore, in this embodiment, since the correspondence information between pixels of a pair of two-dimensional images is obtained using an algorithm for determining pixel correspondence between images, such as LoFTR, there is no need to align the images using homography transformation or affine transformation. Therefore, precise correspondence information between pixels can be obtained even if the multiple images taken around the object are not images that connect distant views, such as a panorama.
[0037] [Embodiment 2] In the first embodiment, the efficiency of tagging is improved by assigning the same tag to points in other images that correspond to points included in tagged positions in an image to which a tag has been assigned.
[0038] On the other hand, there is a need to be able to check objects in multiple images while simultaneously viewing them. Therefore, in the second embodiment, when multiple images are displayed, if the mouse is placed over a point in one image, the corresponding point in the other images is displayed in the correspondence information, thereby improving the efficiency of the visual confirmation work.
[0039] In the following description of the second embodiment, the overlapping description with the first embodiment will be omitted, and the differences will be mainly described.
[0040] (Corresponding point display processing according to the second embodiment) 5 is a flowchart showing a corresponding point display process according to the second embodiment. First, in step S21, the tag assignment unit 13 accepts a user's designation of a specified position on an image from a display screen (an input / output screen D2 described below), for example, designation of the specified position by dragging the mouse. Next, in step S22, the tag assignment unit 13 refers to the correspondence information table T1 and determines, among the points within the specified position, points that have corresponding points in other images in the correspondence information. If there are multiple points that have corresponding points, an arbitrary point is determined.
[0041] Next, in step S23, the tagging unit 13 determines whether or not a corresponding point in another image exists for the point determined in step S22. If a corresponding point in another image exists for the point determined in step S22 (step S23 YES), the tagging unit 13 proceeds to step S24, and if not (step S23 NO), the corresponding point display process ends.
[0042] In step S24, the tagging unit 13 generates a list of other images having the corresponding points determined in step S22 (hereinafter referred to as a "corresponding point image list") based on the correspondence information table T1. The corresponding point image list is displayed on a display screen (an input / output screen D2 described later). Note that the corresponding point image list also includes images for which the designated positions were specified in step S21.
[0043] Next, in step S25, the tagging unit 13 accepts an image selection by the user from the list of corresponding point images generated in step S24. Next, in step S26, the tagging unit 13 highlights the corresponding points in the selected image selected in step S25 on the display screen (input / output screen D2 described below). That is, the pixels (corresponding points) of the image are distinguishably displayed on the display screen according to the reliability of the correspondence information for each pixel.
[0044] Next, in step S27, the tag assignment unit 13 determines whether the reliability of the corresponding point is lower than a threshold. If there are multiple corresponding points, the tag assignment unit 13 determines the reliability of each corresponding point based on the threshold. If there is a corresponding point with a reliability lower than the threshold (step S27 YES), the tag assignment unit 13 proceeds to step S28, and if there is no corresponding point with a reliability lower than the threshold (step S27 NO), the tag assignment unit 13 ends the corresponding point display process.
[0045] In step S28, the tag assignment unit 13 highlights on the display screen (input / output screen D2 described later) an area including corresponding points whose reliability determined in step S27 is less than a threshold. The tag assignment unit 13 may highlight a range consisting of points whose reliability of the correspondence with the corresponding points is equal to or greater than a threshold in the image including the specified position whose specification was accepted in step S21.
[0046] (Input / output screen D2 according to the second embodiment) 6 is a diagram showing the configuration of an input / output screen D2 according to embodiment 2. The input / output screen D2 has an image list display area D21 and an image display area D22. The image list display area D21 displays a corresponding point image list D21a generated in step S24 of the corresponding point display process (FIG. 5).
[0047] In the corresponding point image list D21a of the image list display area D21, the image names "image1" and "image2" are highlighted (hatched in FIG. 6). The image with the image name "image1" in the corresponding point image list D21a corresponds to the image whose position was specified in step S21 of the corresponding point display process (FIG. 5). The image with the image name "image2" in the corresponding point image list D21a corresponds to the image newly selected by the user in step S25 of the corresponding point display process.
[0048] In the image display area D22, an image D22a with the image name "image1" and an image D22d with the image name "image2" are displayed. An area D22b in the image D22a is an area where the reliability of the corresponding points of each point in the correspondence information table T1 is equal to or greater than a threshold. A specific point D22c is selected by the user from the area D22b. In this way, by displaying a transparent color mask, such as an overlay, for identification purposes on the area including pixels with correspondence information whose reliability is equal to or greater than a threshold, the user can select the appropriate pixel.
[0049] The corresponding point image list D21a is generated when a specific point D22c of the image D22a is selected in the image display area D22. By displaying the corresponding point image list D21a, the user can easily select an image having a point corresponding to the specific point D22c. In the initial state when the corresponding point image list D21a is generated, only the image name "image1" of the image D22a is highlighted.
[0050] Then, when the user selects an image name, for example, "image2", from the initial corresponding point image list D21a, the image name "image2" is switched to a highlighted display, and image D22d corresponding to image name "image2" is displayed. Then, corresponding point D22e in image D22d corresponding to specific point D22c in image D22a, and arrow D22f pointing from specific point D22c to corresponding point D22e are displayed. The display of specific point D22c and corresponding point D22e allows the user to easily recognize the pixel of the image to which a tag is to be added.
[0051] Furthermore, if the reliability of the corresponding point D22e is less than the threshold, a display D22g is displayed indicating that the reliability of the corresponding point D22e and / or the specific point D22c is less than the threshold.
[0052] (Effects of the second embodiment) In the above-described second embodiment, a reliability indicating the accuracy of the correspondence information for each pixel of correspondence information between images is calculated, and the image is displayed on a display screen together with an identification display of a predetermined area of the image based on the reliability of the correspondence information. Furthermore, if the reliability of the correspondence information is below a threshold, a message indicating that the reliability is below the threshold is displayed on the display screen. Therefore, according to the second embodiment, a user can recognize the reliability of the corresponding points and select appropriate pixels from pixels having correspondence information with a reliability equal to or higher than the threshold, thereby improving the accuracy of tagging when automatically assigning tags to corresponding points in other images.
[0053] [Embodiment 3] In the above-described first and second embodiments, tags are assigned to points on an image, and corresponding points in the correspondence information are displayed on the image. However, there is a need to search for points on an image to which corresponding tags have been assigned using tag information as a clue. Therefore, in the third embodiment, a list of tag information is displayed, and points on the image that correspond to the tag information are searched for.
[0054] In the following description of the third embodiment, the overlapping description with the first and second embodiments will be omitted, and the differences will be mainly described.
[0055] (Tag search process according to the third embodiment) FIG. 7 is a flowchart showing a tag search process according to the third embodiment.
[0056] First, in step S31, the tagging unit 13 accepts a user's selection of a group of images. Next, in step S32, the tagging unit 13 refers to the tag information table T2 (FIG. 2B) and tallies the tags assigned to each image in the group of images selected in step S31. Tags can be free-form tags or tags that represent categories. Category tags can be tallied by category, for example, "Damaged: 3, Defaced: 2, etc."
[0057] Next, in step S33, the tag assigning unit 13 displays the list of tags and the counting results on a display screen (an input / output screen D3, which will be described later).
[0058] Next, in step S34, the tagging unit 13 accepts a tag search by the user and executes a search for tags that match the search criteria. In step S34, for example, a category tag is selected and the search is narrowed down to tags in that category. Alternatively, in step S34, for free-text tags, a character string within the tag is searched for an arbitrary word and the search is narrowed down to tags that contain that word.
[0059] Next, in step S35, the tag assignment unit 13 displays a list of tags narrowed down by the search executed in step S34 on a display screen (input / output screen D3 described below). Next, in step S36, the tag assignment unit 13 accepts a tag selected by the user from the list of tags displayed in step S35. Next, in step S37, the tag assignment unit 13 displays a list of images linked to the tag selected in step S36 on a display screen (input / output screen D3 described below) based on tag information table T2.
[0060] (Input / output screen D3 according to the third embodiment) 8 is a diagram showing the configuration of an input / output screen D3 according to embodiment 3. The input / output screen D3 has an image list display area D31, an image display area D32, and a list display area D33.
[0061] In the image list D31a of the image list display area D31, the image names "image1", "image2", and "image3" are highlighted (hatched in FIG. 8). The image names "image1" to "image3" in the image list D31a correspond to the images selected in step S31 of the tag search process (FIG. 7).
[0062] The list display area D33 displays a list display D33a in which tags assigned to each image "image1" to "image3" selected from the image list D31a are aggregated by category or keyword in step S32 of the tag search process (FIG. 7).
[0063] In addition, when the user inputs a keyword into a search box D33b and presses a search button D33c in the list display area D33, the search of step S34 of the tag search process is executed, and a tag list display D33d is displayed. The tag list display D33d displays the results narrowed down to tags that have categories or free words that match the keyword.
[0064] The image display area D32 displays tagged images showing a list of images to which tags selected by the user from the tag list display D33d have been added. In the example of Fig. 8, a tagged image D32a with the image name "image1" and a tagged image D32b with the image name "image2" are displayed.
[0065] Also displayed in the image display area D32 are a tag D32c selected by the user from the tag list display D33d, a point D32d in the tagged image D32a to which the tag D32c is assigned, and a point D32e in the tagged image D32b.
[0066] That is, in the third embodiment, tag information relating to tags assigned to a group of images among a plurality of images designated by a user is aggregated. A list of the aggregated tag information is then displayed on a display screen. Images to which tags relating to the tag information designated by the user are assigned are then searched for, and a list of images to which the tags are assigned is displayed on the display screen together with the searched tags.
[0067] (Effects of the third embodiment) In the above-described third embodiment, images tagged with tags that match the category and text information input by the user, the tagged positions, and the tags themselves are displayed on the display screen, allowing the user to narrow down to the desired tag, efficiently check the tags, and add new tags.
[0068] [Embodiment 4] In the above-described first to third embodiments, tags are displayed on two-dimensional images. However, objects to which tags are attached are often three-dimensional objects, and there is a need to check the tag attachment status in three dimensions.
[0069] Therefore, in the fourth embodiment, when a three-dimensional image can be reconstructed from a certain number of two-dimensional images using a three-dimensional reconstruction method such as SfM (Structure from Motion), the three-dimensional image is reconstructed from the two-dimensional images and displayed, and the tags attached to the two-dimensional images are displayed on the three-dimensional image. The three-dimensional image is a mesh or a three-dimensional point cloud.
[0070] In this embodiment, the number of images for each structure stored in the image data storage unit 15 is sufficient to enable a three-dimensional image to be reconstructed.
[0071] In the following description of the fourth embodiment, the overlapping description with the first to third embodiments will be omitted, and the differences will be mainly described.
[0072] (Configuration of information processing device according to embodiment 4) 9 is a diagram showing the configuration of an information processing system 1D according to embodiment 4. The information processing system 1D differs from the information processing system 1 according to embodiment 1 in that it further includes a 3D reconstruction unit 18 implemented by a CPU that executes a program. The information processing system 1D also differs from the information processing system 1 according to embodiment 1 in that it further includes a 2D-3D correspondence information storage unit 19 and a 3D image data storage unit 20 that are provided in predetermined storage areas.
[0073] The 3D reconstruction unit 18 constructs a 3D image from multiple 2D images of the target structure stored in the image data storage unit 15. Pixels in the tagged images are pixels that have a correspondence relationship between images stored in a correspondence information table T1. Multiple pixels with such a correspondence relationship are converted into a single point in 3D space using a 3D reconstruction method such as SfM described in Non-Patent Document 2.
[0074] Therefore, tag information added to a 2D image can be displayed on a 3D image. Furthermore, a dense 3D image is reconstructed using MVS (Multi-View Stereo). The 3D reconstruction unit 18 manages the correspondence between each coordinate of the 2D image reconstructed in this way and the coordinate of the 3D image. The user can grasp the position of the tag in three dimensions in the 3D image, and when checking the details of the object, they can refer to the high-resolution 2D image, allowing them to efficiently observe the target structure.
[0075] (2D-3D Correspondence Information Table T3 According to the Fourth Embodiment) 10 is a diagram showing the configuration of a 2D-3D correspondence information table T3 according to embodiment 4. The 2D-3D correspondence information table T3 is stored in the 2D-3D correspondence information storage unit 19. The 2D-3D correspondence information table T3 manages, as coordinate correspondence information for each target structure, coordinates of multiple 2D images and coordinates of 3D images corresponding to each coordinate.
[0076] (3D display processing according to the fourth embodiment) 11 is a flowchart showing a three-dimensional display process according to the fourth embodiment. The three-dimensional display process according to the fourth embodiment is executed at any timing designated by the user. When the three-dimensional display process is executed, it is assumed that the user's images received by the image input unit 11 are stored in the image data storage unit 15. Furthermore, when the three-dimensional display process is executed, it is assumed that correspondence information for the images stored in the image data storage unit 15 has been generated and stored in the correspondence information table T1, and that tags have been assigned to image pairs having the correspondence information.
[0077] First, in step S41, the 3D reconstruction unit 18 accepts the user's selection of a group of 2D images displayed on a display screen (input / output screen D4, described later). Next, in step S42, the 3D reconstruction unit 18 performs 3D reconstruction using SfM, with the group of images selected in step S41 as input. When SfM is executed, the correspondence between the coordinates of corresponding points in the 2D images and the coordinates of the 3D images is obtained.
[0078] Next, in step S43, the three-dimensional reconstruction unit 18 acquires three-dimensional coordinates corresponding to the coordinates of the tag of each image in the image group from the SfM output.
[0079] Next, in step S44, the 3D reconstruction unit 18 performs 3D reconstruction using MVS with the image group selected in step S41 as input, and stores the reconstructed 3D image in the 3D image data storage unit 20. A dense mesh or a 3D point cloud is output by MVS as the 3D image.
[0080] Next, in step S45, the 3D reconstruction unit 18 assigns a tag whose 3D coordinates were acquired in step S43 to the output of the MVS in step S44. Next, in step S46, the 3D reconstruction unit 18 displays the output of the MVS in step S45 and the tag on a display screen (input / output screen D4 described below).
[0081] Next, in step S47, the 3D reconstruction unit 18 accepts a user's selection of an image from the group of 2D images selected in step S41. Next, in step S48, the 3D reconstruction unit 18 projects the tag assigned to the 3D image in step S46 onto the coordinates of the 2D image selected in step S47 on the display screen (input / output screen D4 described below).
[0082] Next, in step S49, the 3D reconstruction unit 18 displays the tag in the image selected in step S47, and also displays the 3D occlusion state of the tag. In step S49, if a 3D position in the 3D image corresponding to a 2D position in the image group is occluded in the 2D display of the 3D image on the display screen (input / output screen D4 described below), the 3D reconstruction unit 18 displays that the tag is occluded. "Occluded" means that the tag is present on the side of a structure different from the display side in the 2D display of the 3D image.
[0083] In step S49, if the output of the MVS in step S44 is a mesh, it is possible to determine whether each tag is occluded when projecting its coordinates onto the coordinates of the selected image. This allows the occlusion status of the tag to be displayed in the 3D image. On the other hand, if the output of the MVS in step S44 is a 3D point cloud, the occlusion status is left undefined.
[0084] (Input / output screen D4 according to the fourth embodiment) 12A and 12B are diagrams showing the configuration of an input / output screen D4 according to embodiment 4. The input / output screen D4 has an image list display area D41, an image display area D42, and a 3D reconstruction button D43.
[0085] In the image list D41a of the image list display area D41, the image names "image1", "image2", and "image3" are highlighted (hatched in FIG. 12A). The image names "image1" to "image3" in the image list D41a correspond to the group of images selected in step S41 of the 3D display process (FIG. 11).
[0086] The image display area D42 displays an image selected by the user from the image list D41a. In the example of Fig. 12A, an image D42a with the image name "image1" and an image D42b with the image name "image2" are displayed.
[0087] When a group of images is selected from the image list D41a and the 3D reconstruction button D43 is pressed, the 3D display process (FIG. 11) is executed. Then, based on the group of images selected from the image list D41a, a 3D image D44 as shown in FIG. 12B is displayed. The 3D image D44 displays tags D44a and D44b at three-dimensional positions of the 3D image D44 that correspond to the two-dimensional positions of the group of images ("image1" to "image3") selected from the image list D41a, as specified based on the 2D-3D correspondence information table T3.
[0088] Although not shown in the figure, if the output of the MVS in step S44 is a mesh, when the user selects a two-dimensional image from the group of images (three images "image1" to "image3") selected from the image list D41a, the selected two-dimensional image is displayed. Then, tags D44a and D44b assigned to the mesh in step S46 are projected onto the selected two-dimensional image. Furthermore, tags whose three-dimensional positions in the mesh output from the MVS are obstructed are displayed together with an obstruction notification display indicating that they are obstructed. Tags whose three-dimensional positions are obstructed and displayed together with an obstruction notification display are also displayed together with an obstruction notification display when projected from the mesh onto the two-dimensional image.
[0089] (Effects of the fourth embodiment) In the above-described fourth embodiment, a three-dimensional image such as a three-dimensional point cloud or mesh is reconstructed from a group of images selected from a plurality of images using a three-dimensional reconstruction technique. Two-dimensional-three-dimensional correspondence information relating to the correspondence between the two-dimensional positions of tags in the group of images and the three-dimensional positions of tags in the three-dimensional images is then calculated. Based on the two-dimensional-three-dimensional correspondence information, three-dimensional positions in the three-dimensional images corresponding to the two-dimensional positions in the group of images are identified, and the tags are arranged at the identified three-dimensional positions in the three-dimensional images and displayed together with the three-dimensional images on a display screen.
[0090] Therefore, according to embodiment 4, the same tag assigned to corresponding points at two-dimensional positions on multiple two-dimensional images is displayed at a three-dimensional position on a single three-dimensional image, making it easy to confirm and understand the position of the tag.
[0091] Furthermore, in the above-described fourth embodiment, when the three-dimensional image is a mesh, if a three-dimensional position in the three-dimensional mesh image corresponding to a two-dimensional position of the image group is occluded in the display of the three-dimensional mesh image on the display screen, a message indicating that the position is occluded is displayed. Thus, according to the third embodiment, tags that are occluded and do not appear on the display side in the three-dimensional display and their assigned positions can be easily confirmed.
[0092] (Hardware configuration of computer 100) 13 is a diagram showing the hardware configuration of the computer 100. By executing a predetermined program, the computer 100 realizes the information processing systems 1 and 1D and a terminal communicably connected to the information processing systems 1 and 1D via a network.
[0093] The computer 100 has a CPU 101, which is an example of a processor, which are interconnected via an internal communication line 107 such as a bus, a RAM (Random Access Memory) 102, which is an example of a main memory device, a GPU (Graphics Processing Unit) 103, which is an example of a coprocessor of the CPU 101, an SSD (Solid State Drive) 104, which is an example of an auxiliary memory device, an input / output I / F (Interface) 105, and a communication I / F 106.
[0094] The CPU 101 controls the overall operation of the computer 100. The RAM 102 functions as a work memory for the CPU 101. The GPU 103 performs some of the processing of the CPU 101, such as graphics processing. The SSD 104 is a large-capacity nonvolatile storage device used to store various programs and data for long periods of time.
[0095] The executable programs stored in the SSD 104 are loaded into the RAM 102 when the computer 100 is started up or when necessary, and are executed by the CPU 101. In this way, the various processing function units of the information processing systems 1 and 1D are realized.
[0096] The executable program stored in SSD 104 may be recorded on a non-transitory recording medium, read from the non-transitory recording medium by a media reading device, and loaded into RAM 102. Alternatively, the executable program may be obtained from an external computer via a network and loaded into RAM 1002.
[0097] The SSD 104 stores executable programs that realize the processing function units of the information processing systems 1 and 1D.
[0098] The input / output I / F 105 includes an interface device for connecting an input device and an output device. The input device connected to the computer 100 via the input / output I / F 105 is composed of a keyboard, a pointing device such as a mouse, and the like, and is used by the user to input various instructions and information to the computer 100.
[0099] In addition, the output device connected to the computer 100 via the input / output I / F 105 is composed of, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display, or an audio output device such as a speaker, and is used to present necessary information to the user when necessary.
[0100] The communication I / F 106 includes a communication interface device for connecting the computer 100 to each network within the system or for communicating with other computers. The communication I / F 106 includes, for example, a network interface card (NIC) for a wired local area network (LAN) or a wireless LAN.
[0101] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, part or all of the configuration of one embodiment may be combined with part or all of the configuration of another embodiment to the extent that they are not inconsistent. Furthermore, part of the configuration of each embodiment may be added to, deleted from, or replaced with the configuration of another embodiment. [Explanation of symbols]
[0102] 1, 1D: information processing system, 12: correspondence information calculation unit, 13: tag assignment unit, 14: input / output unit, 15: image data storage unit, 16: correspondence information storage unit, 17: tag information storage unit, 18: 3D reconstruction unit, 19: 2D-3D correspondence information storage unit, 20: 3D image data storage unit, 100: computer, 101: CPU, 102: RAM.< / url> < / url>
Claims
1. An information processing method executed by an information processing system that assigns tags to images, comprising: the information processing system includes a processor and a memory; the processor: Accepts input of multiple images of an object taken from multiple viewpoints; calculating, for each pair of images among the plurality of images, correspondence information relating to pixel correspondence between the images according to a resolution required for assigning the tag; assigning the tag to a designated position on the image designated by a user; Based on the correspondence information, a corresponding point is identified, which is a pixel of another image corresponding to a pixel included in the specified position; The specified corresponding point is assigned the same tag as the tag assigned to the pixel corresponding to the specified position. An information processing method characterized by comprising each process.
2. 2. The information processing method according to claim 1, a storage unit for storing tag information relating to the tag assigned to the image; the processor: Accepting input of a new image of the object; calculating the correspondence information for each image pair between the new image and any one of the plurality of images; Identifying pixels of any one of the plurality of images that correspond to pixels of the new image based on the correspondence information; Based on the tag information, a tag identical to the tag assigned to a pixel of any one of the plurality of images corresponding to the pixel of the new image is assigned to the pixel of the new image. An information processing method characterized by comprising each process.
3. 2. The information processing method according to claim 1, the processor: calculating a reliability indicating the likelihood of the correspondence information for each of the correspondence information of the pixels between the images; The pixels of the image are displayed on a display screen in a distinguishable manner according to the reliability of the corresponding information for each pixel. An information processing method characterized by comprising each process.
4. 4. The information processing method according to claim 3, the processor: and displaying, on the display screen, a visual indication that the reliability of the corresponding information of the corresponding point of the pixel is less than a threshold.
2. An information processing method comprising:
5. 5. The information processing method according to claim 4, the processor: searching for another image having the corresponding point corresponding to the designated position of the image designated by the user based on the corresponding information; A list of the other images found is displayed on the display screen. An information processing method characterized by comprising each process.
6. 6. The information processing method according to claim 5, the processor: Identifying the corresponding point of another image that corresponds to the designated position of the image selected by the user from the list of images based on the correspondence information; The identified corresponding points of the other images are displayed on the display screen. An information processing method characterized by comprising each process.
7. 7. The information processing method according to claim 6, the processor: aggregating tag information related to the tags assigned to a group of images among the plurality of images designated by a user; Displaying a list of the collected tag information on the display screen; Searching for the image to which the tag related to the tag information designated by the user has been assigned; The searched tag and a list of the images to which the tag is assigned are displayed on the display screen. An information processing system having each process.
8. 2. The information processing method according to claim 1, the processor: reconstructing a three-dimensional image from a group of images selected from the plurality of images using a predetermined three-dimensional reconstruction technique; calculating two-dimensional-three-dimensional correspondence information relating to the correspondence between the two-dimensional position of the tag in the image group and the three-dimensional position of the tag in the three-dimensional image; identifying the three-dimensional position in the three-dimensional image corresponding to the two-dimensional position in the image group based on the two-dimensional-three-dimensional correspondence information; The tag is arranged at the specified three-dimensional position in the three-dimensional image and displayed together with the three-dimensional image on a display screen. An information processing method characterized by comprising each process.
9. 9. The information processing method according to claim 8, the processor: When the three-dimensional position in the three-dimensional image corresponding to the two-dimensional position in the image group is obstructed in the display of the three-dimensional image on the display screen, a message indicating that the position is obstructed is displayed.
2. An information processing method comprising:
10. An information processing system for tagging images, comprising: the information processing system includes a processor and a memory; The processor: Accepts input of multiple images of an object taken from multiple viewpoints; calculating, for each pair of images among the plurality of images, correspondence information relating to pixel correspondence between the images according to a resolution required for assigning the tag; assigning the tag to a designated position on the image designated by a user; Based on the correspondence information, a corresponding point is identified, which is a pixel of another image corresponding to a pixel included in the specified position; The specified corresponding point is assigned the same tag as the tag assigned to the pixel corresponding to the specified position. An information processing system comprising:
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
JP2023085999A