Information processing method, information processing device, and information processing program

JPWO2025095033A1Undetermined Publication Date: 2025-05-08
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025555028
Authority / Receiving Office
JP · JP
Patent Type
Applications
Priority Date
2024-08-09
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately set the travel time from the reference point to each predetermined point, resulting in the inability to effectively extract the moving image.

Method used

By displaying a bird's-eye view on the display interface of the terminal device and displaying multiple camera point icons, when the user selects a first camera point icon, the moving image associated with the camera point is displayed, and the corresponding moving image is displayed at the front and rear camera points.

Benefits of technology

It realizes effective grasp of the dynamic situation from the selected camera points to the front and rear camera points, and solves the accuracy problem of users when setting driving time.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This information processing method is used in a computer and includes: displaying a bird's-eye view on a display of a terminal device; displaying, on the bird's-eye view, a plurality of imaging point icons indicating a series of multiple imaging points; and upon detection of a selection of one first imaging point icon among the plurality of imaging point icons, displaying a first moving image linked to a first imaging point indicated by the first imaging point icon. The first moving image is a moving image between imaging times at each of two second imaging points indicated by two second imaging icons displayed before and after the first imaging point icon.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and information processing program

[0001] The present disclosure relates to a technique for displaying images captured at multiple capture locations.

[0002] Patent Document 1 discloses a technology for cutting out and displaying video of a set time before and after the time a utility pole passes from video recorded on a drive recorder installed in a vehicle, based on the time the pole passes.

[0003] In the technology of Patent Document 1, in order to extract video from footage recorded in a drive recorder, the user needs to set the time before and after the time of passing a reference point. Therefore, even if the user tries to extract video between specific points before and after the reference point, it is difficult to accurately set the travel time from the reference point to each specific point, which causes a problem in that the video cannot be extracted appropriately.

[0004] Japanese Patent Application Laid-Open No. 2023-144527

[0005] The present disclosure has been made to solve such problems, and aims to provide a technology that can properly grasp the situation from a selected shooting location to the two shooting locations before and after it.

[0006] An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes displaying a bird's-eye view on a display of a terminal device, displaying a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view, and when selection of a first shooting location icon from among the plurality of shooting location icons is detected, displaying a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between the shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

[0007] FIG. 1 is an overall configuration diagram of an information processing system. FIG. 2 is a diagram showing an example of a design drawing displayed on the display of an information terminal. FIG. 3 is a diagram showing an example of a screen displayed when a first shooting location icon is selected. FIG. 4 is a diagram showing an example of a display of a first video. FIG. 5 is a flowchart showing a first example of processing of the information processing system performed during shooting operation. FIG. 6 is a flowchart showing an example of peripheral video display processing. FIG. 7 is a flowchart showing a second example of processing of the information processing system performed during shooting operation. FIG. 8 is an explanatory diagram of a modified example of a method for creating partial videos corresponding to each shooting location.

[0008] (Background to the present disclosure) A user interface is being developed that displays, on a two-dimensional blueprint of a construction site or the like, multiple photography location icons indicating a series of multiple photography locations where the construction site was actually photographed, and when one photography location icon is selected, displays an image photographed at the photography location indicated by that one photography location icon. This allows a construction site manager to understand the situation at the construction site without visiting the construction site.

[0009] Furthermore, a function is being considered for such a user interface that displays not only the image captured at the selected shooting location but also video footage of the two shooting locations before and after the selected shooting location. This allows the supervisor to properly grasp the progress of the construction work not only at the selected shooting location but also from the selected shooting location to the two shooting locations before and after it. With the technology of Patent Document 1, it is difficult for the user to accurately set the travel time from the selected shooting location to the two shooting locations before and after it, making it difficult to realize the above function.

[0010] Therefore, the inventors have conducted extensive research into technology that can properly grasp the situation from a selected shooting location to the two shooting locations before and after it, and have come up with the present disclosure described below.

[0011] (1) An information processing method in one aspect of the present disclosure is an information processing method in a computer, which includes displaying a bird's-eye view on a display of a terminal device, displaying a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view, and, upon detecting selection of a first shooting location icon from among the plurality of shooting location icons, displaying a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between the shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

[0012] According to this configuration, video images are displayed between the shooting times at the two second shooting locations before and after the selected first shooting location, allowing the viewer to properly grasp the situation from the selected first shooting location to the two second shooting locations before and after it.

[0013] (2) In the information processing method described in (1) above, when the selection of the first shooting location icon is detected, the display mode of the two second shooting location icons displayed on the bird's-eye view may be made different from the display mode of the other shooting location icons.

[0014] In this configuration, when selection of a first image capture location icon is detected, the display mode of the two second image capture location icons displayed on the bird's-eye view will be different from the display mode of the other image capture location icons, so that the user can easily see by looking at the bird's-eye view which two image capture locations are displaying videos between the image capture times.

[0015] (3) In the information processing method described in (1) or (2) above, displaying the first video may include displaying time information that associates the shooting time of the currently displayed image included in the first video with the shooting time at the first shooting location and the shooting time at each of the two second shooting locations.

[0016] In this case, the time information is displayed, so that the user can understand the relative positional relationship between the shooting location of the currently displayed image included in the first video and each of the first shooting location and the two second shooting locations.

[0017] (4) In the information processing method described in (3) above, displaying the time information may include displaying the shooting time at the first shooting location using an icon in the same display format as the first shooting location icon displayed on the bird's-eye view, and displaying the shooting time at each of the two second shooting locations using two icons in the same display format as each of the two second shooting location icons displayed on the bird's-eye view.

[0018] In this case, the shooting time at the first shooting location included in the time information is displayed using an icon with the same display style as the first shooting location icon displayed on the bird's-eye view. Also, the shooting times at each of the two second shooting locations included in the time information are displayed using two icons with the same display style as the two second shooting location icons displayed on the bird's-eye view. This allows the user to intuitively grasp the relative positional relationship between the shooting location of the currently displayed image included in the first video and each of the first and two second shooting locations.

[0019] (5) In the information processing method described in any one of (1) to (4) above, the method may further include extracting a first partial video, which is a video between the shooting times at each of the shooting locations before and after each shooting location, from an overall video, which is a video between the shooting times at each of the start and end points of shooting at the multiple shooting locations, and storing the first partial video in association with each shooting location, and displaying the first video may include displaying the first partial video corresponding to the first shooting location as the first video.

[0020] With this configuration, when selection of the first photographing location icon is detected, the first partial video stored in association with the first photographing location is displayed as the first video. Therefore, the first image can be displayed more quickly when the first photographing location icon is selected than when videos between the photographing times at each of the two second photographing locations indicated by the two second photographing icons displayed before and after the first photographing location icon are extracted from the full video and displayed each time selection of the first photographing location icon is detected.

[0021] (6) In the information processing method described in any one of (1) to (4) above, the method further includes extracting a second partial video, which is a video from a first average time, which is the average of the shooting times at each shooting location before each shooting location and each shooting location, to a second average time, which is the average of the shooting times at each shooting location and each shooting location after each shooting location, from an overall video, which is a video between the shooting times at each of the start and end points of shooting at the multiple shooting locations, and storing the second partial video in association with each shooting location, and displaying the first video may include displaying as the first video a video that is a combination of a video of the second partial video corresponding to a third shooting location, which is a second shooting location before the first shooting location, a video of the second partial video from after the shooting time at the third shooting location, the second partial video corresponding to the first shooting location, and a video of the second partial video corresponding to a fourth shooting location, which is a second shooting location after the first shooting location, before the shooting time at the fourth shooting location.

[0022] In this case, second partial videos, which are videos from the first average time to the second average time, are stored in association with each shooting location. Then, a video synthesized using the second partial videos associated with each of the third shooting location, the first shooting location, and the fourth shooting location is displayed as the first video. Therefore, second partial videos corresponding to each shooting location required to compose the first video can be stored without overlapping with second partial videos corresponding to other shooting locations.

[0023] (7) In the information processing method described in any one of (1) to (4) above, a third partial video, which is a video between the shooting times at each shooting location and the shooting location immediately after each shooting location, is extracted from an entire video, which is a video between the shooting times at each shooting start location and the shooting end location of the plurality of shooting locations; and when it is detected that the third partial video contains a photographed image of a predetermined object, an object video, which is a video consisting only of images including the photographed image of the predetermined object and has a period overlapping with the third partial video, is extracted from the entire video, and the third partial video and the object video are combined. and storing a fourth partial video, which is a video obtained by combining a photographed image of the specified object with the third partial video corresponding to each shooting location so that their periods do not overlap, in association with each shooting location, and if it is detected that the third partial video does not include a photographed image of the specified object, storing the third partial video in association with each shooting location; and displaying the first video may include displaying as the first video a video obtained by combining the third partial video or the fourth partial video corresponding to the second shooting location immediately before the first shooting location with the third partial video or the fourth partial video corresponding to the first shooting location so that their periods do not overlap.

[0024] In this configuration, a video obtained by combining the third partial video or the fourth partial video corresponding to the second imaging location immediately before the first imaging location with the third partial video or the fourth partial video corresponding to the first imaging location so that their periods do not overlap is displayed as the first video. The fourth partial video is a video obtained by combining the third partial video and the object video so that their periods do not overlap.

[0025] Therefore, the user can view a video including images of the predetermined object taken not only during the period between the image capture times at the two second image capture locations, one before and one after the first image capture location, but also during the period immediately before and / or after the first image capture location, thereby allowing the user to focus on the predetermined object.

[0026] (8) In the information processing method described in (7) above, detecting whether or not the third partial video contains a photographed image of the specified object includes detecting that the third partial video contains a photographed image of the specified object when the one or more photographed times obtained by inputting the entire video into a model that has machine-learned the relationship between a second video, which is a video shot in a space including the multiple shooting locations and includes one or more photographed images of an annotation object present in the space, and each of the photographed times of the one or more photographed images of the annotation object included in the second video, are included within the period of the third partial video, and the annotation object may be an object designated by a user as a target to be annotated.

[0027] In this case, the user can view a video including images of the annotation object captured not only during the time period between the two second image capturing locations before and after the first image capturing location, but also during the time period immediately before and / or after the first image capturing location, allowing the user to focus on the object that can be designated as the target for annotating.

[0028] (9) In the information processing method described in any one of (1) to (6) above, among the plurality of photographing locations, a photographing location before one of the photographing locations is a photographing location that is n locations before the one of the photographing locations, and a photographing location after the one of the photographing locations is a photographing location that is m locations after the one of the photographing locations, and the n and the m may be natural numbers greater than or equal to 1.

[0029] In this case, video images are displayed between the shooting times at each of the second shooting locations that are n locations before and m locations after the selected first shooting location, allowing the user to properly grasp the situation from the selected first shooting location to the two second shooting locations that are n locations before and m locations after the selected first shooting location.

[0030] Furthermore, the present disclosure can be realized not only as an information processing method that executes the characteristic processes described above, but also as an information processing device or the like that has a characteristic configuration corresponding to the characteristic processes executed by the information processing method. Furthermore, the present disclosure can also be realized as a computer program that causes a computer to execute the characteristic processes included in such an information processing method. Therefore, the same effects as those of the above information processing method can also be achieved in the following other aspects.

[0031] (10) In another aspect of the present disclosure, an information processing device includes a processor, wherein the processor executes the following operations: displaying a bird's-eye view on a display of a terminal device; displaying a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view; and, upon detecting selection of a first shooting location icon from among the plurality of shooting location icons, displaying a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between the shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

[0032] (11) In yet another aspect of the present disclosure, an information processing program causes a computer to display a bird's-eye view on a display of a terminal device, display a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view, and, upon detecting selection of a first shooting location icon from among the plurality of shooting location icons, display a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between the shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

[0033] The present disclosure can also be realized as an information processing system operated by such an information processing program. Needless to say, such a computer program can be distributed on a computer-readable non-transitory recording medium such as a CD-ROM or via a communication network such as the Internet.

[0034] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components. Furthermore, in all embodiments, the respective contents can be combined.

[0035] 1 is an overall configuration diagram of an information processing system 1. The information processing system 1 is a system that extracts and displays videos between the shooting times at two shooting locations before and after a selected shooting location from videos between the shooting times at the first and last shooting locations in a series of multiple shooting locations.

[0036] 1, the information processing system 1 includes a server 10, an information terminal 20 (terminal device), an imaging device 30, and a communication device 40. The server 10 (information processing device and computer), the information terminal 20, and the communication device 40 are connected to each other so as to be able to communicate with each other via a network NT. An example of the network NT is the Internet.

[0037] The server 10 is, for example, a cloud server configured with one or more computers. However, this is just one example, and the server 10 may be configured as an edge server or may be implemented in the information terminal 20. The aspect in which the server 10 is implemented in the information terminal 20 is an example of the aspect in which the information terminal 20 is configured as an information processing device.

[0038] The information terminal 20 is carried by a user. The user is, for example, a manager of a predetermined space in which the image capturing device 30 captures images. The predetermined space is, for example, a construction site. However, this is just one example, and the predetermined space may be a construction site, factory, store, office, etc. The information terminal 20 may be configured as a portable computer such as a smartphone or tablet computer, or as a stationary computer. The information terminal 20 displays various images and screens on the display in accordance with display instructions from the server 10. Although one information terminal 20 is illustrated in the example of FIG. 1, multiple information terminals may be connected to the server 10 via the network NT. The information terminal 20 includes a communication unit 21, a processor 22, a display 23, and an operation unit 24.

[0039] The communication unit 21 is a communication interface that connects the information terminal 20 to the network NT. The communication unit 21 transmits various instructions received from the user by the operation unit 24 to the server 10. The communication unit 21 receives various image and screen display instructions from the server 10.

[0040] The processor 22 is configured by, for example, a central processing unit, and displays various images and screens on the display 23 in accordance with display instructions for various images and screens received by the communication unit 21 .

[0041] The display 23 is configured with various display devices such as a liquid crystal display, an organic EL (Electro-Luminescence) display, etc., and displays various images and screens under the control of the processor 22.

[0042] The operation unit 24 is composed of a keyboard, a touch panel, a mouse, etc., and accepts various instructions input by the user.

[0043] The communication device 40 is configured as a mobile information terminal such as a smartphone or a tablet computer, and connects the photographing device 30 to the network NT. The communication device 40 and the photographing device 30 are connected via a short-range wireless communication path such as Bluetooth (registered trademark).

[0044] The image capturing device 30 is, for example, an omnidirectional camera that captures video at a predetermined frame rate. An omnidirectional camera is also called a 360-degree camera, and is a camera that can capture images in all directions of 360 degrees. The image capturing device 30 is, for example, a portable image capturing device carried by a photographer. The photographer is, for example, a worker or site supervisor at a construction site. The image capturing device 30 may also be a regular camera.

[0045] The photographer moves through the construction site while photographing the construction site with the camera device 30. At the starting point of the photographing, the photographer presses the photographing button on the camera device 30, pointing it in a direction that uses the photographing direction of the camera device 30 as a reference (hereinafter referred to as the reference direction). In the following description, the reference direction is assumed to be north. However, the reference direction is not limited to north and may be another direction. When the photographer reaches the end point of the photographing, he presses the photographing button again. This causes the camera device 30 to end photographing. When the photographing ends, the camera device 30 transmits the photographed video to the communication device 40.

[0046] The communication device 40 receives the video transmitted by the camera device 30. In this case, the communication device 40 displays a blueprint of the construction site and accepts the cameraman's specification of the start and end positions of the video shooting on the blueprint. The start and end positions of the video shooting are specified by the cameraman inputting instructions specifying the positions on the blueprint screen of the construction site displayed on the display of the communication device 40.

[0047] Two-dimensional coordinate axes are defined on the blueprint screen. For example, the east-west direction on the blueprint is defined as the horizontal axis or X-axis, and the north-south direction is defined as the vertical axis or Y-axis. Therefore, the shooting start point and shooting end point are each defined by two-dimensional coordinate values. The communication device 40 transmits shooting information including the shot video to the server 10 via the network NT.

[0048] The shooting information is generated each time a shooting operation is performed. One shooting operation refers to a series of operations performed by a worker holding the camera device 30 at a construction site from the start to the end of shooting. The shooting information includes video captured in one shooting operation between the shooting times at the shooting start point and the shooting end point, and meta information for the video. The meta information includes the shooting ID, shooting start time, shooting end time, blueprint ID, shooting start point, and shooting end point.

[0049] The shooting ID is an identifier for identifying the shooting operation. The shooting start time is the shooting time at the shooting start point. The shooting end time is the shooting time at the shooting end point. The shooting time is obtained, for example, by a clock provided in the shooting device 30. The shooting time may include the year, month, and day when the shooting was performed. Similarly, various times in the following description may include the year, month, and day. The blueprint ID is an identifier for identifying the blueprint of the specified space where the video was shot. The shooting start point and shooting end point are used to identify the shooting location and shooting direction of each image that makes up the video.

[0050] The server 10 includes a processor 11, a memory 12, and a communication unit 13. The processor 11 is configured, for example, by a central processing unit (CPU). The processor 11 includes an instruction receiving unit 111 and a display control unit 112. The instruction receiving unit 111 and the display control unit 112 may be realized by the processor 11 executing an information processing program, or may be configured by a dedicated hardware circuit such as an ASIC. The information processing program may be recorded on a non-transitory computer-readable recording medium.

[0051] The communication unit 13 is a communication interface that connects the server 10 to the network NT. The communication unit 13 receives shooting information transmitted by the communication device 40. The communication unit 13 receives various instructions from the user from the information terminal 20. The communication unit 13 transmits various image and screen display instructions to the information terminal 20.

[0052] The memory 12 is configured as a non-volatile rewritable storage device such as a hard disk drive or a solid state drive, etc. The memory 12 includes a design drawing information storage unit 121, an image information storage unit 122, and a video information storage unit 123.

[0053] The blueprint information storage unit 121 stores blueprint information. Blueprint information is information that associates a blueprint ID with a blueprint of a specific space identified by the blueprint ID. A blueprint is a drawing that shows the design of a construction site, and may be a floor plan, blueprint, map, perspective drawing, etc. of the construction site. A blueprint is an example of a bird's-eye view. A bird's-eye view can also be called a bird's-eye view, and may be a view from above or a view from a high place.

[0054] The image information storage unit 122 stores image information. The image information is information that associates multiple images included in a video captured by the imaging device 30 with meta information for each image. The meta information includes a shooting ID, shooting time, blueprint ID, shooting location, and shooting direction. Specifically, the processor 11 creates image information based on the shooting information received by the communication unit 13, and stores the created image information in the image information storage unit 122.

[0055] More specifically, the processor 11 acquires the shooting start point and shooting end point from meta information included in the shooting information received by the communication unit 13. The processor 11 uses Visual SLAM (Simultaneous Localization and Mapping) technology to identify the shooting start point, shooting end point, and shooting point of each image from multiple images (frames) constituting the video included in the shooting information. The shooting point of each image is represented by two-dimensional coordinate values ​​on the blueprint. The processor 11 also identifies the shooting direction of each image using Visual SLAM technology. The shooting direction of each image is represented, for example, by a three-dimensional polar coordinate vector with north as the reference direction.

[0056] The image capturing device 30 captures video at a predetermined frame rate, so the number of images that make up the video is specified in frame cycles. Storing image information including all of the images that make up the captured video and meta information for each image requires a huge storage capacity in the memory 12. Therefore, the processor 11 determines, using a predetermined method, multiple shooting locations for the target to be displayed on the blueprint.

[0057] For example, the processor 11 may use Visual SLAM technology to extract a preset keyframe image from multiple images constituting a video and determine the shooting location corresponding to the image. However, the method for determining the shooting location is not limited to this. For example, the processor 11 may divide a blueprint of a predetermined space in which the video was shot into multiple regions and determine multiple shooting locations such that one shooting location is included in each of a predetermined number of regions. Alternatively, the processor 11 may extract images from the multiple images constituting the video at time intervals (e.g., 1 second, 5 seconds, 10 seconds, 1 minute, etc.) that are sufficiently longer than the frame rate, and determine the shooting locations corresponding to the images.

[0058] The processor 11 extracts images taken at each of the determined shooting locations (hereinafter, "taken images") from the video included in the shooting information. The processor 11 acquires meta information from the shooting information and adds to the meta information each shooting location and the shooting direction of the image taken at each shooting location. The processor 11 creates image information that associates the meta information with the extracted images taken at each shooting location, and stores the created image information in the image information storage unit 122.

[0059] The video information storage unit 123 stores a video (hereinafter, the entire video) between the shooting times at the shooting start point and the shooting end point, which is included in the shooting information received by the communication unit 13. Specifically, the processor 11 stores the entire video included in the shooting information received by the communication unit 13 in the video information storage unit 123.

[0060] Furthermore, the video information storage unit 123 stores partial videos in association with each shooting location of the object to be displayed on the blueprint. The partial videos are videos that include multiple images taken at each shooting location and its surroundings. Specifically, the processor 11 extracts, from the entire video, videos (first partial videos) between the shooting times at each of two shooting locations, one before and one after each shooting location, as partial videos, and stores the partial videos in the video information storage unit 123 in association with each shooting location.

[0061] The instruction receiving unit 111 acquires an instruction by the user input via the information terminal 20. In particular, when the communication unit 13 receives an instruction by the user from the information terminal 20, the instruction receiving unit 111 detects the input of the instruction. The instruction includes, for example, an instruction to select a blueprint, an instruction to specify a shooting time, an instruction to select a shooting location icon, an instruction to display a selected image, an instruction to display a surrounding video, an instruction to annotate an object, and the like.

[0062] The instruction to select a blueprint is an instruction to select a blueprint corresponding to the design ID input by the user as the blueprint to be processed. The instruction to specify a photography time is an instruction to specify the photography time input by the user as the photography time to be processed. A photography location icon is an icon superimposed on the position of a photography location included in a blueprint. The instruction to select a photography location icon is an instruction to select a photography location icon selected by the user from multiple photography location icons superimposed on the blueprint as the photography location icon to be processed.

[0063] The instruction to display a selected image is an instruction to display an image taken at the shooting location indicated by the shooting location icon selected as the processing target. The instruction to display a surrounding video is an instruction to display video taken between the shooting times at two shooting locations before and after the shooting location indicated by the shooting location icon selected as the processing target. The instruction to annotate an object is an instruction to add an annotation to an object included in a captured image specified by the user.

[0064] When the instruction receiving unit 111 detects input of an instruction to select a blueprint, the display control unit 112 reads out the blueprint selected by the selection instruction from the blueprint information storage unit 121. The display control unit 112 displays the read out blueprint on the display 23 of the information terminal 20.

[0065] When the instruction receiving unit 111 detects input of an instruction to specify a shooting time, the display control unit 112 displays, superimposed on the blueprint, multiple shooting location icons indicating a series of multiple shooting locations, including the shooting location where the photo was taken at the shooting time specified by the instruction.

[0066] 2 is a diagram showing an example of a design drawing 200 displayed on the display 23 of the information terminal 20. FIG. 2 shows an example in which a series of N photography location icons 210 are superimposed on the design drawing 200. In this example, the photography location icons 210 are composed of white circular images. Photography location icon 210(1) indicates the first (first) photography location of the series of N photography locations. Photography location icon 210(N) indicates the last (Nth) photography location of the series of N photography locations.

[0067] Here, it is assumed that the user selects a photography location icon 210 indicating one photography location from the multiple photography location icons 210 superimposed and displayed on the blueprint 200, and the instruction receiving unit 111 detects input of an instruction to select the photography location icon 210. Hereinafter, the photography location icon 210 selected by the user will be referred to as a first photography location icon 210(x). The photography location indicated by the first photography location icon 210(x) will be referred to as the first photography location x.

[0068] FIG. 3 is a diagram showing an example of a screen displayed when the first image capture location icon 210(x) is selected. In this case, the display control unit 112 changes the display mode of the first image capture location icon 210(x). FIG. 3 shows an example in which the color of the first image capture location icon 210(x) is changed from white to black. Note that the method of changing the display mode of the first image capture location icon 210(x) is not limited to this. For example, the shape of the first image capture location icon 210(x) may be changed, or the fill pattern may be changed to hatching, dots, or the like.

[0069] Similarly, the display control unit 112 causes the display mode of the photography location icon 210(x−1) indicating the photography location x−1 that is one location before the first photography location x indicated by the first photography location icon 210(x) to differ from that of the other photography location icons 210. The display control unit 112 causes the display mode of the photography location icon 210(x+1) indicating the photography location x+1 that is one location after the first photography location x to differ from that of the other photography location icons 210.

[0070] Hereinafter, the photography location icon 210(x-1) indicating the photography location x-1 that is one location before the first photography location x will be referred to as the second photography location icon 210(x-1), and the photography location icon 210(x+1) indicating the photography location x+1 that is one location after the first photography location x will be referred to as the second photography location icon 210(x+1). Furthermore, the photography location x-1 indicated by the second photography location icon 210(x-1) will be referred to as the second photography location x-1, and the photography location x+1 indicated by the second photography location icon 210(x+1) will be referred to as the second photography location x+1.

[0071] 3 shows an example in which the fill pattern of the second image capturing location icon 210(x-1) and the second image capturing location icon 210(x+1) has been changed to dots. Note that the method for changing the display mode of the second image capturing location icon 210(x-1) and the second image capturing location icon 210(x+1) is not limited to this. For example, the shape of the second image capturing location icon 210(x-1) and the second image capturing location icon 210(x+1) may be changed, or the fill color may be changed.

[0072] Then, the display control unit 112 displays the image display screen 300 adjacent to the design drawing 200. However, without being limited to this, the display control unit 112 may display the image display screen 300 at a position on the display 23 of the information terminal 20 that is spaced apart from the design drawing 200. Alternatively, the display control unit 112 may display the image display screen 300 so as to be superimposed on a part of the design drawing 200.

[0073] The display control unit 112 acquires the image captured at the first imaging point x from the image information storage unit 122 and displays a thumbnail image 310 that is a reduced version of the image and a video display button 320 on the image display screen 300 .

[0074] When the user clicks on the thumbnail image 310, the instruction receiving unit 111 detects the input of an instruction to display the selected image. In this case, the display control unit 112 obtains the image captured at the first image capturing point x from the image information storage unit 122 and displays the image on the display 23 of the information terminal 20.

[0075] When the user clicks the video display button 320, the instruction receiving unit 111 detects input of an instruction to display a surrounding video. In this case, the display control unit 112 generates a first video linked to the first shooting point x and displays the first video on the display 23 of the information terminal 20. The first video linked to the first shooting point x refers to a video included in the overall video as a video representing the situation around the first shooting point x.

[0076] Specifically, when the display control unit 112 detects an input instruction to display a peripheral video, it acquires a partial video associated with the first image capturing location x from the video information storage unit 123. The partial video is a video captured between the times of capture at two second image capturing locations x-1 and x+1, which are one location before and one location after the first image capturing location x. As a result, the display control unit 112 generates the acquired partial video as a first video linked to the first image capturing location x. The display control unit 112 displays the first video on the display 23 of the information terminal 20.

[0077] 4 is a diagram showing an example of how the first video is displayed. Specifically, the display control unit 112 displays a display screen 400 on the display 23 of the information terminal 20. The display screen 400 includes a display field 401, a play button 410, a pause button 420, a stop button 430, a playback time adjustment unit 440, and a volume adjustment unit 450.

[0078] The display control unit 112 displays the first video in the display field 401. The play button 410, pause button 420, stop button 430, and volume control unit 450 have the same configuration as the play button, pause button, stop button, and volume control unit provided in a general video playback player, and therefore detailed description thereof will be omitted.

[0079] The playback time adjustment unit 440 includes a progress bar 441 for displaying and changing the playback position of the first video, and a display field 442. The progress bar 441 has a configuration similar to that of a progress bar (seek bar) provided in a general video playback player, and therefore a detailed description thereof will be omitted. The display control unit 112 superimposes time information on the progress bar 441, associating the shooting time of the currently displayed image included in the first video, the shooting time at the first shooting location x, and the shooting times at each of the two second shooting locations x-1 and x+1.

[0080] Specifically, the display control unit 112 superimposes an icon 443 (x-1) on the left end of a progress bar 441 indicating the start position of playback of the first video, and displays the icon 443 (x-1) in the same display format as the second shooting location icon 210 (x-1) displayed on the blueprint 200 as the shooting time at the second shooting location x-1.

[0081] The display control unit 112 superimposes an icon 443 (x+1) on the right end of the progress bar 441 indicating the end position of playback of the first video, and displays the icon 443 (x+1) in the same display format as the second shooting location icon 210 (x+1) displayed on the blueprint 200 as the shooting time at the second shooting location x+1.

[0082] The display control unit 112 superimposes an icon 443(x) in the same display format as the first photographing location icon 210(x) displayed on the blueprint 200, indicating the photographing time at the first photographing location x, at the playback position of the image photographed at the first photographing location x on the progress bar 441. The playback position of the image photographed at the first photographing location x on the progress bar 441 is a position that is spaced from the left end of the progress bar 441 to the right by the product of the ratio of the elapsed time from the photographing time at the second photographing location x-1 to the photographing time at the first photographing location x to the total playback time of the first video and the length of the progress bar 441 in the longitudinal direction. The total playback time of the first video is the elapsed time from the photographing time at the second photographing location x-1 to the photographing time at the second photographing location x+1.

[0083] The display control unit 112 displays an icon 444 indicating the shooting time of the currently displayed image included in the first video at the current playback position of the first video on the progress bar 441 .

[0084] The display control unit 112 displays the current playback time of the first video (in the example of Figure 4, "00:02") and the total playback time of the first video (in the example of Figure 4, "00:06") in the display field 442.

[0085] Therefore, by looking at the playback time adjustment unit 440, the user can intuitively grasp the relative positional relationship between the shooting location of the currently displayed image included in the first video and each of the first shooting location x and the two second shooting locations x-1 and x+1.

[0086] In addition, instead of the display control unit 112 superimposing the icon 443(x-1), the icon 443(x), and the icon 443(x+1) on the progress bar 441, the display control unit 112 may display the shooting time at the second shooting location x-1, the shooting time at the first shooting location x, and the shooting time at the second shooting location x+1 at the bottom or top of the progress bar 441, etc.

[0087] (Processing Performed During Photographing Operation) Next, a description will be given of the processing performed by the information processing system 1 during one photographing operation by the photographer. Fig. 5 is a flowchart showing a first example of the processing performed by the information processing system 1 during photographing operation.

[0088] In step S1, the operation unit of the communication device 40 accepts an operation to input a shooting start point (hereinafter, starting point). For example, the photographer inputs an operation to specify the starting point on a blueprint screen displayed on the display of the communication device 40. The photographer may specify the starting point by tapping on the blueprint screen, or may specify the coordinate value of the starting point.

[0089] Next, the photographer performs a shooting operation in a predetermined space at the site by pointing the shooting direction of the camera device 30 toward north and inputting a shooting start command to the camera device 30. As a result, in step S2, the camera device 30 acquires an entire video, which is a video between the shooting times at each of the shooting start point and end point of the shooting operation.

[0090] Next, in step S3, the operation unit of the communication device 40 accepts an operation to input a shooting end point (hereinafter, end point). For example, the photographer inputs an operation to specify the end point on the blueprint screen displayed on the display of the communication device 40. The photographer may specify the end point by tapping on the blueprint screen, or may specify the shooting end point by inputting the coordinate values ​​of the end point.

[0091] Next, in step S4 , the communication device 40 acquires, from the image capturing device 30 , image capturing information including the entire video captured by the image capturing device 30 , and uploads (transmits) the acquired image capturing information to the server 10 .

[0092] Next, in step S5, the display control unit 112 calculates the shooting points of each of the multiple images that make up the full video included in the shooting information acquired in step S4 using a self-location estimation process. As the self-location estimation process, VSLAM (Visual Simultaneous Localization and Mapping) can be adopted.

[0093] Next, in step S6, the display control unit 112 determines, using a predetermined method, multiple shooting locations of the object to be displayed on the blueprint. For example, the display control unit 112 uses Visual SLAM technology to extract images of preset key frames from multiple images that make up the entire video, and determines the shooting locations corresponding to those images.

[0094] Next, in step S7, the display control unit 112 creates a list of the photographing times at each of the photographing locations determined in step S6.

[0095] Next, in step S8, the display control unit 112 sets each of the multiple photography locations determined in step S6 as a processing target and starts processing for the photography locations to be processed. Here, the explanation is given assuming that there are N (N is an integer of 2 or more) photography locations. The display control unit 112 sets the photography locations to be processed in order of the photography time. j is an index that specifies the photography location to be processed, and takes on a value of j=1 to N.

[0096] Next, in step S9, the display control unit 112 generates a partial video corresponding to the shooting location j, associates the partial video with the shooting location j, and stores the partial video in the video information storage unit 123. Specifically, the display control unit 112 performs the following process by referring to the list of shooting times created in step S7.

[0097] The display control unit 112 extracts, as a partial video corresponding to the shooting location j, videos between the shooting times at the shooting location j-1, which is one location before the shooting location j, and the shooting location j+1, which is one location after the shooting location j, from the entire video included in the shooting information acquired in step S4. The display control unit 112 stores the extracted partial video in the video information storage unit 123 in association with the shooting location j.

[0098] Next, in step S10, the display control unit 112 determines whether or not processing has been completed for all of the photography locations j determined in step S6. If processing has not been completed for all of the photography locations j, the process returns to step S8, and the next photography location j is set as the photography location j to be processed. On the other hand, if processing has been completed for all of the photography locations j, that is, if processing has been completed for photography location j (= N), the process ends.

[0099] (Peripheral Video Display Process) Next, a process for displaying a peripheral video of a shooting location indicated by a shooting location icon selected by a user (hereinafter, peripheral video display process) will be described. Fig. 6 is a flowchart showing an example of the peripheral video display process. This process is triggered by the instruction receiving unit 111 detecting input of an instruction to select a blueprint.

[0100] In step S21, the display control unit 112 displays the design drawing 200 selected by the selection instruction received by the instruction receiving unit 111 on the display 23 of the information terminal 20.

[0101] Next, in step S22, the instruction accepting unit 111 detects whether an instruction to specify a shooting time has been input. While the instruction accepting unit 111 does not detect the input of an instruction to specify a shooting time (NO in step S22), the process of step S22 is repeated. If the instruction accepting unit 111 detects the input of an instruction to specify a shooting time (YES in step S22), the process proceeds to step S23.

[0102] In step S23, the display control unit 112 displays, superimposed on the design drawing 200 displayed in step S21, multiple shooting location icons 210 indicating a series of multiple shooting locations, including the shooting location where the shooting was performed at the shooting time specified by the specified instruction detected by the instruction receiving unit 111.

[0103] Next, in step S24, the instruction receiving unit 111 detects whether or not an instruction to select the photographing location icon 210 has been input. While the instruction receiving unit 111 does not detect the input of an instruction to select the photographing location icon 210 (NO in step S24), the processing of step S24 is repeated. If the instruction receiving unit 111 detects the input of an instruction to select the photographing location icon 210 (YES in step S24), the processing proceeds to step S25.

[0104] Next, in step S25 , the display control unit 112 displays an image display screen 300 including a thumbnail image 310 and a video display button 320 on the display 23 of the information terminal 20 .

[0105] At this time, the display control unit 112 further changes the display mode of the first photographing location icon 210(x), the second photographing location icon 210(x-1) indicating the second photographing location x-1 that is one location before the first photographing location x indicated by the first photographing location icon 210(x), and the second photographing location icon 210(x+1) indicating the second photographing location x+1 that is one location after the first photographing location x to a display mode different from that of the other photographing location icons 210.

[0106] Next, in step S26, the instruction receiving unit 111 detects whether an instruction to display a peripheral video has been input. If the instruction receiving unit 111 does not detect an input of an instruction to display a peripheral video (NO in step S26), the process proceeds to step S27. If the instruction receiving unit 111 detects an input of an instruction to display a peripheral video (YES in step S26), the process proceeds to step S28.

[0107] In addition, in step S25, the process of changing the display mode of the first shooting location icon 210(x) and the two second shooting location icons 210(x-1) and 210(x+1) may be performed when the instruction receiving unit 111 detects input of an instruction to display the surrounding video (YES in step S26).

[0108] In step S27, the instruction receiving unit 111 detects whether or not an instruction to select the photography location icon 210 has been input. If the instruction receiving unit 111 does not detect the input of an instruction to select the photography location icon 210 (NO in step S27), the processing returns to step S26. If the instruction receiving unit 111 detects the input of an instruction to select the photography location icon 210 (YES in step S27), the processing returns to step S25.

[0109] In step S28, the display control unit 112 generates a first video linked to the first image capturing location x indicated by the first image capturing location icon 210(x) indicated by the selection instruction detected in step S24 or step S27.

[0110] Next, in step S29, display control unit 112 displays display screen 400 on display 23 of information terminal 20, and displays the first moving image generated in step S28 in display field 401 of display screen 400. After step S29, the process returns to step S24.

[0111] In this manner, in the above embodiment, the first video, which is a video between the shooting times at each of the two second shooting locations x-1 and x+1, which are located one location before and one location after the first shooting location x selected by the user, is displayed on the display 23 of the information terminal 20. This allows the user to properly grasp the situation from the selected first shooting location to the two second shooting locations before and after it.

[0112] Furthermore, when the selection of the first photographing location icon 210(x) is detected, the display mode of the two second photographing location icons 210(x-1) and 210(x+1) displayed on the bird's-eye view becomes different from the display mode of the other photographing location icons. Therefore, by looking at the design plan 200, the user can easily understand which two photographing locations have videos between the photographing times that are displayed on the display screen 400.

[0113] The present disclosure can employ the following modifications.

[0114] (1) In the above embodiment, an example was described in which the display control unit 112 extracts, in step S9 (Figure 5), videos between the shooting times at the shooting location j-1, which is one location before the shooting location j, and the shooting location j+1, which is one location after the shooting location j, from the entire video as a partial video corresponding to the shooting location j, and stores the videos in the video information storage unit 123.

[0115] However, instead of this, the display control unit 112 may extract from the entire video as a partial video a video (second partial video) from a first average time, which is the average of the shooting times at each shooting location j−1 immediately before each shooting location j and each shooting location j, to a second average time, which is the average of the shooting times at each shooting location j and each shooting location j+1 immediately after each shooting location j. Then, the display control unit 112 may store the partial video in the video information storage unit 123 in association with each shooting location j.

[0116] In addition to this, in step S28 (FIG. 6), the display control unit 112 may acquire partial videos corresponding to not only the first image capturing location x but also two second image capturing locations x-1 and x+1, which are located one location before and one location after the first image capturing location x, from the video information storage unit 123. Then, the display control unit 112 may generate a first video linked to the first image capturing location x using the acquired three partial videos.

[0117] Specifically, from the partial video corresponding to the second imaging point x-1 (third imaging point) that is one point before the first imaging point x, the display control unit 112 extracts, as the first segmented video, a video from the time of shooting at the second imaging point x-1 onwards. Furthermore, from the partial video corresponding to the second imaging point x+1 (fourth imaging point) that is one point after the first imaging point x, the display control unit 112 extracts, as the second segmented video, a video from the time of shooting at the second imaging point x+1 onwards. The display control unit 112 then generates, as the first segmented video, a video that combines the first segmented video, the partial video corresponding to the first imaging point x, and the second segmented video.

[0118] In this case, the video from the first average time to the second average time is stored in the video information storage unit 123 as a partial video corresponding to each shooting point j necessary to compose the first video. This makes it possible to prevent a part or the whole of a partial video corresponding to each shooting point j from overlapping with a partial video corresponding to another shooting point. This makes it possible to prevent the storage capacity of the memory 12 from becoming too large.

[0119] (2) Unlike the above embodiment and the above modification (1), the display control unit 112 may extract partial videos corresponding to each shooting location from the entire video, as described below, and detect whether or not the partial videos contain captured images of a predetermined object. Then, when the display control unit 112 detects that the partial videos contain captured images of a predetermined object, the display control unit 112 may extend the partial videos corresponding to each shooting location extracted from the entire video so that the captured images of the predetermined object are displayed without interruption. This configuration can be realized, for example, as follows.

[0120] Fig. 7 is a flowchart showing a second example of the processing of the information processing system 1 performed during a shooting operation. Specifically, the processing shown in Fig. 5 is modified as shown in Fig. 7. More specifically, after step S7, the processing proceeds to step S11. In step S11, the display control unit 112 uses a trained model constructed by machine learning to detect the shooting time (hereinafter, detection time) of each of one or more captured images, including captured images of annotation objects included in the entire video, and creates a list of the one or more detected detection times.

[0121] In more detail, the memory 12 stores in advance a model that has been machine-learned to describe the relationship between a second video, which is a video shot in a space including a series of multiple shooting locations and includes one or more captured images of an annotation object present in the space, and each of the captured images of the one or more captured images of the annotation object included in the second video. The annotation object is an object included in a captured image stored in the image information storage unit 122 and designated by the user as a target to be annotated, as indicated by an annotation instruction for the object acquired by the instruction receiving unit 111.

[0122] The display control unit 112 acquires (detects) one or more shooting times obtained by inputting the entire video included in the shooting information acquired in step S4 into the model stored in the memory 12 as detection times for each of the one or more captured images including the captured image of the annotation object. The display control unit 112 creates a list of detection times for each of the one or more captured images.

[0123] Thereafter, in step S9a, which is a modification of step S9, the display control unit 112 generates partial videos corresponding to each shooting point j, taking into account the list of detection times created in step S11, and stores the partial videos in association with each shooting point j in the video information storage unit 123. Specifically, similar to step S9, the display control unit 112 refers to the list of shooting times created in step S7 (hereinafter, first list) and also refers to the list of detection times created in step S11 (hereinafter, second list), and performs the following processing.

[0124] 8 is an explanatory diagram of a modified example of a method for creating partial videos corresponding to each of the seven shooting locations. In this example, when the first list includes shooting times t1 to t7 at each of the seven shooting locations, the display control unit 112 creates partial videos corresponding to each of the first to sixth shooting locations by taking into account a plurality of consecutive detection times from time ta to time tb, which are included in the second list, and a plurality of consecutive detection times from time tc to time td, which are included in the second list, and which are included in the second list, and which are included in the second list, and taking into account a plurality of consecutive detection times from time tc to time td, ...

[0125] In step S9a, the display control unit 112 extracts from the overall video a video (third partial video) between the shooting times at each shooting location j and the shooting location j+1 one location after each shooting location j, as a video (hereinafter, partial candidate video) that is a candidate for the partial video corresponding to each shooting location j.

[0126] 8 , the display control unit 112 extracts from the entire video a video of period a, from shooting time t1 at the first shooting location to shooting time t2 at the second shooting location, as a partial candidate video corresponding to the first shooting location. Similarly, the display control unit 112 extracts from the entire video video of periods b, c, d, e, and f as partial candidate videos corresponding to the second, third, fourth, fifth, and sixth shooting locations, respectively.

[0127] Next, if one or more detection times included in the second list are included within the period of the part candidate video corresponding to each shooting point j, the display control unit 112 detects that the part candidate video includes a captured image of the annotation object. In this case, the display control unit 112 extracts from the entire video, as object videos corresponding to each shooting point j, videos that are made up only of images that include captured images of the annotation object and that have a period that overlaps with the part candidate video.

[0128] In the example of Figure 8, the period a of the part candidate video corresponding to the first shooting location includes multiple consecutive detection times from time ta to time tb included in the second list, which are spaced at a predetermined frame interval. Therefore, the display control unit 112 detects that the part candidate video includes captured images of the annotation object. In this case, a video consisting only of images including captured images of the annotation object captured at multiple consecutive detection times from time ta to time tb, which are spaced at a predetermined frame interval, overlaps with the part candidate video corresponding to the first shooting location in the period from time ta to time tb. Therefore, the display control unit 112 extracts the video between time ta and time tb from the overall video as an object video corresponding to the first shooting location.

[0129] The display control unit 112 detects that the portion candidate video corresponding to the second shooting location includes captured images of the annotation object because the period b of the portion candidate video corresponding to the second shooting location includes some of the multiple detection times that are consecutive at a predetermined frame cycle from time tc to time td and are included in the second list. In this case, a video consisting only of images that include captured images of the annotation object taken at multiple detection times that are consecutive at a predetermined frame cycle from time tc to time td overlaps with the portion candidate video corresponding to the second shooting location in the period from time tc to time t3. Therefore, the display control unit 112 extracts the video between time tc and time td from the overall video as the object video corresponding to the second shooting location.

[0130] The display control unit 112 detects that the portion candidate video corresponding to the third shooting location includes captured images of the annotation object because the period c of the portion candidate video corresponding to the third shooting location includes some of the multiple consecutive detection times from time tc to time td included in the second list. In this case, a video consisting only of images including captured images of the annotation object captured at multiple consecutive detection times from time tc to time td overlaps with the portion candidate video corresponding to the third shooting location in the period from time t3 to time t4. Therefore, the display control unit 112 extracts the video between time tc and time td from the overall video as the object video corresponding to the third shooting location.

[0131] The display control unit 112 detects that the portion candidate video corresponding to the fourth shooting location includes captured images of the annotation object because the period d of the portion candidate video corresponding to the fourth shooting location includes some of the multiple detection times that are consecutive at a predetermined frame cycle from time tc to time td and are included in the second list. In this case, a video consisting only of images that include captured images of the annotation object taken at multiple detection times that are consecutive at a predetermined frame cycle from time tc to time td overlaps with the portion candidate video corresponding to the fourth shooting location in the period from time t4 to time td. Therefore, the display control unit 112 extracts the video between time tc and time td from the overall video as the object video corresponding to the fourth shooting location.

[0132] The display control unit 112 detects that the partial candidate video corresponding to the fifth and sixth shooting locations does not include a detection time included in the second list within periods e and f, and therefore detects that the partial candidate video does not include a captured image of the annotation object.

[0133] Then, when the display control unit 112 detects that a captured image of the annotation object is included in the partial candidate video corresponding to each shooting point j, the display control unit 112 creates a video (fourth partial video) by combining the partial candidate video and the object video so that their periods do not overlap, as a partial video corresponding to each shooting point j. The display control unit 112 stores the partial video in the video information storage unit 123 in association with each shooting point j. When the display control unit 112 detects that a captured image of the annotation object is not included in the partial candidate video corresponding to each shooting point j, the display control unit 112 stores the partial candidate video corresponding to each shooting point j in association with each shooting point j.

[0134] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the first shooting location includes a captured image of an annotation object. Therefore, the display control unit 112 combines the video of the period a from shooting time t1 to shooting time t2, which is the partial candidate video corresponding to the first shooting location, with the video of the period from shooting time ta to shooting time tb, which is the object video corresponding to the first shooting location, so that the periods do not overlap. The display control unit 112 creates the video of the period a from shooting time t1 to shooting time t2 obtained by this combination as a partial video corresponding to the first shooting location, and stores the partial video in the video information storage unit 123 in association with the first shooting location.

[0135] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the second shooting location includes a captured image of the annotation object. Therefore, the display control unit 112 combines the video of the period b from shooting time t2 to shooting time t3, which is the partial candidate video corresponding to the second shooting location, with the video of the period from shooting time tc to shooting time td, which is the object video corresponding to the second shooting location, so that the periods do not overlap. The display control unit 112 creates the video of the period b1 from shooting time t2 to shooting time td obtained by this combination as a partial video corresponding to the second shooting location, and stores the partial video in the video information storage unit 123 in association with the second shooting location.

[0136] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the third shooting location includes a captured image of the annotation object. Therefore, the display control unit 112 combines the video of the period c from shooting time t3 to shooting time t4, which is the partial candidate video corresponding to the third shooting location, with the video of the period from shooting time tc to shooting time td, which is the object video corresponding to the third shooting location, so that the periods do not overlap. The display control unit 112 creates the video of the period c1 from shooting time tc to shooting time td obtained by this combination as a partial video corresponding to the third shooting location, and stores the partial video in the video information storage unit 123 in association with the third shooting location.

[0137] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the fourth shooting location includes a captured image of the annotation object. Therefore, the display control unit 112 combines the video of the period d from shooting time t4 to shooting time t5, which is the partial candidate video corresponding to the fourth shooting location, with the video of the period from shooting time tc to shooting time td, which is the object video corresponding to the fourth shooting location, so that the periods do not overlap. The display control unit 112 creates the video of the period d1 from shooting time tc to shooting time t5 obtained by this combination as a partial video corresponding to the fourth shooting location, and stores the partial video in the video information storage unit 123 in association with the fourth shooting location.

[0138] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the fifth shooting location does not include a captured image of the annotation object. Therefore, the display control unit 112 creates a video of the period e from shooting time t5 to shooting time t6, which is the partial candidate video corresponding to the fifth shooting location, as a partial video corresponding to the fifth shooting location, and stores the partial video in the video information storage unit 123 in association with the fifth shooting location.

[0139] 8 , as described above, the display control unit 112 detects that the partial candidate video corresponding to the sixth shooting location does not include a captured image of the annotation object. Therefore, the display control unit 112 creates a video of the period f from shooting time t6 to shooting time t7, which is the partial candidate video corresponding to the sixth shooting location, as a partial video corresponding to the sixth shooting location, and stores the partial video in the video information storage unit 123 in association with the sixth shooting location.

[0140] In this modified example, in step S28 shown in FIG. 6, the display control unit 112 generates a first video linked to the first shooting point x in the following manner.

[0141] Specifically, the display control unit 112 generates a first video by combining a partial video corresponding to the second shooting location x-1, which is one location before the first shooting location x, and a partial video corresponding to the first shooting location x, so that the periods do not overlap.

[0142] In the example of Figure 8, when the first shooting location x is the first shooting location, the display control unit 112 generates the first video, which is a partial video corresponding to the first shooting location x, that is, a video of the period a from shooting time t1 to shooting time t2.

[0143] 8 , when the first photographing location x is the second photographing location, the display control unit 112 combines a partial video corresponding to the first photographing location immediately before the first photographing location x, that is, a video of a period a from photographing time t1 to photographing time t2, with a partial video corresponding to the first photographing location x, that is, a video of a period b1 from photographing time t2 to photographing time td, so that the periods do not overlap. The display control unit 112 generates the video of the period a+b1 from photographing time t1 to photographing time td obtained by this combination as the first video linked to the second photographing location.

[0144] 8 , when the first photographing location x is the third photographing location, the display control unit 112 combines a partial video corresponding to the second photographing location immediately before the first photographing location x, which is a video of a period b1 from photographing time t2 to photographing time td, with a partial video corresponding to the first photographing location x, which is a video of a period c1 from photographing time tc to photographing time td, so that the periods do not overlap. The display control unit 112 generates the video of the period b1 from photographing time t2 to photographing time td obtained by this combination as a first video linked to the third photographing location.

[0145] 8 , when the first image capturing location x is the fourth image capturing location, the display control unit 112 combines a partial video corresponding to the third image capturing location immediately before the first image capturing location x, which is a video of a period c1 from image capturing time tc to image capturing time td, with a partial video corresponding to the first image capturing location x, which is a video of a period d1 from image capturing time tc to image capturing time t5, so that the periods do not overlap. The display control unit 112 generates the video of the period d1 from image capturing time tc to image capturing time t5 obtained by this combination as a first video linked to the fourth image capturing location.

[0146] 8 , when the first image capturing location x is the fifth image capturing location, the display control unit 112 combines a partial video corresponding to the fourth image capturing location immediately before the first image capturing location x, which is a video of a period d1 from image capturing time tc to image capturing time t5, with a partial video corresponding to the first image capturing location x, which is a video of a period e from image capturing time t5 to image capturing time t6, so that the periods do not overlap. The display control unit 112 generates the video of the period d1+e from image capturing time tc to image capturing time t6 obtained by this combination as the first image capturing location associated with the fifth image capturing location.

[0147] 8 , when the first image capturing location x is the sixth image capturing location, the display control unit 112 combines a partial video corresponding to the fifth image capturing location immediately before the first image capturing location x, that is, a video of a period e from image capturing time t5 to image capturing time t6, with a partial video corresponding to the first image capturing location x, that is, a video of a period f from image capturing time t6 to image capturing time t7, so that the periods do not overlap. The display control unit 112 generates the video of the period e+f from image capturing time t5 to image capturing time t7 obtained by this combination as the first video associated with the sixth image capturing location.

[0148] According to this modification, not only the images captured at the second image capturing locations x-1 and x+1, which are located immediately before and after the first image capturing location x, but also the images captured immediately before and / or after the image capturing locations are displayed on the display screen 400. This allows the user to focus on the objects that can be designated as the objects to be annotated.

[0149] (3) In the processing performed in the above embodiment and the above modification (1), among a series of multiple shooting locations, the shooting location one location before a certain shooting location may be replaced with the shooting location n locations before that shooting location, where n is a natural number greater than or equal to 2. Similarly, the shooting location one location after that shooting location may be replaced with the shooting location m locations after that shooting location, where m is a natural number greater than or equal to 2.

[0150] Specifically, in the above embodiment, in step S9 (Figure 5), the display control unit 112 may extract from the entire video the video between the shooting times at shooting point j-n, which is n points before shooting point j, and shooting point j+m, which is m points after shooting point j, as a partial video corresponding to shooting point j.

[0151] In the above-described modified example (1), in step S9 ( FIG. 5 ), the display control unit 112 may calculate, as a first average time, the average of the shooting times at each shooting location j and the shooting location j-n that is n locations before each shooting location j. Alternatively, the display control unit 112 may calculate, as a second average time, the average of the shooting times at each shooting location j and the shooting location j+m that is m locations after each shooting location j. Then, the display control unit 112 may extract, from the entire video, a video from the first average time to the second average time as a partial video.

[0152] Accordingly, in step S28 (FIG. 6), the display control unit 112 may acquire partial videos corresponding not only to the first photographing location x but also to two second photographing locations x-n and x+m that are n locations before and m locations after the first photographing location x from the video information storage unit 123. Then, the display control unit 112 may use these partial videos to generate a first video linked to the first photographing location x.

[0153] More specifically, the display control unit 112 may extract, from a partial video corresponding to a second imaging point x-n that is n images before the first imaging point x, a video from the time after the imaging time at the second imaging point x-n as the first segmented video. Furthermore, the display control unit 112 may extract, from a partial video corresponding to a second imaging point x+m that is m images after the first imaging point x, a video from the time before the imaging time at the second imaging point x+m as the second segmented video. The display control unit 112 may then generate, as the first segmented video, a video that combines the first segmented video, the partial video corresponding to the first imaging point x, and the second segmented video.

[0154] (4) In step S11 (FIG. 7) of the above variant (2), the display control unit 112 may use a known image recognition process to detect the detection time, which is the shooting time of each of one or more captured images, including a captured image of a specified object included in the overall video.

[0155] The present disclosure is useful in the technical field of remotely managing the progress of construction work at a construction site.

Claims

1. An information processing method in a computer, comprising: displaying a bird's-eye view on a display of a terminal device; displaying a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view; and, when selection of a first shooting location icon from among the plurality of shooting location icons is detected, displaying a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

2. The information processing method of claim 1, further comprising: when detecting selection of the first shooting location icon, making the display mode of the two second shooting location icons displayed on the bird's-eye view different from the display mode of other shooting location icons.

3. An information processing method as described in claim 1 or 2, wherein displaying the first video includes displaying time information that associates the shooting time of the currently displayed image included in the first video with the shooting time at the first shooting location and the shooting time at each of the two second shooting locations.

4. The information processing method of claim 3, wherein displaying the time information includes: displaying the shooting time at the first shooting location with an icon in the same display format as the first shooting location icon displayed on the bird's-eye view; and displaying the shooting time at each of the two second shooting locations with two icons in the same display format as the two second shooting location icons displayed on the bird's-eye view.

5. The information processing method of claim 1, further comprising: extracting a first partial video, which is a video between the shooting times at each of the shooting locations before and after each shooting location, from an overall video, which is a video between the shooting times at each of the start and end points of shooting at the multiple shooting locations; and storing the first partial video in correspondence with each shooting location; and displaying the first video includes displaying the first partial video corresponding to the first shooting location as the first video.

6. The information processing method of claim 1, further comprising: extracting a second partial video, which is a video from a first average time, which is the average of the shooting times at each shooting location and each shooting location before each shooting location, to a second average time, which is the average of the shooting times at each shooting location and each shooting location after each shooting location, from an entire video, which is a video between the shooting times at each of the start and end points of shooting at the multiple shooting locations; and storing the second partial video in correspondence with each shooting location; and displaying the first video includes displaying as the first video a video that is a combination of a video of the second partial video corresponding to a third shooting location, which is the second shooting location before the first shooting location, a video of the second partial video after the shooting time at the third shooting location, the second partial video corresponding to the first shooting location, and a video of the second partial video corresponding to a fourth shooting location, which is the second shooting location after the first shooting location, before the shooting time at the fourth shooting location.

7. Extracting a third partial video, which is a video between the shooting times at each shooting location and the shooting location one location after each shooting location, from an overall video, which is a video between the shooting times at each of the start and end points of shooting at the multiple shooting locations; when it is detected that the third partial video contains a shooting image of a predetermined object, extracting an object video from the overall video, which is a video consisting only of images including the shooting image of the predetermined object and which has a period that overlaps with the third partial video, and storing a fourth partial video, which is a video obtained by combining the third partial video and the object video so that their periods do not overlap, in association with each shooting location; and when it is detected that the third partial video does not contain a shooting image of the predetermined object, storing the third partial video in association with each shooting location; and displaying the first video includes displaying as the first video a video obtained by combining the third partial video or the fourth partial video corresponding to the second shooting location one location before the first shooting location and the third partial video or the fourth partial video corresponding to the first shooting location so that their periods do not overlap. The information processing method according to claim 1 .

8. The information processing method of claim 7, wherein detecting whether or not the third partial video contains a photographed image of the specified object includes detecting that the third partial video contains a photographed image of the specified object when the one or more photographing times obtained by inputting the entire video into a model that has machine-learned the relationship between a second video, which is a video shot in a space including the multiple shooting locations and includes one or more photographed images of an annotation object present in the space, and each of the photographing times of the one or more photographed images of the annotation object included in the second video, are included within the period of the third partial video, wherein the annotation object is an object designated by a user as a target to be annotated.

9. An information processing method according to any one of claims 1, 5 and 6, wherein, among the plurality of shooting locations, a shooting location preceding one of the shooting locations is a shooting location n locations preceding the one of the shooting locations, and a shooting location following the one of the shooting locations is a shooting location m locations following the one of the shooting locations, and the n and the m are natural numbers greater than or equal to 1.

10. An information processing device including a processor, wherein the processor executes the following operations: displaying a bird's-eye view on a display of a terminal device; displaying a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view; and, when it detects selection of a first shooting location icon from among the plurality of shooting location icons, displaying a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.

11. An information processing program that causes a computer to perform the following steps: display a bird's-eye view on a display of a terminal device; display a plurality of shooting location icons indicating a series of a plurality of shooting locations on the bird's-eye view; and, when selection of a first shooting location icon from among the plurality of shooting location icons is detected, display a first video linked to the first shooting location indicated by the first shooting location icon, wherein the first video is a video between shooting times at each of two second shooting locations indicated by two second shooting icons displayed before and after the first shooting location icon.