Image Processing System
The image processing system addresses inaccuracies in selecting past scenes and heavy processing loads by using key frames and two-dimensional transformation to overlay past videos onto current scenes, enhancing VR/AR experiences.
Patent Information
- Application Number
- JP2022084849
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-12-22
- Estimated Expiration
- 2042-05-24
AI Technical Summary
Conventional VR/AR systems face challenges in accurately selecting appropriate past scenes for overlay due to location information inaccuracies and heavy data processing loads when using three-dimensional models for object tracking.
An image processing system that extracts key frames from past videos, associates them with location information, and uses two-dimensional affine transformation to overlay pre-processed past videos onto current scenes, reducing processing load and improving accuracy.
Effectively selects appropriate past scenes for overlay, reducing processing load and enhancing realism in VR/AR experiences.
Smart Images

Figure 0007789622000001 
Figure 0007789622000002 
Figure 0007789622000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to VR (Virtual Reality) / AR (Augmented Reality) technology, and in particular to technology that is effective when applied to an image processing system that realizes image / video-based VR / AR. [Background technology]
[0002] Systems are being investigated and developed that use VR / AR technology to allow users to have simulated experiences that make them feel as if they are in another space or time. For example, technologies are being investigated and developed that allow users to experience traveling to a distant place (another space) without leaving the comfort of their own home, or to experience what a scene they are currently seeing while traveling looked like in the past (another time) with a sense of presence.
[0003] As a technology related to the latter, for example, U.S. Patent No. 10,127,730 (Patent Document 1) describes a mechanism that uses VR / AR technology to display content simulating attractions that existed in the past at a user's current location based on the current location information of the user's terminal. This technology makes it possible to superimpose past footage of the location onto the scene currently being viewed.
[0004] As an example of a technology for displaying other content superimposed on the current scene, Patent Publication No. 6420605 (Patent Document 2) describes an image processing device that uses image recognition-based AR technology to track objects of any shape, enabling robust tracking with a small DB size and processing load, even when there are large changes in viewpoint such as shooting angle and distance. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] U.S. Patent No. 10,127,730 [Patent Document 2] Patent No. 6420605 Summary of the Invention [Problem to be solved by the invention]
[0006] By using the conventional technology described in Patent Document 1, for example, it is possible to realize a VR / AR system that displays a past scene at a location superimposed on an image of the current scene.
[0007] However, in the case of a system that acquires content related to past scenes at a location based solely on location information, such as the technology described in Patent Document 1, there is a problem in that, depending on the means for acquiring location information, it may be difficult to select or determine an appropriate past scene that corresponds to the current scene, for example, when accurate location information cannot be acquired indoors or underground, or when an error of several tens of centimeters to several meters occurs, or when the scene is completely different depending on the direction the user is facing even at the same location.
[0008] Furthermore, when a VR / AR system displays another image superimposed on a video of the current scene, it is common to treat the object of the image to be superimposed as a three-dimensional model in order to follow the user's movements and changes in the direction they are facing, as in the conventional technology described in Patent Document 2. However, when using a three-dimensional model, the object is deformed using CG, resulting in a less realistic scene and a heavy data processing load.
[0009] Therefore, the object of the present invention is to provide an image processing system that effectively and efficiently selects appropriate past scenes to overlay on images of current scenes using VR / AR, and reduces the load on the system when overlaying the past scenes.
[0010] The above and other objects and novel features of the present invention will become apparent from the description of this specification and the accompanying drawings. [Means for solving the problem]
[0011] Among the inventions disclosed in this application, the outline of representative inventions will be briefly explained as follows.
[0012] An image processing system that is a representative embodiment of the present invention is an image processing system that uses an image processing device to overlay a past video in which a corresponding past scene was captured on a current video related to a current scene captured by a shooting device and display it on a display device.The image processing device has a pre-processing unit that extracts one or more key frames from each of a plurality of past videos, associates them with the past videos from which they were extracted, and records them as pre-processed past videos, an image search unit that identifies a first key frame similar to the current video from the pre-processed past videos by image search, and identifies a first past video corresponding to the first key frame, and an image synthesis unit that converts the first past video on the current video so as to minimize the deviation of feature points in each video, and overlays the first past video on the display device, displaying it on the display device. [Effects of the Invention]
[0013] The effects obtained by the representative inventions disclosed in this application can be briefly explained as follows.
[0014] That is, according to the representative embodiment of the present invention, when a past scene is superimposed on a video of a current scene using VR / AR, it is possible to effectively and efficiently select an appropriate past scene for superimposition, and also to reduce the load of the superimposition. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating an overview of an example of the configuration of an image processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating an example of image processing according to an embodiment of the present invention. [Figure 3]FIG. 1 is a diagram showing an overview of an example of a data structure according to an embodiment of the present invention. [Figure 4] 10 is a flowchart outlining an example of a preprocessing flow according to an embodiment of the present invention. [Figure 5] 1 is a flowchart outlining an example of the flow of image search and synthesis processing in one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In all drawings used to explain the embodiments, the same parts are generally designated by the same reference numerals, and repeated explanations will be omitted. However, parts that have been designated and explained in one drawing may be referred to by the same reference numerals in the explanation of other drawings, although they will not be shown again.
[0017] <Summary> Systems that use VR / AR technology to enable a simulated travel experience from the comfort of your own home have been explored and developed in response to the self-restraint on going out due to the spread of COVID-19. However, from the perspective of supporting local communities in the with-COVID / post-COVID era and revitalizing local areas through digitalization, it is necessary to aim for VR / AR systems that focus on promoting travel to local areas.
[0018] That is, for example, as an experience before going on a trip, by allowing the user to view local scenery in real time from the comfort of their own home or participate in an event, it is possible to foster a desire to actually visit the local area next time. Also, as an experience before the trip or while the user is actually traveling, for example, by superimposing past local scenery (video) that cannot be viewed directly in real time at that time on video of the current local scenery to create a sense of realism, it is possible to foster a desire to return on another occasion.
[0019] The past scenes (videos) that are overlaid here could be, for example, scenes that no longer exist, such as abandoned railways or demolished buildings, events that were held at specific times in the past, such as legendary live events or famous sports matches, or commentary by local reporters, guides, or famous commentators, etc. They could also be scenes that exist today but in a different state than they are now, such as the best scenery for a particular season, weather, or time of day that is ideal for sightseeing, such as rows of cherry blossoms in full bloom, a clear day, or a unique view that can only be seen at dusk or during high or low tide.
[0020] The image processing system, which is one embodiment of the present invention described below, uses VR / AR to overlay (combine) past scenes from the same location onto a captured image of the current scene, allowing the user to have the experience described above.
[0021] <System configuration> 1 is a diagram showing an overview of an example of the configuration of an image processing system according to one embodiment of the present invention. The image processing system 1 includes various devices, such as an image processing device 10 that performs image processing related to VR / AR, a photographing device 20 that photographs a scene in response to a user's instruction and acquires a video image of the scene, and a display device 30 that has a function of displaying the video image photographed by the photographing device 20 and presenting it to the user. The photographing device 20 and the display device 30 are connected to the image processing device 10 by wired or wireless communication means or connection means (not shown).
[0022] The image capturing device 20 can be configured as, for example, a digital video camera, web camera, action camera, drive recorder, or other device with a video capturing function. If the display device 30 described below has a camera function, the image capturing device 20 may be configured as an integrated device to use this function. The display device 30 can be configured as, for example, a device with a video image display function, such as a monitor or a PC (Personal Computer) display. The image capturing device 20 may also be a device with a camera function that can function as the image capturing device 20 in addition to a display function, such as a smartphone, tablet terminal, or VR / AR goggles.
[0023] The image processing device 10 is configured, for example, by a server device, a virtual server built on a cloud computing service, a PC, etc., and realizes various functions related to VR / AR image processing by using a CPU (Central Processing Unit) not shown to execute middleware such as an OS (Operating System), DBMS (DataBase Management System), and web server program that are deployed on memory from a storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive), as well as software running on top of them.
[0024] This image processing device 10 has various units, such as a pre-processing unit 11, an image search unit 12, and an image synthesis unit 13, which are implemented as software. It also has image data such as a past video set 14 and a pre-processed past video set 15, which are recorded and stored as a database, a file, or the like.
[0025] The preprocessing unit 11 has a function of performing preprocessing such as foreground / background separation and scene division, as described below, in advance on past videos of past scenes captured in various places and recorded in the past video set 14, in order to provide them for VR / AR processing, and recording the generated image data as the preprocessed past video set 15. The image search unit 12 has a function of acquiring similar past videos from the preprocessed past video set 15 by image search, based on the video of the current scene captured by the imaging device 20. The image synthesis unit 13 has a function of overlaying (synthesizing) the past video obtained by the image search unit 12 on the video of the current scene captured by the imaging device 20, and displaying it on the display device 30.
[0026] All or part of the image search unit 12 and the image synthesis unit 13 may be implemented on a display device 30 configured by a smartphone, VR / AR goggles, etc., rather than on the image processing device 10. The contents of the preprocessing, the composition of the preprocessed past video set 15, the search for past videos, and the process of overlaying (synthesizing) them with the current video will be described later.
[0027] <Image processing example> Figure 2 is a diagram outlining an example of image processing in one embodiment of the present invention. The diagram in the upper left of Figure 2 shows an example of a video (hereinafter sometimes referred to as "current video") of a current scene at a certain location captured by the camera device 20, and the diagram in the upper right of Figure 2 shows an example of a video (hereinafter sometimes referred to as "past video") of a past scene at the same location (for convenience, both are shown as one of consecutive still images (frames) in the video). The past video shows a situation in which the same "castle" as in the current video is captured, but the "cherry blossom branch" captured in the current video is not captured, and instead, "two horses" that existed in the past but are not captured in the current video are captured.
[0028] In this embodiment, as shown in the lower left diagram of FIG. 2, a past video is transformed onto a current video, overlaid, and composited before playback. Using two-dimensional moving images captured live rather than computer graphics as the past video to be overlaid reduces the processing load for selecting and overlaying the past video, speeding up processing, and producing a more realistic image. Note that when overlaying, for example, corresponding feature points (e.g., objects in the image (e.g., vertices of the "castle" in the example of FIG. 2)) are extracted from the current video and the past video to be overlaid, and a two-dimensional transformation matrix (affine transformation matrix) is calculated to match the corresponding feature points by rotating and moving the image. This transformation matrix is applied to the past video to transform (deform) the past video, and the corresponding feature points are overlaid and composited onto the current video. Note that the calculation and computation of the affine transformation matrix can be performed using a known method (e.g., libraries such as OpenCV) in which the source and destination are represented as matrices and solved using the least squares method.
[0029] Instead of overlaying the entire past video onto the current video, only objects that appear in the foreground in the past video may be cut out and composited together, as shown in the diagram at the bottom right of FIG. 2. This diagram shows an example in which only the "two horses" portion of the past video is cut out and composited together. It may be possible to switch between a method of overlaying the entire past video and a method of overlaying only foreground objects in the past video. For example, the method may be switched by a user's instruction, or may be switched depending on the type of display device 30 used by the user (e.g., a smartphone or VR / AR goggles).
[0030] <Data Structure> 3 is a diagram outlining an example of a data structure in one embodiment of the present invention. In this embodiment, when superimposing a video of a past scene (past video 14a) onto a video of a current scene (current video 21), it is necessary to identify the past video 14a that captured the same subject as the current video 21. In this case, the past video 14a is identified by performing a similar image search for the current video 21. This increases the number of cases where the past video 14a to be superimposed can be identified, even when sufficient location information about the location where the current video 21 was captured or the location of the subject cannot be obtained.
[0031] In this embodiment, when performing a similar image search, the search target is not the past video 14a itself, but rather key frames 15c, which are one or more characteristic still images (frames) that represent the past video 14a. This enables efficient similar image search. The key frames 15c are extracted from the past video 14a by the preprocessing unit 11, and are recorded and stored in association with the past video 14a as one piece of data in the preprocessed past video group 15.
[0032] To extract keyframe 15c, preprocessing unit 11 first separates foreground objects from the background of past video 14a, obtaining foreground video 15b, from which only foreground objects have been extracted, and background video 15a, from which only the background has been extracted. These are also recorded and stored in association with past video 14a as one piece of data in preprocessed past video group 15. Because foreground objects in past video 14a are typically objects that existed in the past but do not appear in current video 21, separating them and then performing a similar image search with current video 21 can more efficiently and effectively identify past video 14a.
[0033] Then, for the background video 15a from which the foreground object has been separated, characteristic frames (frames to be superimposed) are extracted as key frames 15c. In this embodiment, the scenes of the background video 15a are divided at the point where there is a major scene change, and the first frame of each scene is extracted as key frame 15c, but this is not limited to this. A frame at the end or in the middle of a scene may also be used as key frame 15c, as long as it can be treated as a characteristic frame representing each scene.
[0034] The decision to split a scene is made, for example, when the change in the six degrees of freedom (6DoF, consisting of three translational degrees of freedom for movement along the axes of a three-dimensional Cartesian coordinate system and three rotational degrees of freedom around the axes) between frame images exceeds a predetermined threshold, but this is not limited to this.
[0035] In the similar image search using the current video 21, a search is performed on key frames 15c extracted from multiple past videos 14a recorded in the past video set 14, and the past video 14a corresponding to the matching key frame 15c is identified. Then, the identified past video 14a or the foreground video 15b separated and extracted from the past video 14a is transformed using the above-mentioned two-dimensional affine transformation matrix and superimposed on the current video 21 for synthesis.
[0036] <Processing flow> 4 is a flowchart outlining an example of the flow of preprocessing in one embodiment of the present invention. In this preprocessing, preprocessing unit 11 extracts key frames 15c for similar image search from each past video 14a stored in advance in past video set 14, and records the extracted key frames as one data item in preprocessed past video set 15.
[0037] First, the pre-processing unit 11 separates the foreground and background of the past video 14a to be processed (S01). That is, a foreground object (object) is cut out from the past video 14a and extracted as the foreground video 15b, and this is then separated from the past video 14a to obtain the background video 15a (Image Matting). This process can be implemented using a number of existing technologies, methods, libraries, etc. that have been put into practical use, such as AI image processing technologies such as CNN (Convolutional Neural Network). The separated background video 15a and foreground video 15b are associated with the past video 14a and recorded as the pre-processed past video group 15. Note that the subsequent processes are performed on the separated background video 15a.
[0038] Once the background video 15a is acquired, the type of video is then determined (S02), and information on the six degrees of freedom for each frame is acquired according to the type of video, and the amount of change between frames is acquired. If the background video 15a is a video with six degrees of freedom information, the amount of change in the six degrees of freedom between frames is acquired directly (S05). If the background video 15a is a stereo video captured with a stereo camera, the six degrees of freedom are estimated using, for example, a known triangulation method (S03), and the amount of change between frames is acquired (S05). If the background video 15a is a monocular video captured with a monocular camera, the six degrees of freedom are estimated by tracking feature points in the video using, for example, a known V-SLAM (Visual Simultaneous Localization and Mapping) method (S04), and the amount of change between frames is acquired (S05).
[0039] Then, when the amount of change in the six degrees of freedom between frames acquired in step S05 exceeds a predetermined threshold, the scene is divided between those frames (S06). The amount of change in the six degrees of freedom may be a simple sum of the amounts of change in each of the six degrees of freedom, or may be a weighted sum of the amounts of change in one or more specific degrees of freedom. For example, among the six degrees of freedom (moving forward / backward, left / right, up / down, tilting forward / backward (pitch), rotating the head left / right (yaw), and tilting the head left / right (roll)) that are common behaviors of a user viewing VR / AR images, yaw is often the most central movement, so a large weight may be assigned to the amount of change in yaw. The degrees of freedom to be weighted and the weighting value may be changed depending on the content of the video. Instead of the sum of the amounts of change in the six degrees of freedom, a scene may be divided when the amount of change in one or more specific degrees of freedom exceeds a predetermined threshold.
[0040] As a simpler method, instead of the amount of change in six degrees of freedom between frames, for example, a simple amount of change (difference) between image data between frames may be obtained, and when the amount of change exceeds a predetermined threshold, the scene may be divided between the frames.
[0041] When the background video 15a is divided into scenes, the first frame of each scene is extracted as a key frame 15c (S07). Then, data of a key frame group consisting of one or more extracted key frames 15c is associated with information on the previous video 14a to be processed from which it was extracted, and recorded as a preprocessed previous video group 15 (S08). The above series of processes are performed on all previous videos 14a to be processed, and the preprocessing is completed.
[0042] 5 is a flowchart outlining an example of the flow of image retrieval and synthesis processing in one embodiment of the present invention. In this image retrieval and synthesis processing, image retrieval unit 12 and image synthesis unit 13 superimpose and synthesize corresponding past video 14 onto video images of a current scene captured by imaging device 20, and display the synthesized image on display device 30.
[0043] First, the image search unit 12 acquires location information of the local area (S11). Here, the local area refers to the location of the image capture device 20 currently capturing the video 21, or the location of the subject currently captured in the video 21, or locations nearby these. If the image capture device 20 is an information processing terminal such as a smartphone equipped with a camera function and a GPS (Global Positioning System) function, the image capture device 20 acquires location information as latitude and longitude information using the GPS function and transmits this information to the image processing device 10 via communication means (not shown), which is then acquired by the image search unit 12. The location information may be acquired from an information processing terminal equipped with a GPS function that is separate from the image capture device 20 located at the local area, or the photographer or user may input the name of a local facility or landmark and acquire the corresponding location information via the Internet or the like.
[0044] Next, the image search unit 12 acquires one or more candidate past videos 14a from the preprocessed past video set 15 based on the acquired local location information (S12). For example, past videos 14a captured within a predetermined distance from the local location information are acquired as candidates. Then, key frames 15c associated with each acquired past video 14a are acquired from the preprocessed past video set 15 (S13). This narrows down the key frames 15c to be used in the similar image search described below. To enable this narrowing down, each past video 14a is assumed to be recorded in association with location information about the shooting location or the subject. Note that if there is no candidate past video 14a corresponding to the current location information, the process ends and the image of the current video 21 captured by the shooting device 20 is displayed as is on the display device 30.
[0045] Once the narrowed-down key frames 15c are acquired, key frames 15c similar to the current video 21 being captured by the imaging device 20 are identified from among them by a similar image search (S14), and the past video 14a corresponding to the identified key frame 15c is selected from the candidates acquired in step S12 (S15). Note that the method of similar image search is not particularly limited, and a known library using AI technology or the like can be used as appropriate.
[0046] Once the past video 14a is selected, the image synthesis unit 13 then extracts corresponding feature points between the key frame 15c identified in step S14 and the image of the current video 21, and calculates a two-dimensional transformation matrix (affine transformation matrix) for image synthesis that matches each corresponding feature point by rotating and moving (S16). The past video 14a selected in step S15 is then transformed (deformed) by applying the affine transformation matrix calculated in step S16, and this is superimposed on the video of the current video 21 for synthesis and playback (S17). The image played back with the past video 14a superimposed on the video of the current video 21 is displayed on the display device 30.
[0047] When playing back the past video 14a, the entire past video 14a may be played back from the beginning, or the past video 14a may be played back from a point corresponding to the key frame 15c identified in step S14. The point corresponding to the key frame 15c may be, for example, the beginning of a scene including the key frame 15c, or the key frame 15c itself.
[0048] Then, in step S17, it is determined whether the amount of deviation (the deviation of each feature point after applying the affine transformation matrix) when the current video 21 and the past video 14a are superimposed and combined exceeds a predetermined threshold (S18). The amount of deviation compared to the threshold may be the sum of the values at each feature point, or may be the deviation for any one feature point. The determination may also be made using both of these. If the amount of deviation is equal to or less than the threshold (No in step S18), the contents of the current video 21 and the past video 14a still match, and the process returns to step S17 to continue playback by combining the past video 14a with the current video 21. On the other hand, if the amount of deviation exceeds the threshold (Yes in step S18), the current video 21 and the past video 14a no longer match, and the process of combining the past video 14a is terminated.
[0049] Thereafter, the current video 21 is displayed as is on the display device 30, and the series of processes from step S11 are repeated to reselect the past video 14a corresponding to the current video 21. Through the series of processes described above, it is possible to continue displaying the current video 21 while appropriately selecting and superimposing an appropriate past video 14a for display.
[0050] As described above, image processing system 1, which is one embodiment of the present invention, uses VR / AR to display a past scene at a location by overlaying (combining) it with a captured image of the current scene. This system effectively and efficiently selects a past scene that is appropriate for overlaying, and reduces the load involved in the overlay process.
[0051] That is, in the image processing system 1 according to one embodiment of the present invention, a 2D past video 14a captured live rather than using CG is used as a past scene to be overlaid on a current video 21. A key frame 15c is extracted and associated with each past video 14a in advance based on the amount of change in three-dimensional position and orientation (six degrees of freedom). Then, candidate past videos 14a are narrowed down from the past video set 14 using position information related to the current video 21, and key frames 15c related to each narrowed-down past video 14a that are similar to the current video 21 are extracted by performing a similar image search between 2D images, and the corresponding past video 14a is identified. A 2D affine transformation matrix is then generated so that the extracted key frames 15c and the feature points of the current video 21 match (i.e., minimize the deviation), and the identified past video 14a is transformed (deformed) using the transformation matrix and overlaid on the current video 21.
[0052] These techniques reduce the processing load of selecting and overlaying the past video 14a to be overlaid on the current video 21, speeding up the processing and producing realistic VR / AR images.
[0053] The invention made by the inventor has been specifically described above based on the embodiments, but it goes without saying that the present invention is not limited to the above embodiments and can be modified in various ways without departing from the spirit of the invention. Furthermore, the above embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the above embodiments with other configurations.
[0054] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a storage device such as a memory, hard disk, or SSD, or in a storage medium such as an IC card, SD card, or DVD.
[0055] In addition, in the above figures, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily show all the control lines and information lines that are actually implemented. In reality, it can be assumed that almost all components are interconnected. [Industrial Applicability]
[0056] The present invention can be used in image processing systems that realize image- and video-based VR / AR. [Explanation of symbols]
[0057] 1...Image processing system, 10...image processing device, 11...preprocessing unit, 12...image search unit, 13...image synthesis unit, 14...past video group, 14a...past video, 15...preprocessed past video group, 15a...background video, 15b...foreground video, 15c...key frame, 20...Filming equipment, 21...Current video, 30…Display device
Claims
1. An image processing system in which a past video in which a corresponding past scene is shot is superimposed on a current video relating to a current scene shot by a shooting device and displayed on a display device by an image processing device, The image processing device includes: a preprocessing unit that extracts one or more keyframes from each of a plurality of past videos, associates the extracted keyframes with the past videos, and records the preprocessed past videos; an image search unit that identifies a first key frame similar to the current video from the preprocessed past video by image search, and identifies a first past video corresponding to the first key frame; an image synthesis unit that converts and superimposes the first past video onto the current video so as to minimize deviations between feature points in the respective video images, plays back the first past video from a position corresponding to the first key frame, and displays the first past video on the display device; An image processing system comprising:
2. 2. The image processing system according to claim 1, The pre-processing unit acquires values of six degrees of freedom, consisting of three degrees of freedom of translation and three degrees of freedom of rotation, for each frame image of each of the past videos, and when the amount of change in the values of the six degrees of freedom between frame images exceeds a predetermined threshold, divides the scenes in each of the past videos between the frame images and extracts frame images representing each scene as the key frames.
3. 3. The image processing system according to claim 2, The pre-processing unit calculates the amount of change in the six degrees of freedom between each frame image of each of the past videos by weighting the amount of change in yaw in the three rotational degrees of freedom among the six degrees of freedom more heavily than the other degrees of freedom.
4. 2. The image processing system according to claim 1, The image search unit identifies the first key frame by image search from among the multiple key frames recorded as the pre-processed past video, narrowing down the key frames based on location information related to the current video.
Citation Information
Patent Citations
Foil-wound transformer
JP1989020605A
Method and device for retrieving image, and storage medium
JP1999338876A
Video synchronizing apparatus and video synchronizing method
JP2021044849A
Augmented reality and virtual reality location-based attraction simulation playback and creation system and processes for simulating past attractions and preserving present attractions as location-based augmented reality and virtual reality attractions
US10127730B2
Image processing device, image processing method, image processing program and recording medium
US20150229840A1