Information processing apparatus, information processing method, and program

By acquiring and embedding feature point groups from a broader space as metadata, the method addresses the challenge of insufficient feature points in image arrangement, enhancing positional accuracy in head-mounted displays.

JP2025099665APending Publication Date: 2025-07-03CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216513
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for determining the arrangement of images in space, such as those used in head-mounted displays, struggle when insufficient feature points are obtained from videos or depth maps, particularly when the subject is prominently shown.

Method used

The method involves acquiring video data and a feature point group from a wider space than the shooting range, and associating this information with the video data as metadata to facilitate accurate positioning and orientation of the image in space.

Benefits of technology

This approach allows for easier determination of the image's position and orientation in space, even when feature points are scarce, by utilizing a wider range of feature points beyond the immediate shooting area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099665000001_ABST
    Figure 2025099665000001_ABST
Patent Text Reader

Abstract

To facilitate determination of the position of an image to be arranged in a space.SOLUTION: An information processing apparatus 100 acquires video data for viewing, and acquires a feature point group from an image in a range of a space wider than a photographing range at a position where the video data for viewing is photographed. The information processing apparatus 100 adds position information of the feature point group associated with the position where the video data for viewing is photographed, to the video data for viewing, and stores it as a video file.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] In recent years, as a display device for viewers to enjoy stereoscopic video, a display device such as a head-mounted display (HMD) that is worn on the head and arranges and displays a video in space according to the position and orientation of the viewer has been proposed. In such a display device, a process of determining where in space to arrange and display the video is performed. Patent Document 1 discloses a method of determining the arrangement of a video by adjusting the position and orientation of the video so that a plane or feature point specified from a depth video given to the video matches the plane or feature point in the reproduction environment.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the method disclosed in Patent Document 1, when sufficient feature points cannot be obtained from the video or depth map, such as when the subject is prominently shown in the captured video, it may be difficult to determine the arrangement of the video.

[0005] Therefore, an object of the present invention is to make it easier to determine the position of an image arranged in space.

Means for Solving the Problems

[0006] The present invention is characterized by comprising: a first acquisition means for acquiring a moving image; a second acquisition means for acquiring a feature point group from an image of a range within a space wider than a shooting range at a position where the moving image is shot; and an imparting means for imparting position information of the feature point group associated with the shot position to the moving image.

Effect of the Invention

[0007] According to the present invention, it is possible to easily determine the position of an image arranged in a space.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0009] Hereinafter, this embodiment will be described with reference to the accompanying drawings.

[0010] FIG. 1 is a block diagram showing the hardware configuration of an information processing apparatus 100 according to this embodiment. The information processing apparatus 100 includes a CPU 101, a RAM 102, and a ROM 103. Further, the information processing apparatus 100 includes an HDD (hard disk drive) interface (hereinafter, the interface is referred to as "I / F") 104, an input I / F 106, an output I / F 108, and a network I / F 110. These respective blocks are interconnected via a system bus 112.

[0011] The CPU 101 controls the overall operation of the information processing apparatus 100. By the CPU 101 using the RAM 102 as a work memory and executing programs stored in the ROM 103, the HDD 105, etc., the processing of the flowchart described later is realized. The HDD I / F 104 connects a secondary storage device such as the HDD 105 or an optical disk drive. The HDD I / F 104 is an I / F such as Serial ATA (SATA). The CPU 101 can read data from the HDD 105 and write data to the HDD 105 via the HDD I / F 104. The CPU 101 expands the data stored in the HDD 105 into the RAM 102, and conversely, saves the data expanded in the RAM 102 to the HDD 105.

[0012] The input I / F 106 connects an input device 107 such as a keyboard or a mouse and a camera 113, etc. The input I / F 106 is an I / F such as USB or IEEE 1394. The CPU 101 receives input data from the input device 107 via the input I / F 106. Also, the CPU 101 receives video data captured by the camera 113 via the input I / F 106.

[0013] The output I / F 108 connects to an output device 109 such as a display. The output I / F 108 is a video output I / F such as DVI or HDMI (registered trademark). The CPU 101 causes the output device 109 to display video data via the output I / F 108. The network I / F 110 is an I / F for communicating with an external device connected to the network. The CPU 101 transmits and receives data to and from an external server 111 or the like connected to the network via the network I / F 110.

[0014] FIG. 2 is a block diagram showing the functional configuration of the information processing apparatus 100 according to the present embodiment. The information processing apparatus 100 functions as a first acquisition unit 201, a calculation unit 202, a second acquisition unit 204, a third acquisition unit 205, a generation unit 206, and an output unit 207 when the CPU 101 executes programs stored in the ROM 103, the HDD 105, and the like.

[0015] The third acquisition unit 205 acquires video data for viewing. The third acquisition unit 205 acquires, for example, video data captured by a camera attached to an HMD (head-mounted display) via the input I / F 106 while a person wearing the HMD is moving within a space. The first acquisition unit 201 acquires video data within the space where the video data for viewing is captured via the input I / F 106. The video data acquired by the first acquisition unit 201 is used to acquire information on a feature point group.

[0016] In this embodiment, it is assumed that the imaging device that captures video data for viewing (first video data) and the imaging device that captures video data (second video data) for acquiring information on a feature point group are configured with different cameras. As a result, imaging devices with different positions, orientations, and viewing angles can be used according to each application. Note that the imaging device that captures the first video data and the imaging device that captures the second video data may be configured with the same camera. Also, the imaging device that captures the first video data and the imaging device that captures the second video data may each be configured with a plurality of cameras.

[0017] The calculation unit 202 extracts image feature points and image feature amounts at the image feature points from the video frames of the video data acquired by the first acquisition unit 201. Then, based on the extracted information and the three-dimensional feature point map, the calculation unit 202 calculates the camera position and orientation by SLAM (Simultaneous Localization and Mapping) processing and updates the three-dimensional feature point map.

[0018] The second acquisition unit 204 acquires depth data. For the acquisition of depth data, an active sensing depth camera such as a ToF camera may be used, or it may be acquired by other methods such as performing stereo matching with a stereo camera. Depth data is an example of depth information. The generation unit 206 generates video data with metadata attached based on the data obtained by the calculation unit 202, the data obtained by the second acquisition unit 204, and the data obtained by the third acquisition unit 205. The video data to which metadata is attached is video data for viewing and is an example of a moving image.

[0019] The output unit 207 displays the video data with metadata on the output device 109 or stores it as a video file in the HDD 105. When displaying the video data with metadata, by using the feature point group attached to the video data, it becomes easy to determine the position and orientation of the video data to be displayed.

[0020] Using FIG. 4, an example of displaying video data at the shooting location will be described. FIG. 4(a) represents the real space where the video data is not being displayed. FIG. 4(b) represents the state when the video data for viewing is superimposed and displayed at the location of FIG. 4(a). Region 403 is the region where the video data for viewing is superimposed, and it is arranged so that the video data can be seen at the same position as when the video data for viewing was shot.

[0021] FIG. 5 represents the feature point group obtained by performing SLAM processing on the video data for obtaining the information of the feature point group. The shooting range of the video data for obtaining the information of the feature point group is wider than the shooting range of the video data for viewing. Therefore, the feature point group shown in FIG. 5 includes the feature point groups other than the subject of the video data for viewing. In the present embodiment, not only the feature points included in region 403 such as feature point 501 but also the feature points not included in region 403 such as feature point 503 are given as metadata. Thereby, when displaying the video data with metadata, it becomes easier to calculate the shooting position of the video data, and it becomes easier to determine the position and orientation for arranging the video data.

[0022] FIG. 3 is a flowchart showing the processing executed by the information processing apparatus 100 according to the present embodiment. Hereinafter, each step (process) is represented by prefixing S. For example, while a person wearing an HMD equipped with a camera for shooting video data for viewing, a camera for obtaining a feature point group, and a camera for obtaining depth data is moving, the CPU 101 starts by obtaining the video data shot by each camera.

[0023] In S301, the CPU 101 executes SLAM processing based on the video data for obtaining the information of the feature point group. The CPU 101 continues to execute the SLAM processing in a separate thread until it receives an end command, sequentially obtains video frames, and sequentially calculates and saves the camera position and orientation and the three-dimensional feature point group. Details will be described later. Details of this step will be described later with reference to FIG. 6.

[0024] In S302, the CPU 101 waits until there is a command to start shooting video data for viewing. When there is a command to start shooting, it proceeds to S303. In S303, the CPU 101 acquires a video frame and internal camera parameters from a camera that captures video data for viewing.

[0025] In S304, the CPU 101 acquires depth frame data corresponding to the video frame acquired in S303. The depth frame data is acquired from a depth camera. The CPU 101 acquires the depth frame data in combination with internal camera parameters such as the viewing angle of the depth camera. In S305, the CPU 101 repeatedly executes the processes of S303 to S304 until there is a command to stop shooting video data for viewing, and sequentially acquires the video frames and depth frame data captured during movement. When there is a command to stop shooting, the CPU 101 proceeds to S306.

[0026] In S306, the CPU 101 generates video data with metadata. Details of this step will be described later with reference to FIG. 7. In S307, the CPU 101 records the video data with metadata as a video file on the HDD 304.

[0027] In S308, the CPU 101 determines whether the three-dimensional feature point group is sufficient. If it is sufficient, the processing of this flowchart ends. If it is not sufficient, it waits until a certain period of time has elapsed. During the waiting period until a certain period of time has elapsed, the SLAM process continues to be executed and the three-dimensional feature point group is updated. When it becomes sufficient after waiting for a certain period of time, the CPU 101 ends the processing of this flowchart. When it does not become sufficient even after waiting for a certain period of time, it proceeds to S306. As a method for determining whether the three-dimensional feature point group is sufficient, for example, there is a method of determining whether the number of feature points is greater than a threshold value. Note that the determination method is not limited to this, and it may be determined based on whether three-dimensional feature points are detected in a certain proportion or more of the partitions obtained by dividing the surroundings at regular angles from the camera position at a specific frame timing.

[0028] (SLAM process) Figure 6 is a flowchart showing an example of the procedure of the SLAM process executed in S301. In S601, the CPU 101 acquires a video frame and camera internal parameters. For the acquisition of the video frame in this step, a camera different from the camera that captures video data for viewing is used. Specifically, a wide-angle camera is used so that feature points can be extracted from a range wider than the imaging range of the video data for viewing. The camera internal parameters are information including parameters related to the camera, such as internal parameters indicating the characteristics of the camera.

[0029] In S602, the CPU 101 extracts image feature points and image feature amounts from the video frame acquired in S601. The image feature points are points having characteristic information in the image and are represented by two-dimensional positions. The image feature amounts are those in which the features at the image feature points are described by vectors. For the extraction of the image feature points and image feature amounts, ORB, SIFT, SURF, etc. can be used. Also, the CPU 101 performs object recognition on the video frame, and when image feature points are detected from the area where an object is detected, an object ID is assigned to the image feature points. The object ID is an ID for identifying from which object the detected feature points are. The same object ID is assigned to the image feature points detected from the same object. For object recognition, methods such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) can be used.

[0030] In S603, the CPU 101 calculates the camera position and orientation and the three-dimensional feature point group from the information extracted in S602. In this embodiment, Visual SLAM is used to calculate the camera position and orientation and the three-dimensional feature point group. The object ID assigned to the image feature points in S602 is associated with the three-dimensional feature point group obtained by Visual SLAM. As a result, three-dimensional feature point group data associated with the object ID is generated. Since Visual SLAM is a well-known technique, its description is omitted. Note that the shape information of the space may be estimated from the three-dimensional feature point group. For example, the CPU 101 represents a rectangular parallelepiped by parameters of width, depth, height, and rotation, obtains a rectangular parallelepiped overlapping with the point group by RANSAC, and acquires the shape of the room. When the direction of gravity is known, the rotation can also be parameterized by the orientation of rotation around the direction of gravity. In this case, the information obtained from the magnetic sensor can be used as a reference for the orientation. The room is not limited to a rectangular parallelepiped. For example, a plurality of planes may be estimated from the feature point group and combined to represent it.

[0031] In S604, the CPU 101 repeatedly executes the processes of S601 to S603 until a stop command for the SLAM process is given, sequentially acquires video frames captured during movement, and continues to calculate the camera position and orientation and the three-dimensional feature point group. When there is a stop command for the SLAM process, the CPU 101 ends the SLAM process and ends the series of processes in the flowchart of FIG. 3.

[0032] (Video data generation process) FIG. 7 is a flowchart showing an example of a procedure for generating video data including metadata according to this embodiment.

[0033] In S701, the CPU 101 assigns depth data as metadata to the video data. The depth data includes the internal camera parameters of the depth camera in addition to the depth frame data. In S702, the CPU 101 attaches information on the camera position and orientation and the three-dimensional feature point group as metadata to the video data. The information on the three-dimensional feature point group includes the type of image feature quantity, the image feature quantity, and the object ID, in addition to the three-dimensional position information of the three-dimensional feature points. The CPU 101 converts the three-dimensional position information of the three-dimensional feature points based on the camera position and orientation of the initial frame at the start of shooting into a coordinate system with the gravity direction aligned with one axis and stores it. In this way, the CPU 101 associates the position and orientation of the camera that shot the video data with the three-dimensional position information of the three-dimensional feature point group. By the coordinate system conversion, when displaying the video data with metadata, the video data can be aligned with the image in the real space by arranging the video data with the origin of the coordinate system as the camera position.

[0034] Also, the CPU 101 may attach the room shape data as metadata. Also, the CPU 101 may prepare in advance a table in which information on the ease of movement of each object is set, and attach information on the ease of movement of the object to the object ID. For example, the ease of movement is expressed as a value from 0 to 1. Since the wall of the object house is less likely to move, it is set to a value of 0.02, and since the toy is more likely to move, it is set to a value of 0.9. Thereby, when displaying the video data with metadata, the feature points used for alignment can be determined based on the ease of movement of the object.

[0035] Also, the CPU 101 may store the three-dimensional feature point group in the order based on the distance to the object. The distance to the object can be calculated from the point group with the same object ID. For calculating the distance to the object, the closest point belonging to the same object ID may be used, or the farthest point may be used. By arranging the three-dimensional feature point group in the order of distance, when it is desired to arrange the video data according to the nearby object, instead of loading all the feature points, the feature points can be sequentially loaded in the order of proximity, and only the loaded feature points up to the middle can be used for alignment.

[0036] In S703, the CPU 101 attaches the camera position and orientation as metadata to the video data. Similar to the three-dimensional feature point group, the camera position and orientation are converted into a coordinate system with the gravitational direction aligned with one axis based on the camera position and orientation of the initial frame at the start of shooting and stored. Also, when the camera for capturing video data for viewing (the first camera) is different from the camera for capturing video data for acquiring information on the feature point group (the second camera), the position and orientation data of the second camera are converted into the position and orientation data of the first camera.

[0037] (Data Structure of Video File) FIG. 8 shows an example of the data structure of the video file according to the present embodiment and a specific example of the data. The video file 800 shown in FIG. 8 is composed of stored data information 801, feature point group data 802, spatial shape data 803, object data 804, depth data 805, camera parameters 806, and video data 807.

[0038] The stored data information 801 is information on what data is stored in the video file 800. FIG. 9 shows an example of the bit assignment of the stored data information 801. Here, the stored data information 801 has a 32-bit value, and each bit indicates that the target data is stored if it is "1", and indicates that the target data is not stored if it is "0". In the example of FIG. 9, video data is assigned to b0, feature point group data to b1, object data to b3, depth data to b4, and camera parameters to b5. For example, when three types of data, namely video data, feature point group data, and camera parameters, are stored in the video file 800 shown in FIG. 8, the bits of b0, b1, and b5 are "1", and the other bits are "0".

[0039] The feature point group data 802 is information on the three-dimensional feature point group calculated by the SLAM process. The feature point group data 802 holds the object ID to which the feature point belongs and information on the feature amount in association with the three-dimensional position information of each feature point. The spatial shape data 803 is the shape data of the space estimated from the three-dimensional feature point group. The spatial shape data 803 holds information representing the shape of the room, such as width, height, depth, and orientation. The object data 804 is the information of the object obtained by the object recognition process. The object data 804 holds the object name and information on ease of movement (possibility of movement) in association with the object ID.

[0040] The depth data 805 holds the data from the first frame to the Nth frame. N is a natural number of 1 or more. The camera parameters 806 correspond to the information of one video frame per row and hold the camera position and camera orientation obtained by the SLAM process at the time of acquisition of the video frame. The video data 807 holds the data from the first frame to the Nth frame. The video data 807 includes video frames captured by one or more cameras.

[0041] Also, for each piece of data to be stored, metadata indicating what kind of data it is is attached. FIG. 10 shows an example of the metadata of the video data of each camera. The metadata of the video data of each camera includes information on the camera ID, format, width of the image, height of the image, and bit depth.

[0042] FIG. 11(a) shows an example of the metadata of the camera parameters 806. The metadata of the camera parameters 806 includes information on the camera ID, storage parameters, attitude parameters, position parameters, field angle parameters, distance to the shooting target, focal length, width of the image, height of the image, aperture value, shutter speed, and ISO. FIG. 11(b) shows the bit assignment of the camera parameters 806 to be stored. The storage parameters in FIG. 11(a) have a 32-bit value, and each bit indicates that the target data is stored if it is "1" and that the target data is not stored if it is "0". As shown in FIG. 11(b), each piece of information in FIG. 11(a) is assigned to each bit.

[0043] FIG. 12(a) shows an example of the metadata of the feature point group data 802. The metadata of the feature point group data 802 includes information on the number of point groups, storage parameters, and types of feature quantities. FIG. 12(b) shows the bit assignment of the feature point group data 802 to be stored. In the example of FIG. 12(b), the object ID is assigned to b0 and the feature quantity is assigned to b1, respectively. FIG. 13(a) shows an example of the metadata of the object data 804. The metadata of the object data 804 includes information on the number of objects and storage parameters. FIG. 13(b) shows the bit assignment of the object data to be stored. In the example of FIG. 13(b), the object name is assigned to b0 and the possibility of movement is assigned to b1, respectively. FIG. 14(a) shows an example of the metadata of the spatial shape data 803. The metadata of the spatial shape data 803 includes information on storage parameters, width, height, depth, and orientation of the space. FIG. 14(b) shows the bit assignment of the spatial shape data 803 to be stored. As shown in FIG. 13(b), each piece of information in FIG. 14(a) is assigned to each bit.

[0044] As described above, the video data and the metadata are defined and collectively filed. Note that the camera parameter 806 may be one for one video file, or the camera parameter may be provided in units of frames. By having the camera parameter in units of frames, it is possible to cope with cases where the position and orientation of the camera change during shooting.

[0045] (Specific example of video file) Next, a specific example of the video file of this embodiment that complies with the ISO BMFF (ISO Base Media File Format ISO / IEC 14496-12 MPEG-4 Part 12) standard will be described. In the ISO BMFF standard, a file is composed of units called "boxes". FIG. 15(a) shows the structure of a box. The box 1500 is composed of a region 1501 for storing size information, a region 1502 for storing type information, and a region 1503 for storing data. Also, as shown in FIG. 15(b), it is possible to have a structure in which the box 1510 is further included as data inside the box 1500.

[0046] FIG. 16 shows an example of the internal structure of a video file that complies with ISO BMFF. The video file 1600 that complies with ISO BMFF is composed of boxes such as ftyp 1601, moov 1602, camp 1603, feap 1604, obje 1605, spas 1606, Camm 1607, and mdat 1608. Each box will be described below.

[0047] The ftyp box (File Type Compatibility Box) 1601 is the box placed at the beginning of the file. In the ftyp box 1601, information on the file format, information indicating the version of the box, information on compatibility with other file formats, information on the name of the manufacturer that created the file, etc. are stored. The above-mentioned stored data information 801 indicating the type of each data stored in the video file may be stored in the ftyp box 1601.

[0048] The moov box (Movie Box) 1602 is a box that clarifies how data is stored in the file, and information such as a time axis and addresses for managing media data is stored. In the mdat box (Media Data Box) 1608, media data such as moving images and audio is stored. The depth data 805 and video data 807 may be stored in this mdat box. By storing information on how the data is stored in the mdat box 1608 in the moov box 1602, access to these media data becomes possible.

[0049] In the camm (Camera motion Metadata) box 1607, metadata related to the movement of the camera is stored. The camera position and orientation data of the camera parameters 806 may be stored in this camm box. The ftyp box 1601, moov box 1602, and mdat box 1608 are boxes commonly provided in files compliant with ISO BMFF. The camm box 1607 is a box used in the VR180 video format and the like. In contrast, each of the boxes camp1603, feap1604, obje1605, and spas1606 is a box specific to the video file of this embodiment. Hereinafter, specific examples will be given to explain the boxes specific to the video file of this embodiment.

[0050] The camp box (Camera Parameter Data Box) 1603 can be set for each video data, and holds information indicating what data is stored as camera parameters and each value of the corresponding camera parameters. The feap (Feature Points Box) 1604 can be set once for a video file, and holds information indicating what data is stored as feature point group data and each value of the corresponding feature point group. The obje (Object Box) 1605 can be set once for a video file, and holds information indicating what data is stored as object data and each value related to the corresponding object. The spas (Space Shape Box) 1606 holds information indicating what data is stored as data representing the shape of the photographed space, and each value related to the corresponding shape.

[0051] In this embodiment, the ISO BMFF standard is described as an example, but the format of the video file is not limited to this. For example, other standards such as HEIF (High Efficiency Image File Format) and MiAF (Multi-Image Application Format) that are compatible with ISO BMFF may also be used. Also, the values and expressions of the format are not limited to the above description. Further, at least one of the boxes camp1603, feap1604, obje1605, and spas1606 shown in FIG. 16 may be stored in the moov box 1602.

[0052] Using the video data including the metadata according to this embodiment, the specific content of the process when aligning the position of the image of the real space (for example, FIG. 4(a)) and the video data will be described. First, the CPU 101 acquires the current image of the real space from a camera 113 or the like via the input I / F 106. Then, the CPU 101 acquires the camera position and orientation and the position information of the three-dimensional feature points from the acquired image. Then, the CPU 101 reads all of the feature point group data 802 as metadata. Then, the CPU 101 obtains a transformation that matches the position of the three-dimensional feature points acquired from the current image of the real space and the position of the read feature point group data 802 using ICP or the like, and calculates the origin of the read feature point group data 802 in the real space. ICP means Iterative Closest Point. Then, the CPU 101 arranges the video data based on the camera parameters (camera position and orientation) with the calculated origin as a reference.

[0053] As described above, according to the information processing apparatus of the present embodiment, a feature point group obtained from an image of a space wider than the shooting space of the video data to be stored can be embedded in the video data as metadata. Therefore, even when the number of feature points obtained from the shooting space of the video data is small, the feature point group of the shooting space not shown in the video data can be made available, making it easier to determine the arrangement position of the video data to be displayed.

[0054] [Other Embodiments] In the above-described embodiment, monocular video data with depth data is used, but video data with parallax instead of monocular may be used. Video data with parallax is two video data with parallax imaged by a parallax imaging device including two lenses. Also, monocular video data without depth data may be used.

[0055] Also, in the above-described embodiment, in S308 of FIG. 3, it is determined whether or not the three-dimensional feature point group is sufficient, and if it is not sufficient, the process waits until it becomes sufficient and updates the feature point group data of the metadata. However, the flowchart of FIG. 3 may be terminated without performing the above determination.

[0056] Also, some or all of the functions of the blocks and steps in FIGS. 2 and 3 of the above-described embodiment may be implemented by hardware such as an ASIC or an electronic circuit. Also, a program for realizing some or all of the steps in FIG. 3 can be supplied to a system or device via a network or a storage medium, and can also be realized by a process in which one or more processors in the computer of the system or device read and execute the program.

[0057] The disclosure of each of the above-described embodiments includes the following configurations, methods, and programs. (Configuration 1) A first acquisition means for acquiring a moving image, A second acquisition means for acquiring a feature point group from an image of a range within a space wider than the shooting range at the position where the moving image was shot, An imparting means for imparting the position information of the feature point group associated with the photographed position to the moving image; An information processing apparatus characterized by having the above. (Configuration 2) The information processing apparatus according to Configuration 1, wherein the position information of the feature point group is converted so that the position and orientation of the imaging device that captured the moving image are used as a reference. (Configuration 3) The information processing apparatus according to Configuration 1 or 2, wherein the imaging device that captures the moving image is different from the imaging device that captures an image for acquiring the feature point group. (Configuration 4) The second acquisition means acquires object information from the image for acquiring the feature point group, The information processing apparatus according to any one of Configurations 1 to 3, wherein the imparting means associates object information with each feature point of the feature point group. (Configuration 5) The second acquisition means acquires the distance from the imaging device that captures the moving image to the object, The information processing apparatus according to Configuration 4, wherein the imparting means uses an order based on the distance of the object associated with each feature point of the feature point group when imparting the position information of the feature point group. (Configuration 6) The information processing apparatus according to Configuration 4 or 5, wherein the object information includes information regarding the possibility of movement. (Configuration 7) The second acquisition means acquires information on the shape of the shooting space from the image for acquiring the feature point group, The information processing apparatus according to any one of Configurations 1 to 6, wherein the imparting means imparts the information on the shape of the shooting space to the moving image. (Configuration 8) The information processing apparatus according to Configuration 7, wherein the information on the shape of the shooting space includes information on width, height, and depth. (Configuration 9) The information processing apparatus according to configuration 7 or 8, characterized in that the information on the shape of the imaging space includes the information on the orientation. (Configuration 10) A third acquisition means for acquiring the depth information of the moving image, The information processing apparatus according to any one of configurations 1 to 9, characterized in that the attaching means attaches the depth information to the moving image. (Configuration 11) The information processing apparatus according to any one of configurations 1 to 10, characterized in that the second acquisition means acquires the position information of the feature point group by performing SLAM (Simultaneous Localization and Mapping) processing. (Configuration 12) A fourth acquisition means for acquiring an image of the real space, A fifth acquisition means for acquiring the camera position and orientation and the position information of the feature point group from the image of the real space, Control means for controlling to perform alignment of the image of the real space and the moving image based on the position information of the feature point group given by the attaching means and the camera position and orientation and the position information of the feature point group acquired by the fifth acquisition means, The information processing apparatus according to any one of configurations 1 to 11, characterized by having the above. (Configuration 13) The information processing apparatus according to any one of configurations 1 to 12, characterized in that the attaching means attaches the position and orientation of the imaging device that captured the moving image to the moving image. (Configuration 14) The information processing apparatus according to any one of configurations 1 to 13, characterized in that the moving image to which the position information of the feature point group is attached by the attaching means is stored as a moving image file. (Method) A first acquisition step of acquiring a moving image, A second acquisition step of acquiring a feature point group from an image of a range within a space wider than the imaging range at the position where the moving image was captured, An attaching step of attaching the position information of the feature point group associated with the captured position to the moving image, An information processing method characterized by including (Program) causing a computer to function as a first acquisition means for acquiring a moving image, a second acquisition means for acquiring a feature point group from an image of a range within a space wider than a shooting range at a position where the moving image was shot, and an assignment means for assigning position information of the feature point group associated with the shot position to the moving image, and a program for causing the computer to function as such.

Claims

1. a first acquisition means for acquiring a moving image; a second acquisition means for acquiring a feature point group from an image of a range within a space wider than a shooting range at the position where the moving image was shot; an imparting means for imparting position information of the feature point group associated with the shot position to the moving image; An information processing apparatus, characterized by comprising the above.

2. The information processing apparatus according to claim 1, wherein the position information of the feature point group is converted so that the position and orientation of the imaging device that shot the moving image serve as a reference.

3. The information processing apparatus according to claim 1, wherein the imaging device for imaging the moving image and the imaging device for imaging an image for acquiring the feature point group are different.

4. The second acquisition means acquires object information from an image for acquiring the feature point group, The information processing apparatus according to claim 1, wherein the imparting means associates object information with each feature point of the feature point group.

5. The second acquisition means acquires the distance from the imaging device that images the moving image to the object, The information processing apparatus according to claim 4, wherein when imparting the position information of the feature point group, the imparting means uses an order based on the distance of the object associated with each feature point of the feature point group.

6. The information processing apparatus according to claim 4, wherein the object information includes information regarding the possibility of movement.

7. The second acquisition means acquires information on the shape of the shooting space from an image for acquiring the feature point group, The information processing apparatus according to claim 1, wherein the imparting means imparts the information on the shape of the shooting space to the moving image.

8. The information processing apparatus according to claim 7, wherein the information on the shape of the shooting space includes information on width, height, and depth.

9. The information processing apparatus according to claim 7, wherein the information on the shape of the shooting space includes information on orientation.

10. a third acquisition means for acquiring depth information of the moving image; The information processing apparatus according to claim 1, wherein the imparting means imparts the depth information to the moving image.

11. The information processing apparatus according to claim 1, wherein the second acquisition means acquires the position information of the feature point group by performing SLAM (Simultaneous Localization and Mapping) processing.

12. a fourth acquisition means for acquiring an image of the real space; a fifth acquisition means for acquiring camera position and orientation and position information of a feature point group from the image of the real space; control means for controlling to perform alignment between the image of the real space and the moving image based on the position information of the feature point group given by the giving means and the camera position and orientation and the position information of the feature point group acquired by the fifth acquisition means; The information processing apparatus according to claim 1, characterized by comprising the above.

13. The information processing apparatus according to claim 1, wherein the giving means gives the position and orientation of the imaging device that captured the moving image to the moving image.

14. The information processing apparatus according to claim 1, characterized in that the moving image to which the position information of the feature point group is given by the giving means is stored as a moving image file.

15. a first acquisition step of acquiring a moving image; a second acquisition step of acquiring a feature point group from an image of a range in a space wider than the shooting range at the position where the moving image was shot; a giving step of giving the position information of the feature point group associated with the shot position to the moving image; An information processing method characterized by including the above.

16. A computer, a first acquisition means for acquiring a moving image; a second acquisition means for acquiring a feature point group from an image of a range in a space wider than the shooting range at the position where the moving image was shot; a giving means for giving the position information of the feature point group associated with the shot position to the moving image, A program that functions as.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    WO2018131238A1