Information processing device and information processing method
The information processing device and method address the challenge of tracking objects in head-mounted displays by using feature points and key frames to maintain accurate position and orientation, improving safety and immersion in virtual reality.
Patent Information
- Application Number
- JP2022016495
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-04
- Publication Date
- 2026-03-02
- Estimated Expiration
- 2042-02-04
AI Technical Summary
Existing technologies struggle to accurately and continuously track the position and orientation of objects, particularly in head-mounted displays, leading to potential collisions and reduced immersion in virtual reality environments.
An information processing device and method that extracts feature points from captured images, designates key frames for reference, and classifies and discards key frame data using specific discard rules to maintain accurate tracking over time.
Enables continuous and accurate tracking of the position and orientation of objects, enhancing safety and immersion in virtual reality experiences by preventing collisions and ensuring seamless movement within the displayed environment.
Smart Images

Figure 0007822193000001 
Figure 0007822193000002 
Figure 0007822193000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and an information processing method for identifying the state of an object based on a captured image. [Background technology]
[0002] Image display systems that allow users wearing head-mounted displays to view a target space from any viewpoint are becoming widespread. For example, there is known electronic content that realizes virtual reality (VR) by displaying a virtual three-dimensional space on the head-mounted display in accordance with the user's line of sight. In addition, a walk-through system has also been developed that allows users wearing a head-mounted display to virtually walk around within a space displayed as a video by physically moving around.
[0003] Technology that tracks the position and posture of an object that is allowed to move freely and performs corresponding information processing is needed not only in head-mounted displays but also in various fields such as autonomous mobile robots and automobiles (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-149621 Summary of the Invention [Problem to be solved by the invention]
[0005] For users wearing a head-mounted display who cannot see the real world, it is desirable to be able to accurately identify the relative positions of the user and surrounding objects and provide appropriate warnings and movement restrictions so that they can enjoy content without the risk of colliding with them. Avoiding such dangers and guiding the user in the correct direction are necessary in a variety of fields, as mentioned above. Furthermore, in the case of head-mounted displays, it is necessary to continuously and accurately track the position and posture of the user's head in order to enhance the sense of immersion and realism in the displayed world, such as with VR technology, which changes the field of view according to the user's movements.
[0006] The present invention has been made in view of the above problems, and its purpose is to provide a technique that can accurately and continuously track the position and orientation of an object using captured images. [Means for solving the problem]
[0007] To solve the above problems, one aspect of the present invention relates to an information processing device, which includes: a status information acquisition unit that extracts feature points from the latest frame of a video being captured, and acquires status information about the position and orientation of a device equipped with a camera capturing the video based on the relationship between corresponding feature points in a past frame to be compared and points on a subject represented by the feature points; a registration information generation unit that defines a frame from the latest frame that satisfies a predetermined condition as a key frame to be used as a reference for past frames, classifies and registers the frame into one of multiple groups with different discard rules, and discards any of the registered key frame data in accordance with the discard rule; and a registration information storage unit that stores the key frame data together with classification information.
[0008] Another aspect of the present invention relates to an information processing method, which includes the steps of: extracting feature points from a latest frame of a video being captured, and acquiring state information on the position and attitude of a device equipped with a camera capturing the video based on the relationship between corresponding feature points in a past frame to be compared and points on a subject represented by the feature points; defining a frame among the latest frames that satisfies a predetermined condition as a key frame to be used as a reference for past frames, and classifying and registering the key frame in one of a plurality of groups having different discard rules; discarding data of one of the registered key frames in accordance with the discard rule; and storing the key frame data together with classification information in a storage device.
[0009] Any combination of the above components, or conversion of the present invention between a system, a computer program, a recording medium on which a computer program is readably recorded, a data structure, etc., are also valid aspects of the present invention. [Effects of the Invention]
[0010] According to the present invention, it is possible to continue to accurately track the position and orientation of an object using captured images. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram showing an example of the appearance of a head-mounted display according to an embodiment of the present invention; [Figure 2] 1 is a diagram illustrating an example of the configuration of an image display system according to an embodiment of the present invention. [Figure 3] 1 is a diagram for explaining an example of an image world that the image generating device displays on a head-mounted display in this embodiment. FIG. [Figure 4] This is a diagram outlining the principles of Visual SLAM. [Figure 5] FIG. 2 is a diagram showing an internal circuit configuration of the image generating device according to the present embodiment. [Figure 6]FIG. 2 is a diagram showing an internal circuit configuration of a head-mounted display according to the present embodiment. [Figure 7] 1 is a block diagram showing functional blocks of an image generating device according to an embodiment of the present invention. [Figure 8] 1A and 1B are diagrams showing examples of frame images captured by a stereo camera and key frame data obtained therefrom in this embodiment. [Figure 9] 3 is a diagram illustrating an example of a data structure of registration information stored in a registration information storage unit in the present embodiment. FIG. [Figure 10] 10A and 10B are diagrams illustrating examples of space division for evaluating the spatial coverage of key frames in the present embodiment. [Figure 11] 10 is a flowchart showing a processing procedure in which the image generating device sets a play area in the present embodiment. [Figure 12] 12 is a bird's-eye view illustrating an example of a change in state of the head-mounted display during the processing period from S14 to S18 in FIG. [Figure 13] 10 is a flowchart showing a processing procedure in which the image generating device executes an application in the present embodiment. [Figure 14] 10A and 10B are diagrams illustrating transitions of key frames in the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] This embodiment relates to a technology for identifying the position and orientation of a moving object equipped with a camera by analyzing an image captured by the camera. To that extent, the type of moving object is not particularly limited, but hereinafter, a head-mounted display will be described as a representative example. In the following description, at least one of the position and orientation of a head-mounted display may be collectively referred to as the "state" of the head-mounted display.
[0013] 1 shows an example of the appearance of a head-mounted display 100. In this example, the head-mounted display 100 is made up of an output mechanism unit 102 and a wearing mechanism unit 104. The wearing mechanism unit 104 includes a wearing band 106 that, when worn by a user, goes around the head and secures the device in place.
[0014] The output mechanism unit 102 includes a housing 108 shaped to cover the left and right eyes when the user wears the head-mounted display 100, and is provided with a display panel inside so that it faces the eyes when worn. In this example, the display panel of the head-mounted display 100 is not transparent. In other words, the head-mounted display 100 is a light-opaque head-mounted display.
[0015] The housing 108 may further include an eyepiece lens located between the display panel and the user's eyes when the head mounted display 100 is worn, to expand the user's field of view. The head mounted display 100 may further include speakers or earphones at positions corresponding to the user's ears when worn. The head mounted display 100 also includes a built-in motion sensor that detects the translational and rotational movements of the head of the user wearing the head mounted display 100, as well as the position and posture at each time.
[0016] The head-mounted display 100 also includes a stereo camera 110 on the front surface of the housing 108. The stereo camera 110 captures video of the surrounding real space in a field of view that corresponds to the user's line of sight. If the captured image is displayed immediately, the real space in the direction the user is facing can be seen as it is, which is known as video see-through. Furthermore, if a virtual object is drawn on the image of a real object captured in the captured image, augmented reality (AR) can be realized.
[0017] 2 shows an example of the configuration of an image display system according to this embodiment. The image display system 10 includes a head-mounted display 100, an image generation device 200, and a controller 140. The head-mounted display 100 is connected to the image generation device 200 via wireless communication. The image generation device 200 may further be connected to a server via a network. In this case, the server may provide the image generation device 200 with data for an online application such as a game in which multiple users can participate via the network.
[0018] The image generation device 200 is an information processing device that identifies the position of the viewpoint and the direction of the line of sight of the user wearing the head-mounted display 100 based on the position and orientation of the head-mounted display 100, generates a display image so as to provide a corresponding field of view, and outputs the image to the head-mounted display 100. For example, the image generation device 200 may generate a display image of a virtual world in which the game is set as the electronic game progresses, or may display moving images for viewing or providing information, regardless of whether the world is a virtual world or the real world. Furthermore, by displaying a panoramic image with a wide angle of view centered on the user's viewpoint on the head-mounted display 100, the user can be given a deep sense of immersion in the displayed world. The image generation device 200 may be a stationary game console or a PC.
[0019] The controller 140 is a controller (e.g., a game controller) that is held by the user's hand and into which the user's operations are input to control image generation in the image generation device 200 and image display in the head mounted display 100. The controller 140 is connected to the image generation device 200 by wireless communication. As a variation, one or both of the head mounted display 100 and the controller 140 may be connected to the image generation device 200 by wired communication via a signal cable or the like.
[0020] FIG. 3 is a diagram illustrating an example of an image world that the image generation device 200 displays on the head-mounted display 100. In this example, a state is created in which the user 12 is in a room, which is a virtual space. As shown in the figure, objects such as walls, floors, windows, a table, and objects on the table are arranged in a world coordinate system that defines the virtual space. The image generation device 200 defines a view screen 14 in the world coordinate system according to the position of the viewpoint and the direction of the line of sight of the user 12, and draws a display image by representing the images of objects on the view screen.
[0021] The image generation device 200 acquires the state of the head-mounted display 100 at a predetermined rate and changes the position and orientation of the view screen 14 accordingly. This allows the head-mounted display 100 to display an image in a field of view corresponding to the user's line of sight. The image generation device 200 can also generate stereo images with parallax and display the stereo images in the left and right regions of the display panel of the head-mounted display 100, allowing the user 12 to view a virtual space in three dimensions. This allows the user 12 to experience virtual reality as if they were in a room in the displayed world.
[0022] In this embodiment, the image generation device 200 sequentially acquires the state of the head-mounted display 100, and in turn the position and posture of the user's head wearing it, based on at least images captured by the stereo camera 110. The acquired information can be used to set the view screen 14, as well as to set a play area that defines the range of the real world in which the user can move. The play area is the range of the real world in which the user viewing the virtual world through the head-mounted display 100 can move around, and is, for example, a range in which safe movement without colliding with surrounding objects is guaranteed.
[0023] Visual SLAM (Simultaneous Localization and Mapping) is a known technology that simultaneously estimates the self-position of a mobile object equipped with a camera and creates an environmental map using captured images. Figure 4 is a diagram outlining the principles of Visual SLAM. A camera 22 is mounted on the mobile object and captures video of a real space 26 within its field of view while changing its position and orientation. Assume that feature points 28a and 28b representing point 24 on the same object are extracted from frame 20a captured at a certain time and frame 20b captured a time Δt later.
[0024] The deviation in position coordinates of corresponding feature points 28a, 28b (hereinafter sometimes referred to as "corresponding points") on each frame plane depends on the change in the position and orientation of the camera 22 over time Δt. Specifically, if matrices representing the changes due to the rotational and translational motion of the camera 22 are R and T, respectively, and the three-dimensional vectors from the camera 22 to the point 24 at each time are P1 and P2, the following relational expression holds: P1=R P2+T
[0025] Using this relationship, by extracting multiple corresponding points in two frames that are successive in time and solving simultaneous equations, it is possible to identify changes in the position and orientation of the camera 22 between them. Furthermore, by minimizing the error in the derived results through recursive calculation, it is possible to accurately construct three-dimensional information about the surface of the subject in real space 26, such as point 24. If the camera 22 is a stereo camera 110, the three-dimensional position coordinates of point 24 and the like can be calculated independently at each time, making it easier to perform calculations such as extracting corresponding points.
[0026] However, even when the camera 22 is a monocular camera, the Visual SLAM algorithm has been established. Therefore, for the purpose of tracking the state of the head-mounted display 100, the camera provided in the head-mounted display 100 is not limited to the stereo camera 110. Also, many Visual SLAM algorithms have been proposed, and any of them may be adopted. In any case, according to the illustrated principle, it is possible to derive the change in state of the camera 22 from the previous time at the same rate as the frame rate of the video image, but this information is relative.
[0027] Therefore, as the same process is repeatedly performed and the derived state changes are added to the state obtained at the previous time, errors accumulate, and it is conceivable that the tracking accuracy will deteriorate over time and that tracking will eventually fail. Therefore, feature points from past frames in which the state of the camera 22 has already been identified are associated with the state of the camera 22 and saved. These feature points are then used to compare with the feature points of the current frame (latest frame), thereby eliminating the accumulation of errors at a predetermined timing. In other words, by using past state information as a reference and regularly setting a timing for acquiring the state of the camera 22 as a change from that reference, errors that occur during that time are canceled out.
[0028] This makes it possible to suppress an increase in error even if tracking continues for a long period of time. A frame used as a reference for state changes in this way is called a "key frame." When a key frame is used as a reference for comparison to acquire the state of camera 22, the frame at that time (current frame) itself can also function as a key frame. In order to obtain the required number of corresponding points with the key frame by comparison with the current frame, it is desirable to secure a key frame obtained in a state as close as possible to the state of camera 22 capturing the current frame.
[0029] In other words, to accurately track the state of the camera 22, the wider the range of movement allowed for the mobile object equipped with the camera 22, the more key frames must be saved. On the other hand, if key frames are saved frequently in accordance with the movement of the mobile object during operation, the storage area will eventually be depleted. For this reason, it is possible to discard data starting with the oldest key frames, but in this case, a bias may occur in the state of the camera 22 for which key frames are saved. For example, if the position or orientation of the camera 22 does not change significantly for a long period of time, only key frames representing that position or orientation will remain.
[0030] In this case, if the camera 22 suddenly changes direction or moves in a direction different from its previous tendency, there may not be a key frame with sufficient corresponding points, and matching may fail. As an initial process, it is conceivable to acquire a variety of key frames in advance so as to completely cover the range of movement, and then in subsequent processes, track the state of the camera 22 without changing the key frames. However, in the case of the head-mounted display 100 in particular, the user needs to move until a key frame that satisfies the conditions is obtained, which is a lot of work.
[0031] In this embodiment, key frames are basically acquired along with state information when setting the play area or executing an application. The acquired key frames are then classified into multiple groups with different discard rules. The simplest way is to separate key frames into those that cannot be discarded and those that can be discarded. For example, in the former case, one key frame is selected for each of multiple states that cover all the directions that the camera 22 can face.
[0032] This allows a key frame that is close to the camera 22's current state to be selected as the reference frame, regardless of how the camera's orientation changes. As long as a minimum number of such key frames are secured, the state of the camera 22 can be accurately acquired even if the remaining key frames are biased toward key frames obtained in a state close to the current state. In other words, such key frames can be discarded in order of age. Note that once a classification is determined, it may be fixed as is, or may be updated depending on the situation. Specific examples will be described later.
[0033] 5 shows the internal circuit configuration of the image generation device 200. The image generation device 200 includes a CPU (Central Processing Unit) 222, a GPU (Graphics Processing Unit) 224, and a main memory 226. These components are connected to one another via a bus 230. An input / output interface 228 is further connected to the bus 230. A communication unit 232, a storage unit 234, an output unit 236, an input unit 238, and a recording medium drive unit 240 are connected to the input / output interface 228.
[0034] The communication unit 232 includes a peripheral device interface such as USB or IEEE1394, and a network interface such as a wired LAN or wireless LAN. The storage unit 234 includes a hard disk drive, a nonvolatile memory, etc. The output unit 236 outputs data to the head mounted display 100. The input unit 238 accepts data input from the head mounted display 100 and also accepts data input from the controller 140. The recording medium drive unit 240 drives a removable recording medium such as a magnetic disk, an optical disk, or a semiconductor memory.
[0035] The CPU 222 controls the entire image generating device 200 by executing an operating system stored in the storage unit 234. The CPU 222 also executes various programs (e.g., VR game applications) that are read from the storage unit 234 or a removable recording medium and loaded into the main memory 226, or that are downloaded via the communication unit 232. The GPU 224 has the functions of a geometry engine and a rendering processor, performs drawing processing in accordance with drawing commands from the CPU 222, and outputs the drawing results to the output unit 236. The main memory 226 is composed of RAM (Random Access Memory), and stores programs and data necessary for processing.
[0036] 6 shows the internal circuit configuration of the head mounted display 100. The head mounted display 100 includes a CPU 120, a main memory 122, a display unit 124, and an audio output unit 126. These units are connected to one another via a bus 128. An input / output interface 130 is further connected to the bus 128. To the input / output interface 130, a communication unit 132 including a wireless communication interface, a motion sensor 134, and the stereo camera 110 are connected.
[0037] The CPU 120 processes information acquired from each unit of the head mounted display 100 via the bus 128, and supplies display image and audio data acquired from the image generating device 200 to the display unit 124 and audio output unit 126. The main memory 122 stores programs and data necessary for processing by the CPU 120.
[0038] The display unit 124 includes a display panel such as a liquid crystal panel or an organic EL panel, and displays an image in front of the eyes of a user wearing the head mounted display 100. The display unit 124 may achieve stereoscopic vision by displaying a pair of stereo images in areas corresponding to the left and right eyes. The display unit 124 may further include a pair of lenses that are positioned between the display panel and the user's eyes when the head mounted display 100 is worn, and that expand the user's field of view.
[0039] The audio output unit 126 is composed of speakers or earphones provided at positions corresponding to the user's ears when the head mounted display 100 is worn, and allows the user to hear audio. The communication unit 132 is an interface for sending and receiving data to and from the image generation device 200, and realizes communication using known wireless communication technology such as Bluetooth (registered trademark). The motion sensor 134 includes a gyro sensor and an acceleration sensor, and acquires the angular velocity and acceleration of the head mounted display 100.
[0040] As shown in Fig. 1, the stereo camera 110 is a pair of video cameras that capture the surrounding real space from left and right viewpoints in a field of view corresponding to the user's viewpoint. Objects present in the user's line of sight (typically in front of the user) are captured in frames of a moving image captured by the stereo camera 110. Measurement values by the motion sensor 134 and data on images captured by the stereo camera 110 are transmitted to the image generating device 200 via the communication unit 132 as necessary.
[0041] Fig. 7 is a block diagram showing functional blocks of the image generation device. As described above, the image generation device 200 executes general information processing such as the progress of an electronic game and communication with a server, but Fig. 7 mainly shows a function of acquiring state information of the head-mounted display 100 and a function realized using the state information. Note that at least some of the functions of the image generation device 200 shown in Fig. 7 may be implemented in a server connected to the image generation device 200 via a network. Furthermore, among the functions of the image generation device 200, the function of acquiring state information of the head-mounted display 100 from a captured image may be realized as a separate information processing device, or may be implemented in the head-mounted display 100 itself.
[0042] 7 can be realized in hardware by the configuration of the CPU 222, GPU 224, main memory 226, storage unit 234, etc. shown in Fig. 5, and can be realized in software by a computer program that implements the functions of the multiple functional blocks. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various forms by hardware alone, software alone, or a combination thereof, and are not limited to any one of them.
[0043] The image generating device 200 includes a data processing unit 250 and a data storage unit 252. The data processing unit 250 executes various types of data processing. The data processing unit 250 transmits and receives data to and from the head mounted display 100 and the controller 140 via the communication unit 232, output unit 236, and input unit 238 shown in FIG. 5. The data storage unit 252 is realized by the storage unit 234 shown in FIG. 5, and stores data referenced or updated by the data processing unit 250. In other words, the data storage unit 252 functions as a non-volatile storage device.
[0044] The data storage unit 252 includes an app storage unit 254, a play area storage unit 256, and a registration information storage unit 258. The app storage unit 254 stores data such as programs and object models required to execute applications that involve image display, such as VR games. The play area storage unit 256 stores data related to the play area. The data related to the play area includes data indicating the positions of the points that make up the boundary of the play area (for example, the coordinate values of each point in the world coordinate system).
[0045] The registration information storage unit 258 stores registration data for acquiring the head mounted display 100, and furthermore, the position and posture of the head of a user wearing the head mounted display 100. Specifically, the registration information storage unit 258 stores the above-mentioned key frame data in association with data of an environmental map (hereinafter referred to as a "map") that represents the structure of an object surface in three-dimensional real space.
[0046] The map data is, for example, information on the three-dimensional position coordinates of points on the surface of objects that make up the room where the user plays the VR game, and each point is associated with a feature point extracted from a key frame. The key frame data is associated with the state of the stereo camera 110 when it was acquired.
[0047] The data processing unit 250 includes a system unit 260, an app execution unit 290, and a display control unit 292. The functions of these functional blocks may be implemented in a computer program. The CPU 222 and GPU 224 of the image generating device 200 may perform the functions of the functional blocks by reading the computer program from the storage unit 234 or a recording medium into the main memory 226 and executing it.
[0048] The system unit 260 executes system processing related to the head mounted display 100. The system unit 260 provides a common service for multiple applications (e.g., VR games) for the head mounted display 100. The system unit 260 includes a captured image acquisition unit 262, an image analysis unit 272, and a play area control unit 264.
[0049] The captured image acquisition unit 262 sequentially acquires frame data of images captured by the stereo camera 110, which is transmitted from the head mounted display 100. The frequency at which the captured image acquisition unit 262 acquires data may be the same as the frame rate of the stereo camera 110, or may be lower than that.
[0050] The image analysis unit 272 sequentially acquires state information of the head mounted display 100 using the above-described Visual SLAM technique. The state information of the head mounted display 100 is used to set or reset the play area prior to the execution of an application such as a VR game. The state information is also used to set the view screen when the application is executed, to warn the user when approaching the play area boundary, and so on. Therefore, the image analysis unit 272 appropriately supplies the acquired state information to the play area control unit 264, the app execution unit 290, and the display control unit 292 according to the situation at that time.
[0051] In detail, the image analysis unit 272 includes a registration information generation unit 274 and a state information acquisition unit 276. The registration information generation unit 274 generates registration information consisting of key frame data and map data based on the frame data of the captured image acquired by the captured image acquisition unit 262. The registration information generation unit 274 ultimately stores the generated registration information in the registration information storage unit 258 so that it can be read out at any later time.
[0052] The state information acquisition unit 276 acquires the state of the head mounted display 100, i.e., information on the position and posture, at each time, based on the data of each frame of the captured image acquired by the captured image acquisition unit 262 and the registration information. The registration information may be information being generated by the registration information generation unit 274, or may be information read by the registration information generation unit 274 from the registration information storage unit 258. The state information acquisition unit 276 may integrate the registration information with a measurement value by the motion sensor 134 built into the head mounted display 100 to generate state information.
[0053] The registration information generation unit 274 may extract frames that satisfy a predetermined criterion from among the frames used by the state information acquisition unit 276 to acquire the state of the head mounted display 100, and store the frames as new key frames. At this time, the registration information generation unit 274 classifies the key frames into one of a plurality of groups for which different discard rules are set, based on the predetermined criterion.
[0054] Furthermore, the registration information generation unit 274 discards the data of the selected and saved key frames according to the discarding rule set for each group as necessary. For example, when the data size of the key frames reaches the upper limit and new key frames cannot be saved, the registration information generation unit 274 discards the key frames in order, starting with the oldest registered key frame, in a group for which the discarding rule set is FIFO (First-In First-Out).
[0055] Alternatively, once a predetermined number of key frames belonging to a group that cannot be discarded have been saved, the registration information generation unit 274 may finalize the key frames that make up that group and classify all subsequent key frames into a group that can be discarded. In either case, by differentiating the key frame discard rules, the registration information storage unit 258 can be controlled to prevent a capacity shortage while securing the key frames necessary to acquire state information.
[0056] The play area control unit 264 sets an area in real space where the user can move safely as the play area, and then presents an appropriate warning when the user approaches the boundary of the play area during the execution of the application. At either stage, the play area control unit 264 successively acquires status information of the head mounted display 100 from the image analysis unit 272 and performs processing based on the status information.
[0057] For example, when setting a play area, the play area control unit 264 instructs the user wearing the head-mounted display 100 to look around. As a result, real objects around the user, such as furniture, walls, and floors, are photographed by the stereo camera 110. The image analysis unit 272 sequentially acquires frames of the photographed images and generates registration information based on them. The play area control unit 264 refers to the map data in the registration information and automatically determines, as the play area, an area of the floor surface that will not collide with furniture, walls, etc.
[0058] The play area control unit 264 may accept a user's operation to edit the play area by displaying an image representing the boundaries of the play area once determined on the head-mounted display 100. In this case, the play area control unit 264 acquires the content of the user's operation via the controller 140 and changes the shape of the play area accordingly. The play area control unit 264 finally stores data regarding the determined play area in the play area storage unit 256.
[0059] The registration information generation unit 274 of the image analysis unit 272 may set key frames saved during play area setting or key frames selected from the saved key frames according to a predetermined criterion as non-discardable key frames. As described above, when setting the play area, images are intentionally taken to cover the entire area around the user, making it easier to obtain key frames that are not biased in state.
[0060] The app execution unit 290 reads data of an application such as a VR game selected by the user from the app storage unit 254 and executes it. At this time, the app execution unit 290 successively acquires state information of the head mounted display 100 from the image analysis unit 272, sets the view screen at the corresponding position and orientation, and draws the display image. This makes it possible to present the world of the display target in a field of view that corresponds to the movement of the user's head.
[0061] The display control unit 292 sequentially transmits various images generated by the app execution unit 290, such as frame data of VR images and AR images, to the head mounted display 100. When setting the play area, the display control unit 292 also transmits to the head mounted display 100, as necessary, an image instructing the user to look around, an image showing the state of the tentatively determined play area and accepting editing, an image warning the user of approaching the boundary of the play area, and the like.
[0062] FIG. 8 shows an example of frame images captured by stereo camera 110 and keyframe data obtained from them. In reality, a pair of images captured by stereo camera 110 is obtained for each frame, but the figure shows only one of them schematically. State information acquisition unit 276 extracts multiple feature points contained in frame image 40 using known methods such as corner detection or a method based on brightness gradients. As described above, this is compared with a frame image from the previous time to obtain corresponding points, thereby identifying state information of stereo camera 110 at the time frame image 40 was captured.
[0063] On the other hand, when a predetermined number or more, for example, 24 or more, of feature points are extracted from one frame image 40, the registration information generation unit 274 designates it as a key frame. This is because frames with many feature points are easier to detect corresponding points by matching with other frames. However, the selection criteria for a key frame are not limited to the number of feature points, and may include a frame image in which the state of the stereo camera 110 at the time of capture differs by more than a threshold, or a combination of multiple criteria may be used.
[0064] The registration information generation unit 274 generates data for a key frame 42 that is composed of the position coordinates on the image plane of a feature point (e.g., feature point 44) extracted from the selected frame image 40 and an image of a predetermined range around the feature point. The registration information generation unit 274 further associates the data with state information of the stereo camera 110 at the time the image was captured, and generates the final key frame data. However, the components of the key frame data are not limited to these.
[0065] 9 shows an example of the data structure of the registration information stored in the registration information storage unit 258. The registration information has a structure in which multiple key frame data 72a, 72b, 72c, 72d, etc. are associated with one map data 70. In the figure, the map data 70 is represented as "MAP1," and the key frame data 72a, 72b, 72c, 72d, etc. are represented as "KF1," "KF2," "KF3," "KF4," etc., and each has a structure as shown in the example below.
[0066] The map data 70 consists of a "map ID," which is identification information uniquely assigned to a map, and a list of "point information" that constitutes the map. The point information is data that associates a "point ID," which is identification information uniquely assigned to each point, with its position coordinates in three-dimensional space (represented as "3D position coordinates" in the figure).
[0067] Each piece of key frame data 72a, 72b, 72c, 72d, etc. is composed of a "key frame ID," which is identification information uniquely assigned to the key frame, "camera state information," which is the position and orientation of the stereo camera 110 when it was acquired, its "acquisition time," and a list of "feature point information" that constitutes the key frame. Here, the camera state information may be expressed as a relative state obtained by Visual SLAM, for example, with the initial state as the reference, or an absolute value obtained by combining the measurement value of the motion sensor 134 built into the head mounted display 100 and the result of Visual SLAM may be used.
[0068] The feature point information corresponds to the data shown in key frame 42 in Fig. 8, and is data that associates the position coordinates of the extracted feature points on the image plane (indicated as "2D position coordinates" in the figure) with the surrounding image of a predetermined size. Each feature point is also associated with a "point ID," which is information identifying the point in three-dimensional space that it represents.
[0069] The data structure of the registration information is not limited to that shown in the figure, and may include auxiliary information that can efficiently extract corresponding points when matching with the current frame or improve the accuracy of the map, as appropriate. The registration information storage unit 258 may store registration information with a similar structure for multiple map data 70. For example, if there are multiple locations where a user plays a game, the play area control unit 264 sets a play area for each location, and the registration information generation unit 274 acquires and stores the registration information for each location.
[0070] In this case, the registration information is further associated with identification information of the location. The image analysis unit 272 acquires location information input by the user at the start of game play, reads out the corresponding registration information, and uses it to acquire state information of the head mounted display 100. In addition, when the furniture arrangement is changed even in the same location or when the lighting environment differs depending on the time of day, the registration information generation unit 274 may generate registration information for each state.
[0071] In either case, the registration information generation unit 274 stores the key frame data 72a, 72b, 72c, 72d, etc. in each registration information in a state where the data is classified according to a predetermined criterion, such as "Group A," "Group B," etc. The registration information generation unit 274 stores in its internal memory a key frame discard rule set for each group. For example, two groups may be created, one of which cannot be discarded and the other of which is discarded in a FIFO manner.
[0072] In this case, the registration information generation unit 274 first classifies all key frames obtained in multiple states covering all possible directions the user can face into the former category, and then classifies key frames obtained in overlapping states into the latter category. Alternatively, the "quality" of key frames may be evaluated and divided into multiple levels to create groups. For example, the "quality" may be divided into three levels, and key frames belonging to the lowest quality group may be discarded preferentially. The highest quality group cannot be discarded, and groups of intermediate quality may be discarded when certain conditions are met.
[0073] The "quality" of a newly obtained keyframe may be compared with that of existing keyframes, and the group to which the existing keyframe belongs may be changed as appropriate. Here, a keyframe with high "quality" is one that was captured in a state where the quality differs from the others by a predetermined value or more, one with a large number of feature points, a keyframe that was captured recently, or one that has been used for matching many times. Any of these may be used as an index for classification, or two or more of them may be combined to form the basis for classification.
[0074] For example, keyframes may be scored based on these criteria and classified according to the range of their total scores. This allows keyframe data to be discarded in descending order of quality. If the number of times used for matching is used as a classification index, the registration information generation unit 274 also records the number of times matching is used in the keyframe data structure shown in the figure. The number of groups into which keyframes are classified, the classification criteria, and the discarding rules for the groups can be variously considered, and may be optimized depending on the characteristics of the application, the scale of the system, etc.
[0075] 10 shows an example of space division for evaluating the spatial coverage of key frames. Here, the spatial coverage of key frames refers to the extent of the range of the distribution of the states of the head mounted display 100 when each key frame is captured, relative to all states that the head mounted display 100 can take.
[0076] The figure shows three example patterns of division, all of which are examples of equal division in the yaw direction with the direction of gravity as the axis, centered on the user's position in a reference state such as when the play area setting is started. Pattern 50 is composed of divided areas divided into four at a central angle of 90 degrees. Pattern 52 is composed of divided areas that are shifted in phase by 45 degrees from pattern 50, based on the direction the user is facing. Pattern 54 is composed of divided areas divided into 16 at a central angle of 22.5 degrees.
[0077] The registration information generation unit 274 evaluates the range covered by a frame that satisfies the conditions of a key frame by using at least one of the segmented areas of patterns 50, 52, and 54 as a bin. Specifically, the registration information generation unit 274 determines that the segmented area is covered by the key frame by identifying which of the segmented areas the direction of the stereo camera 110 when the key frame was captured belongs to. Then, the registration information generation unit 274 counts the total number of segmented areas covered by the key frame.
[0078] For example, when all three segmentation patterns shown in the figure are used, a threshold value of "10" is set for the total number of segmentation areas covered by key frames. When only pattern 54 is used, a threshold value of "8", for example, is set. These values are determined as values that can be cleared by a user wearing the head mounted display 100 if the user looks around 180 degrees, but cannot be cleared if the user does not look around.
[0079] However, the division pattern and threshold values used may be set in various ways depending on factors such as the range of directions the user can face when the application is running. For example, the central angle does not need to be divided equally. When the number of division areas covered by key frames reaches a threshold, the registration information generation unit 274 determines that the key frames have covered the space. Then, for example, the registration information generation unit 274 classifies the key frame that first covers the division area into a group that cannot be discarded. At the same time, the play area control unit 264 may use the map already constructed at that time to finally determine the shape of the play area.
[0080] By selecting non-discardable keyframes based on this evaluation, the keyframe states are distributed as much as possible, allowing a wide range to be covered with a small number of keyframes. The registration information generation unit 274 then classifies the obtained keyframes into groups that allow discarding. Alternatively, if a newly obtained keyframe for the same segmentation area has higher quality, it will replace the existing non-discardable keyframe.
[0081] The division pattern is not limited to division in the yaw direction, but may also be division in the pitch direction, which determines the angle of elevation, or division in a translation direction such as forward, backward, left, and right of the user, or division in multiple directions by combining these. However, if comprehensiveness is evaluated using division in the yaw direction, which provides feature point information with the greatest number of variations, it becomes easier to avoid situations where there is no frame that can be matched with the current frame, even if the number of key frames that cannot be discarded is small.
[0082] Next, the operation of the image display system configured as described above will be described. Fig. 11 is a flowchart showing the processing procedure for the image generation device 200 to set the play area. The user can select initial setup or resetting of the play area in the system setting menu of the head mounted display 100. When initial setup or resetting of the play area is selected, the captured image acquisition unit 262 of the image generation device 200 establishes communication with the head mounted display 100 and starts acquiring data of frames captured by the stereo camera 110 (S10).
[0083] Next, the play area control unit 264 causes the head mounted display 100 to display a message urging the user to change the state of the stereo camera 110 via the display control unit 292 (S12). In practice, the play area control unit 264 may cause the head mounted display 100 to display an instruction image that prompts the user to look around while wearing the head mounted display 100.
[0084] The image analysis unit 272 then performs Visual SLAM using the frame images captured in this way, and generates registration information for the space in which the user is located (S14). Specifically, the image analysis unit 272 obtains the correspondence between feature points extracted from each frame image and identifies the three-dimensional position coordinates of the points represented by the feature points. It also uses the information on the position coordinates to obtain changes in the state of the stereo camera 110, and then reprojects the points accordingly to correct errors in the position coordinates. During these processes, the image analysis unit 272 selects key frames from the frame images based on predetermined conditions.
[0085] In parallel with this, the play area control unit 264 constructs a play area in the space surrounding the user (S16). For example, the play area control unit 264 estimates the three-dimensional shape of the user's room using map data included in the registration information. At this time, the play area control unit 264 may detect a plane (typically the floor) perpendicular to the direction of gravity indicated by the motion sensor 134 based on the estimated three-dimensional shape of the room, and detect the result of combining multiple detected planes of the same height as the play area.
[0086] The processes of S14 and S16 are repeated until the image analysis unit 272 determines that enough key frames have been obtained to cover the surrounding space, based on an evaluation of the coverage rate of the key frames for the spatial divisions shown in Fig. 10 (N in S18, S14, S16). This allows the boundaries of the play area to be set without omissions, and sufficient map data and key frame data for the play area to be prepared.
[0087] If it is determined that the key frames have encompassed the surrounding space (Y in S18), the play area control unit 264 stores play area data including the coordinate values of the point cloud that constitutes the boundary of the play area in the play area storage unit 256 (S20). Note that, as described above, the play area control unit 264 may give the user an opportunity to edit the play area, and then store the edited data in the play area storage unit 256.
[0088] Furthermore, the image analysis unit 272 associates map data such as that shown in Fig. 9 with a plurality of key frame data and stores them in the registration information storage unit 258 (S22). At this time, the image analysis unit 272 classifies the key frame that first filled each division area shown in Fig. 10, for example, into a group that cannot be discarded, and classifies the other key frames into a group that can be discarded, and stores them. However, as described above, various modes of classification are conceivable.
[0089] 12 is a bird's-eye view illustrating an example of state changes of the head mounted display 100 during the processing period of S14 to S18 in FIG. 11. Here, the position and direction of the head mounted display 100 are represented by an isosceles triangle (for example, isosceles triangle 32). The base of the isosceles triangle corresponds to the imaging surface of the stereo camera 110. In other words, the line of sight of the head mounted display 100, and therefore the user, is in the direction of, for example, arrow 34.
[0090] Following the instructions displayed in S12, the user wears the head-mounted display 100 on his / her head and moves around the room 30 while looking around. This causes the state of the head-mounted display to change as shown in the figure. At this time, the play area control unit 264 may control the head-mounted display 100 to display the moving image captured by the stereo camera 110 in the line of sight direction as it is. The user can move safely thanks to video see-through, which shows the user the state of the real space in the direction the user is facing.
[0091] The image analysis unit 272 of the image generating device 200 generates registration information based on the frames of the video. The play area control unit 264 uses this information to set the play area 36. The image analysis unit 272 may also cause the head mounted display 100 to transmit, along with the frame data of the video, measurements taken by the motion sensor 134, such as the angular velocity and acceleration of the head mounted display 100. By using this information, the image analysis unit 272 can more accurately identify the state information of the stereo camera 110 at each time.
[0092] 13 is a flowchart showing the processing steps for the image generating device 200 to execute an application. This flowchart begins, for example, when a user launches an application on the image generating device 200. As a prerequisite, it is assumed that at least the play area setting processing shown in FIG. 11 has been performed. First, the play area control unit 264 reads play area data stored in the play area storage unit 256, for example, data indicating the shape and size of the play area (S32).
[0093] The image analysis unit 272 also reads out the registration information stored in the registration information storage unit 258 (S34). The registration information includes key frame data that has been previously acquired by setting the play area or executing an application.
[0094] Next, the App execution unit 290 reads the program data of the application from the App storage unit 254 and starts processing (S36). For example, the App execution unit 290 starts a VR game that progresses according to the user's movements and the contents of the user's operations via the controller 140. In response to this, the image analysis unit 272 executes Visual SLAM using frames of images captured by the stereo camera 110 that are sequentially transmitted from the head mounted display 100, and acquires state information of the head mounted display 100 (S38).
[0095] At this time, the image analysis unit 272 uses the key frame read out in S34 at a predetermined timing, such as a predetermined cycle, as a reference for the current frame. This makes it possible to accurately acquire the state of the head mounted display 100 even in the early stages when sufficient frame data is not yet available. In particular, by using the key frame acquired when setting the play area as a reference, the user's position relative to the play area can be acquired with high precision.
[0096] Furthermore, when a newly obtained frame satisfies the requirements for a key frame, the image analysis unit 272 generates key frame data for that frame and stores it in an internal memory, etc. At this time, if a condition for discarding data is met, such as the total size of all key frame data exceeding the capacity of the corresponding storage area in the registration information storage unit 258, the image analysis unit 272 selects and discards a key frame from a group permitted to be discarded.
[0097] For example, the image analysis unit 272 determines which key frames to discard from the group based on a predetermined index, such as a key frame acquired earlier, a key frame that has been used as a match target only a few times, or a key frame acquired in a state far from the current state of the stereo camera 110. The image analysis unit 272 also determines the classification of new key frames based on a predetermined criterion. When a new key frame is registered, the group to which an existing key frame belongs may be changed.
[0098] Note that in the processing of S38, the image analysis unit 272 may, as described above, integrate the measurement values from the motion sensor 134 built into the head mounted display 100 and the results of Visual SLAM to derive the state of the head mounted display 100. State information of the head mounted display 100 is sequentially supplied to the App execution unit 290. Based on the state information of the head mounted display 100, the App execution unit 290 generates frame data of a display image in a field of view corresponding to the user's viewpoint and line of sight, and sequentially transmits the frame data to the head mounted display 100 via the display control unit 292 (S40).
[0099] The head mounted display 100 sequentially displays the transmitted frame data. When the user's position in the real world satisfies the conditions for issuing a warning, for example, when the distance from the head mounted display 100 worn by the user to the boundary of the play area becomes equal to or less than a predetermined threshold (for example, 30 cm) (Y in S42), the play area control unit 264 of the image generating device 200 executes a predetermined warning process for the user (S44).
[0100] For example, the play area control unit 264 supplies an image that represents the boundary of the play area as a three-dimensional object such as a fence to the display control unit 292. The display control unit 292 may superimpose the image that represents the boundary of the play area on the image of the application generated by the app execution unit 290 and display it as a display image on the head-mounted display 100. Furthermore, the play area control unit 264 may cause the head-mounted display 100 to display a video see-through image via the display control unit 292 when the user's position in the real world approaches the boundary of the play area or when the user passes over the boundary of the play area.
[0101] If the user's position in the real world does not satisfy the conditions for issuing a warning (N at S42), the process of S44 is skipped. Unless a predetermined termination condition is satisfied, such as the user stopping execution of an application (N at S46), the processes of S38 to S44 are repeated. If the predetermined termination condition is satisfied (Y at S46), the image analysis unit 272 reflects the new key frame data obtained at S38 and the resulting update content of the entire registered information in the registered information storage unit 258 (S48), and terminates all processing.
[0102] 14 is a diagram illustrating an example of key frame transitions in this embodiment. (a) shows key frames acquired and used when setting a play area, and (b) shows key frames acquired and used when an application is executed, using isosceles triangles that represent the state of the head mounted display 100 when each is acquired. As in FIG. 12, the base of the isosceles triangle corresponds to the imaging surface of the stereo camera 110. In addition, the illustrated example assumes that key frames are classified into two types: "non-discardable" and "discardable," with the former being represented by a black isosceles triangle (e.g., isosceles triangle 80) and the latter by a white isosceles triangle (e.g., isosceles triangle 82).
[0103] The transition (a) is realized when the play area is not set yet or when it needs to be reset, along with the setting of the play area. In the initial stage S50, the user looks around the surrounding space as indicated by the arrows in response to an instruction from the play area control unit 264. This causes the stereo camera 110 of the head-mounted display 100 to capture video images while moving its line of sight in a roughly radial direction. The image analysis unit 272 generates registration information using Visual SLAM using frames from the video images.
[0104] At this time, the registration information generation unit 274 sequentially determines frames that satisfy the requirements from among the frames from which feature points have been extracted as key frames. In the figure, the stage of S50 is the initial state in which all key frames are discardable, but in the subsequent stage of S52, the key frames are classified. For example, the registration information generation unit 274 selects key frames acquired in states corresponding to each of the yaw direction division areas shown in Figure 10 one by one and classifies them into a group that cannot be discarded, and classifies the rest into a group that can be discarded.
[0105] Once the play area has been set, the registration information generation unit 274 associates all key frame data, regardless of classification, with the map data along with classification information and stores them in the registration information storage unit 258. In practice, for example, a total of about 100 key frame data are stored, with about half of them classified as non-discardable. However, this is not intended to limit the scope of the embodiment.
[0106] The transition in (b) is realized together with the processing of the application. In the initial stage S60, the registration information generation unit 274 reads out the registration information stored in the registration information storage unit 258. As a result, key frames covering all directions are obtained as surrounded by the dashed line 84, without the user having to look around. Furthermore, when the application is started, new captured images are transmitted from the head mounted display 100, and key frames are added as appropriate.
[0107] The state information acquisition unit 276 acquires the current state information of the head mounted display 100 by comparing it with a key frame obtained in a state similar to the newly transmitted frame (for example, frame 86). Even in the initial stage when the number of newly transmitted frames is small, a variety of key frames are prepared, so that state information can be acquired efficiently and with stable accuracy regardless of the state of the head mounted display 100.
[0108] The registration information generation unit 274 basically determines the discard rules by inheriting the classification of key frames that was previously performed, such as when setting the play area. The registration information generation unit 274 also classifies new key frames that have been acquired up to that point in S62, when the initial state of the head mounted display 100 is acquired. At this time, the registration information generation unit 274 may increase the number of non-discardable key frames. The registration information generation unit 274 may also change the group configuration, such as by making a key frame that was originally classified into a non-discardable group discardable.
[0109] When the user's position or posture changes significantly, for example, as the application game progresses, the number of key frames in a state similar to that change increases, as shown in S64. When the number of key frames reaches an upper limit set in consideration of the capacity of the registration information storage unit 258, the registration information generation unit 274 sequentially discards key frames that have been determined to be discardable up to that point. This allows key frames that are likely to be matched to be stored preferentially, even if the storage capacity of the registration information storage unit 258 is limited.
[0110] Furthermore, by retaining key frames selected based on predetermined criteria as non-discardable, even if the user's movement patterns suddenly change, matching will not take time or fail, and highly accurate state tracking can be maintained. When application processing is completed, the registration information generation unit 274 associates all key frame data, regardless of classification, with map data along with classification information and stores them in the registration information storage unit 258. As a result, the same transition as in (b) will be repeated the next time the application is processed.
[0111] According to the embodiment described above, in a technology for acquiring the position and orientation of a moving object and an environmental map from captured image frames using Visual SLAM, a reference keyframe is selected and used to compare with the current frame, thereby reducing errors in state information. Here, keyframes are classified into multiple groups with different discard rules based on predetermined criteria. This makes it possible to control whether to discard, the time until discarding, and the priority of storage, depending on the compatibility of the quality and characteristics of the keyframe with the actual state of the moving object.
[0112] As a result, even with limited storage capacity, it is possible to store keyframes that are highly necessary in a readable state, enabling accurate acquisition of status information from the early stages when capture time is short. When applied to a head-mounted display that displays images of a virtual world according to the user's movements, such as in a VR game, keyframes are acquired when the play area is set. This allows keyframes that cover the entire space to be stored according to the user's individual environment, enabling highly accurate acquisition of status information and, ultimately, the generation of display images and warnings of dangers associated with movement.
[0113] While tracking the state, new key frame data is generated and registered, and existing key frames are discarded as appropriate according to rules, so the key frame configuration can be kept constantly optimized. This eliminates the need to prepare many key frames in advance to anticipate a variety of situations, saving the user time and effort in initial setup and reducing the required storage capacity.
[0114] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention. [Explanation of symbols]
[0115] 10 Image display system, 100 Head-mounted display, 200 Image generation device, 222 CPU, 226 Main memory, 258 Registration information storage unit, 264 Play area control unit, 272 Image analysis unit, 274 Registration information generation unit, 276 Status information acquisition unit, 292 Display control unit.
Claims
1. a state information acquisition unit that extracts feature points from the latest frame of a moving image that shows the space around the user, which is being captured by a camera provided in the head-mounted display, and acquires state information about the position and attitude of the head-mounted display based on the relationship between the corresponding feature points in a past frame that is a comparison target and the points on the subject that they represent; a registration information generation unit that defines frames that satisfy predetermined conditions among the latest frames as key frames to be used as references for the past frames, classifies the frames into one of a plurality of groups having different discard rules, and registers the frames, and discards any of the registered key frame data in accordance with the discard rule; a registration information storage unit that stores the key frame data together with the classification information; a play area control unit that sets a play area in which the user can move based on a three-dimensional environmental map made up of the set of points that is constructed together with the state information; Equipped with The information processing device is characterized in that the registration information generation unit selects the key frames to be classified into a group that cannot be discarded from the frames of the video captured when the play area is set.
2. The information processing device according to claim 1, characterized in that the registration information generation unit selects the key frames to be classified into a group that cannot be discarded based on the distribution of the state of the head-mounted display when each of the key frames was captured.
3. The information processing device described in claim 2, characterized in that the registration information generation unit classifies one of the key frames taken in each direction of a partitioned area formed by dividing space in a yaw direction centered on the head-mounted display and with the direction of gravity as its axis, into the group that cannot be discarded.
4. The information processing device according to any one of claims 1 to 3, characterized in that the registration information generation unit classifies the key frames based on at least one of the number of extracted feature points, the acquisition time, the number of times used for matching, and the difference from others in the state of the head-mounted display when photographed.
5. The information processing device according to any one of claims 1 to 4, characterized in that when a new key frame causes the number of registered key frames to exceed the upper limit, the registration information generation unit discards the key frame that was registered earliest among the registered key frames belonging to one of the groups.
6. an application execution unit that generates a display image in a field of view corresponding to a line of sight of a user based on state information of the head-mounted display; 6. The information processing device according to claim 1, wherein the state information acquisition unit acquires state information of the head-mounted display using the key frame data stored in the registration information storage unit before the application execution unit starts processing.
7. extracting feature points from the latest frame of a video image of the user's surroundings that is being captured by a camera provided in the head-mounted display, and acquiring state information on the position and orientation of the head-mounted display based on the relationship between the corresponding feature points in a past frame to be compared and the points on the subject that they represent; a step of defining a frame that satisfies a predetermined condition among the latest frames as a key frame to be used as a reference for the past frames, classifying the frame into one of a plurality of groups for which different discarding rules are set, and registering the key frame; discarding any of the registered key frame data in accordance with the discarding rule; storing the key frame data together with the classification information in a storage device; A step of setting a play area in which the user can move based on a three-dimensional environmental map made up of the set of points constructed together with the state information; Including, An information processing method characterized in that the registration step selects the key frames to be classified into a group that cannot be discarded from the frames of the video captured when the play area is set.
8. a function of extracting feature points from the latest frame of a video image of the user's surroundings that is being captured by a camera provided in the head-mounted display, and acquiring state information on the position and posture of the head-mounted display based on the relationship between the corresponding feature points in a past frame to be compared and the points on the subject that they represent; a function of defining a frame that satisfies a predetermined condition among the latest frames as a key frame to be used as a reference for the past frames, and classifying and registering the frame into one of a plurality of groups for which different discarding rules are set; a function of discarding any of the registered key frame data in accordance with the discarding rules; a function of storing the key frame data together with the classification information in a storage device; A function of setting a play area in which a user can move based on a three-dimensional environmental map made up of a set of points constructed together with the state information; This is realized by a computer, A computer program characterized in that the registration function selects the key frames to be classified into a group that cannot be discarded from the frames of the video captured when the play area is set.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2019149621A
Self-position estimation device, self-position estimation method and program
JP2020144710A
Information processing device, information processing method and program
JP2021184115A
Image generation device and image generation method
WO2019123548A1