Information processing device, re-training method, and re-training program
The information processing device updates a virtual information generation model using real-space images and positional data to ensure that three-dimensional virtual information reflects changes in the real environment, addressing the mismatch between virtual and real spaces.
Patent Information
- Application Number
- PCT/JP2024/013031
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies fail to reflect changes in the real space within three-dimensional virtual information generated by devices like head-mounted displays, leading to a mismatch between the virtual and real environments.
An information processing device that includes an acquisition unit for capturing images and positional information, and a re-learning unit to update a virtual information generation learning model using images and positional data to generate three-dimensional virtual information that reflects changes in the real space.
Enables the generation of three-dimensional virtual information that accurately represents the modified situation in the real space, allowing users to experience a synchronized virtual environment.
Smart Images

Figure JP2024013031_02102025_PF_FP_ABST
Abstract
Description
Information processing device, relearning method, and relearning program
[0001] The present disclosure relates to an information processing device, a relearning method, and a relearning program.
[0002] Devices such as head-mounted displays (HMDs) are known. For example, a user can wear an HMD and experience a virtual space through the HMD. A technology for generating three-dimensional virtual information, which is information about the virtual space, is known (see Non-Patent Document 1).
[0003] Ben Mildenhall et al. "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis"
[0004] For example, three-dimensional virtual information is generated using Neural Radiance Fields (NeRF) or the like. When the three-dimensional virtual information is generated, an image of the real space at a certain time is used. If the situation in the real space changes after that time, the change in the situation in the real space is not reflected in the three-dimensional virtual information. Therefore, the user cannot experience the virtual space in the changed situation.
[0005] The objective of the present disclosure is to enable three-dimensional virtual information to be generated that shows the modified situation.
[0006] According to one aspect of the present disclosure, there is provided an information processing device, the information processing device including: an acquisition unit that acquires learning data, which is at least one of an image including an object moved in real space and one or more multi-view images that are one or more images of the object when the object is viewed from a direction different from a direction in which the object was captured, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information of the virtual space; and a re-learning unit that re-learns the virtual information generation learning model using the learning data and the positional information.
[0007] According to the present disclosure, three-dimensional virtual information showing the changed situation can be generated.
[0008] 1 is a diagram showing an information processing system according to a first embodiment; 2 is a diagram showing hardware included in an information processing device according to the first embodiment; 3 is a block diagram showing functions of the information processing device according to the first embodiment; 4 is a flowchart showing an example of processing executed by the information processing device according to the first embodiment; 5 is a block diagram showing functions of an information processing device according to a second embodiment; and 6 is a flowchart showing an example of processing executed by the information processing device according to the second embodiment.
[0009] Hereinafter, an embodiment will be described with reference to the drawings.
[0010] Embodiment 1. Fig. 1 is a diagram showing an information processing system according to embodiment 1. The information processing system includes an information processing device 100 and an imaging device 200. The information processing system may also include a display device 300. The information processing device 100, the imaging device 200, and the display device 300 communicate with each other via a network.
[0011] For example, the information processing device 100 is a device that executes a relearning method. For example, the imaging device 200 is a see-through HMD. The imaging device 200 may be a camera that captures an image of an object existing in the real space 10. For example, the display device 300 is a fully immersive HMD. The display device 300 can display the virtual space 20.
[0012] A brief description of the first embodiment will be given. Note that the following specific example is merely an example. A user wearing the imaging device 200 moves the handlebar 11 downward. The imaging device 200 captures an image of the handlebar 11 in the downwardly moved state. The imaging device 200 transmits the captured image to the information processing device 100.
[0013] The information processing device 100 can use the virtual information generation learning model to generate three-dimensional virtual information, which is information about the virtual space 20 corresponding to the real space 10. The virtual information generation learning model can generate three-dimensional virtual information indicating the handle 21 in a state where it is not lowered downward. In other words, the virtual information generation learning model generates three-dimensional virtual information that does not reflect a changed situation in the real space 10. Therefore, the information processing device 100 re-trains the virtual information generation learning model. In detail, the information processing device 100 re-trains the virtual information generation learning model using the image. As a result, the virtual information generation learning model can generate three-dimensional virtual information indicating the changed situation.
[0014] The information processing device 100 receives gaze information from the display device 300. The information processing device 100 generates three-dimensional virtual information using the virtual information generation learning model and the gaze information. The information processing device 100 transmits the three-dimensional virtual information to the display device 300. This allows the user wearing the display device 300 to visually recognize the handle 21 in a downwardly tilted state in the virtual space 20.
[0015] The processing performed in the information processing system will be described in detail below.
[0016] The hardware included in the information processing device 100 will now be described. Fig. 2 is a diagram showing the hardware included in the information processing device of embodiment 1. The information processing device 100 is also referred to as a computer. The information processing device 100 includes a processor 101, a volatile storage device 102, and a non-volatile storage device 103.
[0017] The processor 101 controls the entire information processing device 100. For example, the processor 101 is a central processing unit (CPU) or a field programmable gate array (FPGA). The processor 101 may be a multiprocessor. The information processing device 100 may also include a processing circuit.
[0018] The volatile storage device 102 is a main storage device of the information processing device 100. For example, the volatile storage device 102 is a random access memory (RAM). The nonvolatile storage device 103 is an auxiliary storage device of the information processing device 100. For example, the nonvolatile storage device 103 is a hard disk drive (HDD) or a solid state drive (SSD).
[0019] Next, a description will be given of the functions of the information processing device 100. Fig. 3 is a block diagram showing the functions of the information processing device of embodiment 1. The information processing device 100 has a storage unit 110, an acquisition unit 120, a detection unit 130, a relearning unit 140, a calculation unit 150, a virtual information generation unit 160, and a provision unit 170.
[0020] The storage unit 110 may be realized as a storage area secured in the volatile storage device 102 or the non-volatile storage device 103. Some or all of the acquisition unit 120, detection unit 130, relearning unit 140, calculation unit 150, virtual information generation unit 160, and provision unit 170 may be realized by a processing circuit. Also, some or all of the acquisition unit 120, detection unit 130, relearning unit 140, calculation unit 150, virtual information generation unit 160, and provision unit 170 may be realized as program modules executed by the processor 101. For example, the program executed by the processor 101 is also referred to as a relearning program or a relearning program product. For example, the relearning program is recorded on a recording medium.
[0021] The storage unit 110 stores various information. The acquisition unit 120 acquires video from the imaging device 200. When video is acquired, the detection unit 130 detects the movement of an object based on the video. For example, the detection unit 130 detects the movement of the steering wheel 11 based on the video. As a method for detecting the movement of an object, for example, the detection unit 130 detects the movement of an object by detecting the difference between images.
[0022] The acquisition unit 120 acquires from the detection unit 130 an image including an object that has been moved in the real space 10. For example, the acquisition unit 120 acquires from the detection unit 130 an image including the handle 11 in a downwardly tilted state.
[0023] The acquisition unit 120 acquires position information of an object that has been moved in the real space 10 from the image capture device 200. For example, the acquisition unit 120 acquires position information of the handle 11 that is in a downwardly tilted state from the image capture device 200.
[0024] The acquisition unit 120 acquires the virtual information generation learning model. For example, the acquisition unit 120 acquires the virtual information generation learning model from the storage unit 110. Also, for example, the acquisition unit 120 acquires the virtual information generation learning model from an external device. The external device is a device that exists outside the information processing device 100. For example, the external device is a cloud server, an external memory, etc. An illustration of the external device is omitted. Note that, for example, the virtual information generation learning model is a trained model using NeRF. The virtual information generation learning model may also be called an NeRF learning model.
[0025] The relearning unit 140 converts the position information in the real space 10 into position information in the virtual space 20. For example, the relearning unit 140 performs the conversion using information indicating the correspondence between the real space 10 and the virtual space 20. The acquiring unit 120 acquires the position information in the virtual space 20 from the relearning unit 140.
[0026] The re-learning unit 140 re-learns the virtual information generation learning model using an image including an object in a moved state and the position information of the object in the virtual space 20. For example, the re-learning unit 140 re-learns the virtual information generation learning model using an image including the handle 11 in a downwardly tilted state and the position information of the handle in the virtual space 20.
[0027] The acquisition unit 120 may acquire gaze information, which is the direction in which an image of an object in a moved state is captured, from the imaging device 200. For example, the acquisition unit 120 acquires gaze information, which is information indicating the direction in which a user wearing the imaging device 200 is looking at the handle 11 in a lowered state (i.e., the imaging direction), from the imaging device 200. When the gaze information is acquired, the re-learning unit 140 may re-learn the virtual information generation learning model using an image including the object in a moved state, the position information of the object in the virtual space 20, and the gaze information.
[0028] The above describes a case where the acquisition unit 120 acquires an image including an object that has been moved in the real space 10 from the detection unit 130. The acquisition unit 120 may acquire an image including an object that has been moved in the real space 10 from the imaging device 200.
[0029] In the above description, the acquisition unit 120 acquires position information of an object moved in the real space 10 from the image capture device 200. The acquisition unit 120 may acquire position information of the image capture device 200 in the real space 10 and the line-of-sight information from the image capture device 200. The calculation unit 150 may calculate position information of the object moved in the real space 10 based on the position information of the image capture device 200 and the line-of-sight information. When the position information is calculated, the acquisition unit 120 acquires the position information from the calculation unit 150.
[0030] In the above description, the acquisition unit 120 acquires the position information in the virtual space 20 from the relearning unit 140. The acquisition unit 120 may acquire the position information in the virtual space 20 from the imaging device 200.
[0031] When the acquisition unit 120 acquires the gaze information from the display device 300, the virtual information generation unit 160 generates three-dimensional virtual information based on the virtual information generation learning model and the gaze information. The provision unit 170 provides the three-dimensional virtual information to the display device 300.
[0032] Next, the processing executed by the information processing device 100 will be described using a flowchart. Fig. 4 is a flowchart showing an example of the processing executed by the information processing device of embodiment 1. (Step S11) The acquisition unit 120 acquires video from the imaging device 200. (Step S12) The detection unit 130 detects the movement of the steering wheel 11 based on the video. (Step S13) The acquisition unit 120 acquires position information of the moved steering wheel 11 in the real space 10 from the imaging device 200. (Step S14) The acquisition unit 120 acquires a virtual information generation learning model from the storage unit 110.
[0033] (Step S15) The relearning unit 140 converts the position information in the real space 10 into the position information in the virtual space 20. (Step S16) The relearning unit 140 re-learns the virtual information generation learning model using an image including the handle 11 in a moved state and the position information in the virtual space 20. (Step S17) The acquisition unit 120 acquires line-of-sight information from the display device 300. (Step S18) The virtual information generation unit 160 generates three-dimensional virtual information based on the virtual information generation learning model and the line-of-sight information. (Step S19) The provision unit 170 provides the three-dimensional virtual information to the display device 300. This allows the user wearing the display device 300 to visually recognize the handle 21 in a downwardly tilted state in the virtual space 20.
[0034] According to the first embodiment, the information processing device 100 re-learns the virtual information generation learning model. As a result, the virtual information generation learning model can generate three-dimensional virtual information that represents a changed situation. Therefore, the information processing device 100 can generate three-dimensional virtual information that represents a changed situation.
[0035] Second Embodiment Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be mainly described. Furthermore, in the second embodiment, descriptions of the commonalities between the first embodiment and the second embodiment will be omitted.
[0036] 5 is a block diagram showing the functions of the information processing device of embodiment 2. The information processing device 100 further includes a multi-viewpoint image generation unit 180. Part or all of the multi-viewpoint image generation unit 180 may be realized by a processing circuit. Alternatively, part or all of the multi-viewpoint image generation unit 180 may be realized as a module of a program executed by the processor 101.
[0037] The multi-viewpoint image generation unit 180 generates, based on an image including an object in a moved state, one or more images of the object viewed from a direction different from the direction in which the object was captured, as one or more multi-viewpoint images. When multiple multi-viewpoint images are generated, the multi-viewpoint image generation unit 180 may be expressed as generating, based on an image including the object in a moved state, multiple images of the object viewed from multiple directions as multiple multi-viewpoint images. For example, the multi-viewpoint image generation unit 180 generates an image of the object in a moved state viewed from the right, an image of the object in a moved state viewed from the left, etc.
[0038] Furthermore, for example, when generating a multi-viewpoint image, the multi-viewpoint image generating unit 180 generates the multi-viewpoint image using Generative Adversarial Networks (GAN).
[0039] The acquisition unit 120 acquires one or more multi-viewpoint images from the multi-viewpoint image generation unit 180 .
[0040] The re-learning unit 140 re-learns the virtual information generation learning model using an image including an object in a moved state, one or more multi-view images, and the position information in the virtual space 20. The re-learning unit 140 may re-learn the virtual information generation learning model using an image including an object in a moved state in the real space 10, one or more multi-view images, the position information in the virtual space 20, line-of-sight information (i.e., information indicating the imaging direction), and direction information indicating the direction (e.g., the multiple directions) used when generating the one or more multi-view images.
[0041] Next, the processing executed by the information processing device 100 will be described using a flowchart. Fig. 6 is a flowchart showing an example of the processing executed by the information processing device of embodiment 2. The processing in Fig. 6 differs from the processing in Fig. 4 in that steps S15a and 16a are executed. Therefore, steps S15a and 16a will be described in Fig. 6. Further, a description of the processing other than steps S15a and 16a will be omitted.
[0042] (Step S15a) The multi-viewpoint image generation unit 180 generates one or more multi-viewpoint images based on an image including the object in a moved state. (Step S16a) The relearning unit 140 re-learns the virtual information generation learning model using an image including the object in a moved state in the real space 10, one or more multi-viewpoint images, and the position information in the virtual space 20.
[0043] According to the second embodiment, the information processing device 100 retrains the virtual information generation learning model using one or more multi-view images, thereby enabling the virtual information generation learning model to generate three-dimensional virtual information with higher accuracy.
[0044] In the above, the case where the acquisition unit 120 acquires one or more multi-view images from the multi-view image generation unit 180 has been described. The acquisition unit 120 may acquire one or more multi-view images from the imaging device 200. The relearning unit 140 may re-learn the virtual information generation learning model using one or more multi-view images and the position information in the virtual space 20.
[0045] In this way, the information processing device 100 re-learns the virtual information generation learning model using learning data that is at least one of an image including an object that has been moved in the real space 10 and one or more multi-perspective images, and the position information in the virtual space 20.
[0046] Embodiment 3 In the first and second embodiments, the case where the processing up to relearning is performed by one device has been described. The imaging device 200 may have some of the functions of the information processing device 100. For example, the imaging device 200 has some of the functions of the acquisition unit 120 and the detection unit 130. In this way, the processing up to relearning may be performed by multiple devices.
[0047] The embodiments are merely examples, and various modifications are possible within the scope of the present disclosure. Furthermore, the features of the respective embodiments can be combined with each other as appropriate.
[0048] 10 Real space, 11 Handle, 20 Virtual space, 21 Handle, 100 Information processing device, 101 Processor, 102 Volatile storage device, 103 Non-volatile storage device, 110 Storage unit, 120 Acquisition unit, 130 Detection unit, 140 Re-learning unit, 150 Calculation unit, 160 Virtual information generation unit, 170 Provision unit, 180 Multi-viewpoint image generation unit, 200 Imaging device, 300 Display device.
Claims
1. An information processing device having: an acquisition unit that acquires learning data that is at least one of an image including an object moved in real space and one or more multi-view images that are one or more images of the object when viewed from a direction different from the direction in which the object was captured, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information about the virtual space; and a re-learning unit that re-learns the virtual information generation learning model using the learning data and the positional information.
2. The information processing device according to claim 1, wherein the re-learning unit re-learns the virtual information generation learning model using an image including the object and the position information.
3. The information processing device according to claim 1, wherein the re-learning unit re-learns the virtual information generation learning model using the one or more multi-view images and the position information.
4. The information processing device according to claim 2 or 3, wherein the re-learning unit re-learns the virtual information generation learning model using an image including the object, the one or more multi-view images, and the position information.
5. A re-learning method in which one or more devices acquire learning data, which is at least one of an image including an object moved in real space and one or more multi-perspective images which are one or more images of the object when viewed from a direction different from the direction in which the object was captured, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information which is information about the virtual space, and re-learns the virtual information generation learning model using the learning data and the positional information.
6. A re-learning program that causes an information processing device to execute a process of acquiring learning data that is at least one of an image including an object moved in real space and one or more multi-perspective images that are one or more images of the object when viewed from a direction different from the direction in which the object was captured, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information about the virtual space, and re-learning the virtual information generation learning model using the learning data and the positional information.