Information processing device, relearning method, and relearning program

JPWO2025203561A5Active Publication Date: 2026-03-05MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024552482
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2026-03-05
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

Existing three-dimensional virtual information generated using NeRF does not reflect changes in the real space over time, leading to a mismatch between the virtual and real environments.

Method used

An information processing device that includes an acquisition unit to capture images of moved objects, a virtual information generation learning model, a multi-view image generation unit, and a relearning unit to update the model using position and line-of-sight information, enabling generation of three-dimensional virtual information that reflects changes in the real space.

Benefits of technology

Enables generation of three-dimensional virtual information that accurately represents the changed situation in the real space, allowing users to experience a dynamic virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000007_0000
    Figure 00000007_0000
  • Figure 00000007_0001
    Figure 00000007_0001
  • Figure 00000007_0002
    Figure 00000007_0002
Patent Text Reader

Abstract

The information processing device (100) has an acquisition unit (120) that acquires learning data which is at least one of an image including an object moved in real space and one or more multi-view images which are one or more images of an object viewed from a direction different from the direction in which the object was captured, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information which is information of the virtual space, and a re-learning unit (140) that re-learns the virtual information generation learning model using the learning data and the positional information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an information processing device, a relearning method, and a relearning program. [Background technology]

[0002] Devices such as HMDs (Head Mounted Displays) are known. For example, a user can wear an HMD and experience a virtual space through the HMD. A technology for generating three-dimensional virtual information, which is information about the virtual space, is known (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Ben Mildenhall et al. “NeRF:Representing Scenes as Neural Radiance Fields for View Synthesis” Summary of the Invention [Problem to be solved by the invention]

[0004] For example, the three-dimensional virtual information is generated using NeRF (Neural Radiance Fields) or the like. When the three-dimensional virtual information is generated, an image of the real space at a certain time is used. If the situation in the real space is changed after that time, the change in the situation in the real space is not reflected in the three-dimensional virtual information. Therefore, the user cannot experience the virtual space in the changed situation.

[0005] The objective of the present disclosure is to enable three-dimensional virtual information to be generated that shows the changed situation. [Means for solving the problem]

[0006] According to an embodiment of the present disclosure, there is provided an information processing device.image, an acquisition unit that acquires position information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information of the virtual space; a multi-viewpoint image generating unit that generates, based on an image including the object, one or more images obtained when the object is viewed from a direction different from a direction in which the object is captured, as one or more multi-viewpoint images; The device further includes a re-learning unit that re-learns the virtual information generation learning model by using the position information. Effect of the Invention

[0007] According to the present disclosure, three-dimensional virtual information can be generated that shows the changed situation. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an information processing system according to a first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating hardware included in an information processing device according to the first embodiment. [Diagram 3] 1 is a block diagram showing functions of an information processing device according to a first embodiment; [Figure 4] 4 is a flowchart showing an example of processing executed by the information processing device according to the first embodiment. [Diagram 5] FIG. 11 is a block diagram showing the functions of an information processing device according to a second embodiment. [Figure 6] 13 is a flowchart showing an example of processing executed by an information processing device according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, an embodiment will be described with reference to the drawings.

[0010] Embodiment 1 1 is a diagram showing an information processing system according to the first embodiment. The information processing system includes an information processing device 100 and an imaging device 200. The information processing system may also include a display device 300. The information processing device 100, the imaging device 200, and the display device 300 communicate with each other via a network.

[0011] For example, the information processing device 100 is a device that executes a re-learning method. For example, the image capturing device 200 is a see-through HMD. The image capturing device 200 may be a camera that captures an image of an object existing in the real space 10. For example, the display device 300 is a fully immersive HMD. The display device 300 can display the virtual space 20.

[0012] A brief description of the first embodiment will be given. The following specific example is merely an example. A user wearing the imaging device 200 moves the handlebar 11 downward. The imaging device 200 captures an image of the handlebar 11 in the downwardly moved state. The imaging device 200 transmits the image obtained by capturing the image to the information processing device 100.

[0013] The information processing device 100 can generate three-dimensional virtual information, which is information of the virtual space 20 corresponding to the real space 10, by using the virtual information generation learning model. The virtual information generation learning model can generate three-dimensional virtual information indicating the handle 21 in a state where it is not lowered downward. That is, the virtual information generation learning model generates three-dimensional virtual information in which the situation changed in the real space 10 is not reflected. Therefore, the information processing device 100 re-learns the virtual information generation learning model. In detail, the information processing device 100 re-learns the virtual information generation learning model by using the image. As a result, the virtual information generation learning model can generate three-dimensional virtual information indicating the changed situation.

[0014] The information processing device 100 receives gaze information from the display device 300. The information processing device 100 generates three-dimensional virtual information using a virtual information generation learning model and the gaze information. The information processing device 100 transmits the three-dimensional virtual information to the display device 300. This allows a user wearing the display device 300 to visually recognize the handle 21 in a downwardly lowered state in the virtual space 20.

[0015] The processes performed in the information processing system will be described in detail below.

[0016] The hardware included in the information processing device 100 will now be described. 2 is a diagram showing hardware included in the information processing device of embodiment 1. The information processing device 100 is also called a computer. The information processing device 100 includes a processor 101, a volatile storage device 102, and a non-volatile storage device 103.

[0017] The processor 101 controls the entire information processing device 100. For example, the processor 101 is a central processing unit (CPU) or a field programmable gate array (FPGA). The processor 101 may be a multiprocessor. The information processing device 100 may also include a processing circuit.

[0018] The volatile storage device 102 is a main storage device of the information processing device 100. For example, the volatile storage device 102 is a random access memory (RAM). The non-volatile storage device 103 is an auxiliary storage device of the information processing device 100. For example, the non-volatile storage device 103 is a hard disk drive (HDD) or a solid state drive (SSD).

[0019] Next, functions of the information processing device 100 will be described. 3 is a block diagram showing functions of the information processing device of embodiment 1. The information processing device 100 includes a storage unit 110, an acquisition unit 120, a detection unit 130, a relearning unit 140, a calculation unit 150, a virtual information generation unit 160, and a provision unit 170.

[0020] The storage unit 110 may be realized as a storage area secured in the volatile storage device 102 or the non-volatile storage device 103 . A part or all of the acquisition unit 120, the detection unit 130, the relearning unit 140, the calculation unit 150, the virtual information generation unit 160, and the provision unit 170 may be realized by a processing circuit. Also, a part or all of the acquisition unit 120, the detection unit 130, the relearning unit 140, the calculation unit 150, the virtual information generation unit 160, and the provision unit 170 may be realized as a module of a program executed by the processor 101. For example, the program executed by the processor 101 is also called a relearning program or a relearning program product. For example, the relearning program is recorded on a recording medium.

[0021] The storage unit 110 stores various information. The acquisition unit 120 acquires video from the imaging device 200 . When an image is acquired, the detection unit 130 detects the movement of an object based on the image. For example, the detection unit 130 detects the movement of the handle 11 based on the image. As a method for detecting the movement of an object, for example, the detection unit 130 detects the movement of an object by detecting a difference between images.

[0022] The acquisition unit 120 acquires from the detection unit 130 an image including an object in a state in which it has been moved in the real space 10. For example, the acquisition unit 120 acquires from the detection unit 130 an image including the handle 11 in a state in which it is lowered downward.

[0023] The acquisition unit 120 acquires position information of an object that has been moved in the real space 10 from the imaging device 200. For example, the acquisition unit 120 acquires from the imaging device 200 position information of the handle 11 in a downwardly lowered state.

[0024] The acquisition unit 120 acquires the virtual information generation learning model. For example, the acquisition unit 120 acquires the virtual information generation learning model from the storage unit 110. Also, for example, the acquisition unit 120 acquires the virtual information generation learning model from an external device. The external device is a device that exists outside the information processing device 100. For example, the external device is a cloud server, an external memory, etc. A diagram of the external device is omitted. Note that, for example, the virtual information generation learning model is a trained model using NeRF. The virtual information generation learning model may be called a NeRF learning model.

[0025] The relearning unit 140 converts the position information in the real space 10 into position information in the virtual space 20. For example, the relearning unit 140 performs the conversion using information indicating the correspondence between the real space 10 and the virtual space 20. The acquisition unit 120 acquires the position information in the virtual space 20 from the relearning unit 140.

[0026] The re-learning unit 140 re-learns the virtual information generation learning model using an image including an object in a moved state and the position information of the object in the virtual space 20. For example, the re-learning unit 140 re-learns the virtual information generation learning model using an image including the handle 11 in a downwardly lowered state and the position information of the handle in the virtual space 20.

[0027] The acquisition unit 120 may acquire line-of-sight information, which is the direction in which an image of an object in a moved state is captured, from the imaging device 200. For example, the acquisition unit 120 acquires line-of-sight information, which is information indicating the direction in which a user wearing the imaging device 200 is looking at the handlebar 11 in a lowered state (i.e., the imaging direction), from the imaging device 200. When the line-of-sight information is acquired, the re-learning unit 140 may re-learn the virtual information generation learning model using an image including the object in a moved state, the position information of the object in the virtual space 20, and the line-of-sight information.

[0028] In the above, a case has been described in which the acquisition unit 120 acquires an image including an object in a state in which it has been moved in the real space 10 from the detection unit 130. The acquisition unit 120 may acquire an image including an object in a state in which it has been moved in the real space 10 from the imaging device 200.

[0029] In the above, a case has been described in which the acquisition unit 120 acquires position information of an object moved in the real space 10 from the imaging device 200. The acquisition unit 120 may acquire the position information of the imaging device 200 in the real space 10 and the line of sight information from the imaging device 200. The calculation unit 150 may calculate the position information of the object moved in the real space 10 based on the position information of the imaging device 200 and the line of sight information. When the position information is calculated, the acquisition unit 120 acquires the position information from the calculation unit 150.

[0030] In the above, a case has been described in which the acquisition unit 120 acquires the position information in the virtual space 20 from the relearning unit 140. The acquisition unit 120 may acquire the position information in the virtual space 20 from the imaging device 200.

[0031] When the acquisition unit 120 acquires the line-of-sight information from the display device 300, the virtual information generation unit 160 generates three-dimensional virtual information based on the virtual information generation learning model and the line-of-sight information. The providing unit 170 provides the three-dimensional virtual information to the display device 300 .

[0032] Next, the process executed by the information processing device 100 will be described with reference to a flowchart. FIG. 4 is a flowchart illustrating an example of processing executed by the information processing device according to the first embodiment. (Step S11 ) The acquisition unit 120 acquires a video image from the imaging device 200 . (Step S12) The detection unit 130 detects the movement of the handle 11 based on the video. (Step S13) The acquisition unit 120 acquires, from the imaging device 200, position information of the moved handle 11 in the real space 10. (Step S14) The acquisition unit 120 acquires the virtual information generation learning model from the storage unit 110.

[0033] (Step S15) The re-learning unit 140 converts the position information in the real space 10 into position information in the virtual space 20. (Step S16) The re-learning unit 140 re-learns the virtual information generation learning model using an image including the handle 11 in a moved state and the position information in the virtual space 20. (Step S17) The acquiring unit 120 acquires line-of-sight information from the display device 300. (Step S18) The virtual information generator 160 generates three-dimensional virtual information based on the virtual information generation learning model and the line of sight information. (Step S19) The providing unit 170 provides the three-dimensional virtual information to the display device 300. As a result, the user wearing the display device 300 can visually recognize the handle 21 in a downwardly lowered state in the virtual space 20.

[0034] According to the first embodiment, the information processing device 100 re-learns the virtual information generation learning model. As a result, the virtual information generation learning model can generate three-dimensional virtual information that indicates a changed situation. Therefore, the information processing device 100 can generate three-dimensional virtual information that indicates a changed situation.

[0035] Embodiment 2 Next, a description will be given of embodiment 2. In embodiment 2, differences from embodiment 1 will be mainly described. In embodiment 2, descriptions of matters common to embodiment 1 will be omitted.

[0036] 5 is a block diagram showing functions of the information processing device of embodiment 2. The information processing device 100 further includes a multi-view image generating unit 180. A part or the whole of the multi-viewpoint image generating unit 180 may be realized by a processing circuit. Also, a part or the whole of the multi-viewpoint image generating unit 180 may be realized as a module of a program executed by the processor 101.

[0037] The multi-viewpoint image generating unit 180 generates, based on an image including an object in a moved state, one or more images of the object viewed from a direction different from the direction in which the object was captured, as one or more multi-viewpoint images. When multiple multi-viewpoint images are generated, the multi-viewpoint image generating unit 180 may be expressed as generating, based on an image including an object in a moved state, multiple images of the object viewed from multiple directions as multiple multi-viewpoint images. For example, the multi-viewpoint image generating unit 180 generates an image of the object in a moved state viewed from the right direction, an image of the object in a moved state viewed from the left direction, and the like.

[0038] Furthermore, for example, when generating a multi-viewpoint image, the multi-viewpoint image generating unit 180 generates the multi-viewpoint image using Generative Adversarial Networks (GAN).

[0039] The acquisition unit 120 acquires one or more multi-viewpoint images from the multi-viewpoint image generation unit 180 .

[0040] The re-learning unit 140 re-learns the virtual information generation learning model using an image including an object in a moved state, one or more multi-view images, and the position information in the virtual space 20. The re-learning unit 140 may re-learn the virtual information generation learning model using an image including an object in a moved state in the real space 10, one or more multi-view images, the position information in the virtual space 20, line-of-sight information (i.e., information indicating the imaging direction), and direction information indicating the direction (e.g., the multiple directions) used when generating the one or more multi-view images.

[0041] Next, the process executed by the information processing device 100 will be described with reference to a flowchart. Fig. 6 is a flowchart showing an example of a process executed by the information processing device of the embodiment 2. The process in Fig. 6 differs from the process in Fig. 4 in that steps S15a and 16a are executed. Therefore, steps S15a and 16a will be described in Fig. 6. Descriptions of the processes other than steps S15a and 16a will be omitted.

[0042] (Step S15a) Multi-viewpoint image generation section 180 generates one or more multi-viewpoint images based on an image including an object in a moved state. (Step S16a) The re-learning unit 140 re-learns the virtual information generation learning model using an image including an object in a state in which it has been moved in the real space 10, one or more multi-view images, and the position information in the virtual space 20.

[0043] According to the second embodiment, the information processing device 100 retrains the virtual information generation learning model using one or more multi-view images. Therefore, the virtual information generation learning model can generate more accurate three-dimensional virtual information.

[0044] In the above, the case where the acquisition unit 120 acquires one or more multi-view images from the multi-view image generation unit 180 has been described. The acquisition unit 120 may acquire one or more multi-view images from the imaging device 200. The re-learning unit 140 may re-learn the virtual information generation learning model using one or more multi-view images and the position information in the virtual space 20.

[0045] In this way, the information processing device 100 re-learns the virtual information generation learning model using learning data which is at least one of an image including an object that has been moved in the real space 10 and one or more multi-perspective images, and the position information in the virtual space 20.

[0046] Third embodiment In the first and second embodiments, the case where the process up to the relearning is realized by one device has been described. The imaging device 200 may have part of the functions of the information processing device 100. For example, the imaging device 200 has part of the functions of the acquisition unit 120 and the detection unit 130. In this manner, the process up to the relearning may be realized by multiple devices.

[0047] The embodiments are merely examples, and various modifications are possible within the scope of the present disclosure. Furthermore, the features of the respective embodiments can be appropriately combined with each other. [Explanation of symbols]

[0048] 10 real space, 11 handle, 20 virtual space, 21 handle, 100 information processing device, 101 processor, 102 volatile storage device, 103 non-volatile storage device, 110 memory unit, 120 acquisition unit, 130 detection unit, 140 relearning unit, 150 calculation unit, 160 virtual information generation unit, 170 provision unit, 180 multi-viewpoint image generation unit, 200 imaging device, 300 display device.

Claims

1. an acquisition unit that acquires an image including an object moved in a real space, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information of the virtual space; a multi-viewpoint image generating unit that generates, based on an image including the object, one or more images of the object viewed from a direction different from a direction in which the object was captured, as one or more multi-viewpoint images; a re-learning unit that re-learns the virtual information generation learning model using an image including the object, the one or more multi-view images, and the position information; An information processing device having the above.

2. One or more devices Acquire an image including an object moved in a real space, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information of the virtual space; generating, based on an image including the object, one or more images of the object viewed from a direction different from the direction in which the object was captured, as one or more multi-viewpoint images; re-training the virtual information generation learning model using the image including the object, the one or more multi-view images, and the position information; How to relearn.

3. In the information processing device, Acquire an image including an object moved in a real space, positional information of the object in a virtual space corresponding to the real space, and a virtual information generation learning model that generates three-dimensional virtual information that is information of the virtual space; generating, based on an image including the object, one or more images of the object viewed from a direction different from the direction in which the object was captured, as one or more multi-viewpoint images; re-training the virtual information generation learning model using the image including the object, the one or more multi-view images, and the position information; A re-learning program that executes the process.