Information processing apparatus, information processing method, and program
By using image acquisition and processing technology in the information processing device, extracting mask areas and generating display images of virtual objects, the problem of incorrect display of virtual objects in mixed reality technology is solved, and the user experience is improved.
Patent Information
- Application Number
- JP2023182387
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-24
- Publication Date
- 2025-05-09
AI Technical Summary
In the mixed reality technology, the user cannot effectively avoid seeing objects that do not belong to the hand in the virtual space, resulting in poor user experience.
In the information processing device, an image containing a virtual object is generated using an image acquisition unit, a color registration unit, a mask area extraction unit, an object information acquisition unit, and an image generation unit. The device extracts the mask area, determines the boundaries of the mask area based on registered color information or deep learning networks, and generates a display image of the virtual object to ensure that the virtual object is only displayed in front of the hand.
It effectively reduces the situation where users see objects that do not belong to their hands in the virtual space, thereby improving the user experience and reducing user discomfort.
Smart Images

Figure 2025071947000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In recent years, mixed reality (MR) technology has become known as a technology for fusing real space and virtual space in real time. In MR technology, an HMD (head mounted display) displays a composite image in which an image of a real space captured by an imaging device is superimposed with an image of a virtual space corresponding to the position and orientation of the imaging device.
[0003] In this case, depending on the distance between the imaging device and the real object and the distance between the imaging device and the virtual object, the sense of distance between the objects can be expressed by not displaying a virtual object in the area of a specific real object in the captured image. For example, if a virtual object is not drawn in the area of a hand in the captured image, the hand can be displayed as being positioned in front of the virtual object. This makes it easier for the user to grasp the positional relationship between the virtual object and the real object, and makes it easier to work (verify) using one's own hand in the virtual space.
[0004] In Patent Document 1, the device extracts a pre-registered color region of the hand from an image captured by a stereo camera in order to detect the positional relationship between the user's hand and a virtual object. The device also generates a polygon model of the hand based on the contour of the extracted region and the depth calculated from the corresponding points of the stereo camera. When it is determined that the polygon model of the hand is located in front of the virtual object, the device draws the polygon model of the hand in front, thereby appropriately displaying the user's hand. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2008-210276 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, in Patent Document 1, if an object with the same color area as the color area registered in advance by the user is placed in front of the virtual object, an object other than the hand is drawn, which gives the user a sense of incongruity.
[0007] Therefore, an object of the present invention is to reduce the sense of discomfort felt by a user when the user is made to perceive a space in which virtual objects are arranged in real space. [Means for solving the problem]
[0008] One aspect of the present invention is a method for producing a composition comprising the steps of: An information processing device that generates a first image to be displayed on a display device, the first image including a virtual object, an image acquisition means for acquiring a second image obtained by capturing an image of a real space; an extraction means for extracting a mask region estimated to be a region of a specific object from the second image based on the second image; An acquisition means for acquiring object information including information on a position or a shape of an object based on the second image; an image generating means for generating the first image based on the mask region and the object information of the specific object; The information processing device is characterized by having:
[0009] One aspect of the present invention is a method for producing a composition comprising the steps of: 1. An information processing method for generating a first image, the first image including a virtual object, to be displayed on a display device, comprising: an image acquiring step of acquiring a second image capturing a real space; an extraction step of extracting a mask region from the second image, the mask region being estimated to be a region of a specific object based on the second image; acquiring object information including information on a position or a shape of an object based on the second image; an image generating step of generating the first image based on the mask region and the object information of the specific object; The information processing method is characterized by having the following features. Effect of the Invention
[0010] According to the present invention, when the user is made to perceive a space in which a virtual object is arranged in a real space, the sense of discomfort felt by the user can be reduced. [Brief description of the drawings]
[0011] [Figure 1] 1 is a block diagram of an information processing device according to a first embodiment. [Diagram 2] 4 is a flowchart of processing by the information processing device according to the first embodiment. [Diagram 3] FIG. 1 is a diagram illustrating a display system according to a first embodiment. [Figure 4] 4 is a flowchart of a process of a mask region extraction unit according to the first embodiment. [Diagram 5] 4 is a flowchart of a process of an object information acquisition unit according to the first embodiment. [Figure 6] 4 is a flowchart of a process of an image generating unit according to the first embodiment. [Figure 7] FIG. 2 is a diagram for explaining the bones of a hand according to the first embodiment. [Figure 8] 10 is a flowchart of a process of an image generating unit according to the second embodiment. [Figure 9] 13A to 13C are diagrams illustrating generation of a mask model according to Modification 2. [Figure 10] 13 is a flowchart of a process of an object information acquisition unit according to the third embodiment. [Figure 11] 11 is a flowchart of a process of an image generating unit according to the third embodiment. [Figure 12] FIG. 4 is a diagram for explaining selection of a mask region according to the first embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, an embodiment for carrying out the present invention will be described in detail. In an MR space, a user may attempt to touch a virtual object placed in the MR space with his / her hand. For this reason, it is preferable for the user to see his / her own hand in front of the virtual object. Hereinafter, a display system 1 that appropriately displays the user's hand and a virtual object overlapping the hand will be described. However, instead of the "user's hand," any real object (such as a specific part of a person or a controller) existing in the real space may be used.
[0013] <Embodiment 1> A display system 1 (information processing system) according to the first embodiment will be described with reference to FIG. 3. The display system 1 includes an information processing device 1000, an imaging device 1100, and a display device 1200. The information processing device 1000 is, for example, a computer or a smartphone. For example, the imaging device 1100 is a camera incorporated in a video see-through type HMD (Head Mounted Display). The display device 1200 is a display in an HMD or a PC monitor. For this reason, the imaging device 1100 and the display device 1200 together can be said to be an HMD (a display device that can be worn on a user's head).
[0014] FIG. 1 is a block diagram of an information processing device 1000. It comprises an image acquisition unit 1010 , a color registration unit 1020 , a data storage unit 1030 , a mask region extraction unit 1040 , an object information acquisition unit 1050 , and an image generation unit 1060 .
[0015] The information processing device 1000 has two execution modes: a color registration mode and an experience mode. In the color registration mode, the information processing device 1000, for example, registers the color of a hand (real object) that appears in a captured image in which a real space is captured by the imaging device 1100. In the experience mode, the information processing device 1000 extracts, as a mask area, an area that is estimated to be a hand area from the captured image based on the color of the hand that has been registered in advance. Then, the information processing device 1000 controls the display device 1200 to display an image (an image representing a mixed reality space) corresponding to the mask area.
[0016] The processing of the information processing device 1000 according to the first embodiment will be described with reference to the flowchart in Fig. 2. The processing of each component of the information processing device 1000 will be described below. Note that, in the following, it is assumed that the captured image is an image (live view image) of a real space captured in real time. In other words, the user can experience a space in which a virtual object is arranged in a real space in real time by viewing the display of the HMD.
[0017] In step S2010, the captured image acquisition unit 1010 stores the captured image of the hand captured by the imaging device 1100.
[0018] In step S2020, the captured image acquisition unit 1010 stores the captured image in the data storage unit 1030.
[0019] In step S2021, the captured image acquisition unit 1010 determines the execution mode of the information processing device 1000. If it is determined that the execution mode is the color registration mode, the process proceeds to step S2030. If it is determined that the execution mode is the trial mode, the process proceeds to step S2040.
[0020] In step S2030, the color registration unit 1020 registers (stores) information about the color of the hand appearing in the captured image acquired from the captured image acquisition unit 1010 in a color registration table of the data storage unit 1030. For example, the color registration unit 1020 registers information about the color indicated by a pixel designated by the user among all pixels of the captured image in the color registration table as information about the color of the hand.
[0021] In step S2040, the mask region extraction unit 1040 extracts a mask region that is estimated to be a hand region from the captured image. For example, the mask region extraction unit 1040 extracts a mask region from the captured image based on the captured image and a color registration table stored in the data storage unit 1030. As a result, if hand color information is registered in the color registration table, the hand color region is extracted as the mask region. The mask region extraction unit 1040 may use any known method as long as it can extract a region that is estimated to be a hand region. Details of the process of step S2040 will be described later using the flowchart of FIG. 4.
[0022] In step S2050, the object information acquisition unit 1050 acquires information about a hand (real object) as object information based on the captured image stored in the data storage unit 1030. For example, the object information acquisition unit 1050 acquires information about the positions of bones of a person's hand as object information using a deep neural network. The object information acquisition unit 1050 may acquire point cloud information or shape information of the hand as object information. Details of the process of step S2050 will be described later with reference to the flowchart of FIG. 5.
[0023] In step S2060, the image generating unit 1060 acquires the virtual object and the captured image stored in the data storage unit 1030. In addition, the image generating unit 1060 acquires the virtual object and the captured image stored in the data storage unit 1030. Based on the object information acquired from the image generating unit 1060, the mask area is corrected so as to display only the hand area. The image generating unit 1060 generates a mask model (CG model) corresponding to the corrected mask area. The mask model is a model that controls so as not to display a virtual object even if the virtual object is superimposed.
[0024] In step S2060, the image generating unit 1060 generates a composite image by superimposing the captured image, the mask model, and the virtual object in this order. In this way, the image of the hand in the captured image is displayed in the range where the captured image, the mask model, and the virtual object are superimposed in this order. Therefore, a composite image can be generated in which the hand is displayed in the foreground rather than the virtual object. Details of the process in step S2060 will be described later with reference to the flowchart in FIG. 6.
[0025] The flowchart in FIG. 4 shows details of the process of step S2040 (the process of the mask region extraction unit 1040).
[0026] In step S 4010 , the mask region extraction unit 1040 acquires the captured image from the data storage unit 1030 .
[0027] In step S4020, the mask region extraction unit 1040 refers to the user's settings and determines whether or not to use a color registration table to extract the mask region. If it is determined that the color registration table is to be used, the process proceeds to step S4030. If it is determined that the color registration table is not to be used, the process proceeds to step S4040.
[0028] In step S4030, the mask region extraction unit 1040 acquires from the data storage unit 1030 a color registration table in which hand color information is registered.
[0029] In step S4040, the mask region extraction unit 1040 acquires hand region information. For example, the mask region extraction unit 1040 acquires information on a region recognized as a person's hand in a captured image as hand region information using a pre-trained deep neural network. The deep neural network is trained in advance by inputting, as training data, a large number of combinations of one image and the hand region in the image.
[0030] In step S4050, the mask region extraction unit 1040 extracts the region of the hand color registered in the color registration table in the captured image as a mask region.
[0031] In step S4060, the mask region extraction unit 1040 extracts, as a mask region, a region recognized as a hand in the captured image, based on the hand region information acquired in step S4040.
[0032] The flowchart in FIG. 5 shows details of the process of step S2050 (the process of the object information acquisition unit 1050).
[0033] In step S5010 , the object information acquisition unit 1050 acquires the captured image from the data storage unit 1030 .
[0034] In step S5020, the object information acquisition unit 1050 acquires, for example, information on the position (three-dimensional position) of the hand as object information based on the captured image. The object information acquisition unit 1050 acquires, for example, information on the positions of the bones (bones; skeleton) of the hand as object information by using a pre-trained deep neural network. Alternatively, the object information acquisition unit 1050 may acquire point cloud information of the hand as object information. For example, the object information acquisition unit 1050 acquires, when the position of the object is If possible, the three-dimensional position of the object may be obtained by any known method.
[0035] Figure 7 shows an image of hand bones. Hand bones are, for example, line segments connecting the joint points of the fingers. The positions of the hand bones indicate the three-dimensional position of the hand.
[0036] The flowchart in FIG. 6 shows details of the process of step S2060 (the process of the image generating unit 1060).
[0037] In step S6010, the image generating unit 1060 acquires the captured image from the data storage unit 1030.
[0038] In step S6020, the image generation unit 1060 acquires a virtual object from the data storage unit 1030.
[0039] In step S6030, the image generation unit 1060 acquires the mask region from the mask region extraction unit 1040.
[0040] In step S6040, the image generation unit 1060 acquires the object information from the object information acquisition unit 1050.
[0041] In step S6050, the image generating unit 1060 calculates the distance between the mask region and the hand based on the hand position information included in the object information. The image generating unit 1060 also determines whether the distance between the mask region and the hand is equal to or less than a predetermined distance. Here, the distance between the two regions may be the distance between the center points of the two regions, or the shortest distance (interval) between the two regions.
[0042] 12A to 12C are diagrams for explaining the process of step S6050. For example, as shown in FIG. 12A, it is assumed that two mask regions (hand mask region 1201 and box mask region 1202 having the same color information) are extracted by mask region extraction unit 1040. In this case, image generation unit 1060 calculates the distance between each of two mask regions 1201 and 1202 and the hand. Specifically, image generation unit 1060 calculates the distance between hand mask region 1201 and hand bone 1203, and the distance between box mask region 1202 and hand bone 1203. For example, in the case shown in FIG. 12B, hand bone 1203 is determined to be shorter than a predetermined distance from hand mask region 1201, but longer than a predetermined distance from box mask region 1202. That is, only hand mask region 1201 as shown in FIG. 12C is determined to be shorter than a predetermined distance from the hand.
[0043] In step S6060, the image generation unit 1060 selects a mask region whose distance from the hand position is determined to be equal to or less than a predetermined distance, and generates a mask model based only on the selected mask region.
[0044] In addition, if the object information includes information on the shape of a hand, in steps S6050 and S6060, the image generation unit 1060 may generate a mask model based on only the mask area that has a shape closest to the hand shape among the extracted multiple mask areas.
[0045] In step S6070, the image generating unit 1060 superimposes the mask model and the virtual object on the captured image. In this way, the image generating unit 1060 generates a composite image to be displayed on the display device 1200.
[0046] According to the first embodiment, the information processing device generates a composite image by combining a captured image and a virtual object. In this case, an object having color information of the hand can be extracted as a mask area to generate a composite image. In addition, when a mask area other than the user's hand area is extracted, the mask area of the object other than the hand can be excluded based on the object information. This makes it possible to appropriately select the mask area, and therefore it is possible to generate a composite image that reduces the sense of incongruity felt by the user. Therefore, when the user is made to perceive a space in which a virtual object is placed in real space, the sense of incongruity felt by the user can be reduced.
[0047] (Variation 1) In the first embodiment, the object information acquisition unit 1050 acquires information on the position of the hand (object) as object information based on the captured image using a deep neural network. The object information acquisition unit 1050 may acquire hand position information (bone position information) using a device such as a sensor that can measure the depth of an object, not based on the captured image. Specifically, the object information acquisition unit 1050 may acquire object information such as hand position information based on measurement information of a sensor that measures an object (such as the distance to the object or the position of the object) shown in the captured image.
[0048] (Variation 2) 9A, when a hand holds a box of the same color information as the hand, the boundary between the hand and the outline of the box cannot be determined, and the area including the hand and the box may be extracted as a single mask area. Therefore, the image generating unit 1060 may correct the mask area (the size of the mask area) based on the object information so as to include only the area close to the hand.
[0049] According to the object information acquired by the object information acquisition unit 1050, for example, the position of the hand bone can be determined as shown in FIG. 9B. Therefore, when the outline of the mask area and the hand bone are separated by more than a specific distance, the image generation unit 1060 determines that an object other than the hand is also included in the mask area. Then, the image generation unit 1060 excludes the range in which the hand bone is separated by more than a specific distance from the mask area. Alternatively, the image generation unit 1060 may correct the shape of the mask area so as to approach the shape of the hand bone. According to these, the image generation unit 1060 can correct the mask area so as to approach the size of the hand area, and therefore can generate a mask model with a more appropriate range as shown in FIG. 9C.
[0050] <Embodiment 2> In the first embodiment, in step S6050, the image generating unit 1060 generates a mask model based on bone information of the hand. However, when the hands of multiple users (persons) are included in the captured image, even if it is desired to extract only the hand region of the user (hereinafter referred to as the "user") who uses the display system 1 as the mask region, the hand regions of the users other than the user are also extracted as the mask region. This gives a sense of incongruity to the user who views the composite image.
[0051] The flowchart in Fig. 8 shows the processing of the image generating unit 1060 according to the embodiment 2. In steps S6010 to S6070, the same processing as that described in the embodiment 1 is performed. In the embodiment 2, the processing of step S8001 is performed between steps S6040 and S6050.
[0052] In step S8001, the image generating unit 1060 determines information on the hand of the user among the hands captured by the imaging device. For example, the image generating unit 1060 determines that the hand closest to the imaging device 1100 is the hand of the user. Then, the image generating unit 1060 deletes object information other than the hand of the user from the object information of the multiple hands acquired in step S6040. This allows the image generating unit 1060 to select only the object information of the hand of the user, and to use only the object information of the selected hand in determining the distance between the mask region and the hand in step S6050.
[0053] According to the second embodiment, it is possible to reduce the possibility that the hand area of a person other than the user is used to generate a mask model. Therefore, it is possible to reduce the possibility that the hand of a person other than the user is captured in the composite image, thereby reducing the sense of incongruity felt by the user when viewing the composite image.
[0054] <Embodiment 3> In the first and second embodiments, the object information acquisition unit 1050 has been described on the assumption that it acquires only the object information of the hand. However, the object information acquisition unit 1050 may acquire object information of all or a plurality of main objects appearing in the captured image.
[0055] As shown in the flowchart of FIG. 10, in step S5020, the object information acquisition unit 1050 acquires information on the three-dimensional positions of all objects (or any number of objects) included in the captured image. Then, in step S5030, the object information acquisition unit 1050 acquires attribute information of all objects from the captured image. The attribute information includes information on the shape of the object, such as a circle or a square, and on a part of the human body, such as a face or a hand. The attribute information of the object may be extracted by any known method, such as using a deep neural network. Therefore, in the third embodiment, the object information includes information on the position, shape, and the part (part of the human body) to which the object corresponds.
[0056] 11 shows the process of the image generating unit 1060 according to the third embodiment. Steps S6010 to S6070 are the same as those described in the first embodiment.
[0057] In step S6010, the image generating unit 1060 selects only the object information of the hand from the object information of all objects based on the attribute information. The image generating unit 1060 deletes object information other than the hand from the object information of all objects. This allows only the object information of the hand to be used to determine the distance between the mask area and the hand in step S6050.
[0058] According to the third embodiment, it is possible to reduce the possibility that an area other than the hand is used to generate a mask model. Therefore, it is possible to reduce the possibility that an unnecessary real object other than the hand is captured in the synthetic image, thereby reducing the sense of incongruity felt by the user when viewing the synthetic image.
[0059] In each embodiment, a video see-through HMD has been described, but an optical see-through HMD that allows the real space to be viewed through the screen of the display device may be used. In this case, the image generating unit 1060 generates a composite image by combining a virtual object and a mask model.
[0060] Although the present invention has been described in detail based on the preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.
[0061] Also, in the above, "If A is equal to or greater than B, proceed to step S1; if A is smaller (lower) than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1; if A is equal to or less than B, proceed to step S2." Conversely, "If A is greater (higher) than B, proceed to step S1; if A is equal to or less than B, proceed to step S2" may be read as "If A is greater (higher) than B, proceed to step S1; if A is smaller (lower) than B, proceed to step S2." Therefore, unless a contradiction arises, "equal to or greater than A" may be read as "greater (higher; longer; more) than A," and "equal to or less than A" may be read as "smaller (lower; shorter; less) than A." And "Bigger (higher; longer; more) than A" may be read as "A or greater," and "smaller (lower; shorter; fewer) than A" may be read as "A or less."
[0062] Each functional unit in each of the above embodiments (variations) may or may not be individual hardware. The functions of two or more functional units may be realized by common hardware. Each of a plurality of functions of one functional unit may be realized by individual hardware. Two or more functions of one functional unit may be realized by common hardware. Furthermore, each functional unit may or may not be realized by hardware such as an ASIC, FPGA, or DSP. For example, the device may have a processor and a memory (storage medium) in which a control program is stored. Then, the functions of at least some of the functional units of the device may be realized by the processor reading and executing the control program from the memory.
[0063] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions.
[0064] The disclosure of the above embodiments includes the following configurations, methods, and programs. (Configuration 1) An information processing device that generates a first image to be displayed on a display device, the first image including a virtual object, an image acquisition means for acquiring a second image obtained by capturing an image of a real space; an extraction means for extracting a mask region estimated to be a region of a specific object from the second image based on the second image; An acquisition means for acquiring object information including information on a position or a shape of an object based on the second image; an image generating means for generating the first image based on the mask region and the object information of the specific object; 13. An information processing device comprising: (Configuration 2) the extraction means refers to a table in which color information of the specific object is registered in advance, and extracts the mask region based on the color information registered in the table. 2. The information processing device according to configuration 1. (Configuration 3) The extraction means extracts the mask region using a pre-trained deep neural network. 2. The information processing device according to configuration 1. (Configuration 4) The specific object is a specific part of a human body. 4. The information processing device according to any one of configurations 1 to 3. (Configuration 5) The specific object is a specific part of a user who uses the display device. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 6) The specific part is a human hand. 6. The information processing device according to configuration 4 or 5. (Configuration 7) The object information includes shape information and information on a corresponding part of a human body. 7. The information processing device according to any one of configurations 4 to 6. (Configuration 8) The acquisition means acquires the object information of a plurality of objects appearing in the second image based on the second image. 8. The information processing device according to any one of configurations 1 to 7. (Configuration 9) The image generating means 1) selects the object information of the specific object from the object information of the plurality of objects, and 2) generates the first image based on the selected object information and the mask region. 9. The information processing device according to configuration 8. (Configuration 10) The acquisition means acquires the object information using a pre-trained deep neural network. 10. The information processing device according to any one of configurations 1 to 9. (Configuration 11) the acquiring means acquires the object information based on measurement information of a sensor that measures the specific object. 10. The information processing device according to any one of configurations 1 to 9. (Configuration 12) The image generating means includes: correcting a size of the mask region so as to approximate a size of a region of the specific object based on the object information of the specific object; generating the first image based on the corrected mask region; 12. The information processing device according to any one of configurations 1 to 11. (Configuration 13) When a plurality of the mask regions are extracted by the extraction means, the image generation means 1) selects at least one mask region from the plurality of mask regions based on object information of the specific object, and 2) generates the first image based on the selected mask region. 13. The information processing device according to any one of configurations 1 to 12. (Configuration 14) The object information includes location information, When the plurality of mask regions have been extracted by the extraction means, the image generation means selects at least one of the mask regions that is located at a distance from the position of the specific object that is shorter than a predetermined distance. 14. The information processing device according to configuration 13. (Configuration 15) the image generating means generates the first image by combining a model based on the mask region, the second image, and the virtual object. 15. The information processing device according to any one of configurations 1 to 14. (method) 1. An information processing method for generating a first image, the first image including a virtual object, to be displayed on a display device, comprising: an image acquiring step of acquiring a second image capturing a real space; an extraction step of extracting a mask region from the second image, the mask region being estimated to be a region of a specific object based on the second image; acquiring object information including information on a position or a shape of an object based on the second image; an image generating step of generating the first image based on the mask region and the object information of the specific object; 13. An information processing method comprising: (program) A program for causing a computer to function as each of the means of the information processing device according to any one of configurations 1 to 15. [Explanation of symbols]
[0065] 1000: information processing device, 1100: imaging device, 1200: display device, 1010: captured image acquisition unit, 1040: mask region extraction unit, 1050: Object information acquisition unit, 1060: Image generation unit
Claims
1. An information processing device that generates a first image to be displayed on a display device, the first image including a virtual object, an image acquisition means for acquiring a second image obtained by capturing an image of a real space; an extraction means for extracting a mask region estimated to be a region of a specific object from the second image based on the second image; an acquisition means for acquiring object information including information on a position or a shape of an object based on the second image; an image generating means for generating the first image based on the mask region and the object information of the specific object; 13. An information processing device comprising:
2. the extraction means refers to a table in which color information of the specific object is registered in advance, and extracts the mask region based on the color information registered in the table.
2. The information processing apparatus according to claim 1,
3. The extraction means extracts the mask region using a pre-trained deep neural network.
2. The information processing apparatus according to claim 1,
4. The specific object is a specific part of a human body.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
5. The specific object is a specific part of a user who uses the display device.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
6. The specific part is a human hand.
5. The information processing apparatus according to claim 4.
7. The object information includes shape information and information on a corresponding part of a human body.
5. The information processing apparatus according to claim 4.
8. The acquisition means acquires the object information of a plurality of objects appearing in the second image based on the second image.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
9. The image generating means 1) selects the object information of the specific object from the object information of the plurality of objects, and 2) generates the first image based on the selected object information and the mask region.
9. The information processing apparatus according to claim 8,
10. The acquisition means acquires the object information using a pre-trained deep neural network.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
11. the acquiring means acquires the object information based on measurement information of a sensor that measures the specific object.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
12. The image generating means includes: correcting a size of the mask region so as to approximate a size of a region of the specific object based on the object information of the specific object; generating the first image based on the corrected mask region; 4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
13. When a plurality of the mask regions are extracted by the extraction means, the image generation means 1) selects at least one mask region from the plurality of mask regions based on object information of the specific object, and 2) generates the first image based on the selected mask region.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
14. The object information includes location information, When the plurality of mask regions have been extracted by the extraction means, the image generation means selects at least one of the mask regions that is located at a distance from the position of the specific object that is shorter than a predetermined distance.
14. The information processing apparatus according to claim 13,
15. the image generating means generates the first image by combining a model based on the mask region, the second image, and the virtual object.
4. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
16. 1. An information processing method for generating a first image, the first image including a virtual object, to be displayed on a display device, the method comprising: an image acquisition step of acquiring a second image obtained by capturing an image of a real space; an extraction step of extracting a mask region from the second image, the mask region being estimated to be a region of a specific object based on the second image; acquiring object information including information on a position or a shape of an object based on the second image; an image generating step of generating the first image based on the mask region and the object information of the specific object; 13. An information processing method comprising:
17. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and device for generating three-dimensional model information
JP2008210276A