Information processing device, information processing method, and computer program
The information processing apparatus generates and updates virtual content using environmental information to harmonize with the surrounding landscape, addressing the challenge of aesthetic integration in conventional building design methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-22
- Publication Date
- 2026-06-03
Smart Images

Figure 2026091188000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, a computer program, and the like.
Background Art
[0002] For those with high costs for newly creating physical objects, such as new buildings, furniture, artificial trees, etc. (hereinafter referred to as buildings), it is often the case that the design of the object is created in computer graphics in advance to check the appearance.
[0003] However, there is a problem that the cost of the work of human design of the building design is high. Therefore, conventionally, a method of automatically generating the building design by a computer has been proposed.
[0004] For example, Patent Document 1 discloses a method of automatically generating the design of a new 3D model having the same characteristics as an existing group of buildings. Also, Patent Document 2 discloses a method of automatically generating the design of a 3D model in consideration of the regulation conditions set for the land. Also, Patent Document 3 describes a method of constructing and using an object arrangement characteristic database.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0006]
Non-Patent Document 1
[0007] However, with the conventional methods described in Patent Documents 1, 2, etc., the design of the generated structure may not be in harmony with the surrounding landscape.
[0008] One of the objectives of this invention is to provide an information processing device capable of generating virtual content that harmonizes with the surrounding landscape and other elements. [Means for solving the problem]
[0009] An information processing apparatus in an embodiment of the present invention is An environmental information acquisition means for acquiring environmental information related to the environment, A generation means for generating virtual content based on the aforementioned environmental information, An update means for updating the virtual content to harmonize with the environment based on the aforementioned environmental information, characterized by having
Advantages of the Invention
[0010] According to the present invention, it is possible to provide an information processing apparatus capable of generating virtual content that harmonizes with the surrounding landscape and the like.
Brief Description of the Drawings
[0011] [Figure 1] FIG. is a hardware configuration example showing an example of an information processing apparatus according to Embodiment 1 of the present invention. [Figure 2] FIG. is an image diagram showing an example of a scene for generating virtual content in Embodiment 1. [Figure 3] FIG. is a functional block diagram showing an example of a logical configuration example of the information processing apparatus 1 in Embodiment 1. [Figure 4] FIG. is a flowchart showing a processing example of an information processing method using the information processing apparatus 1 in Embodiment 1. [Figure 5] FIG. is a functional block diagram showing an example of a logical configuration of the information processing apparatus 200 in Embodiment 2. [Figure 6] FIG. is a flowchart showing a processing example of an information processing method using the information processing apparatus 200 in Embodiment 2. [Figure 7] FIG. is a flowchart showing a detailed processing example of step S2040 in FIG. 6. <实
Mode for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each figure, the same members or elements are denoted by the same reference numerals, and duplicate explanations are omitted or simplified.
[0013] <Embodiment 1> Figure 1 is a diagram showing an example of a hardware configuration of an information processing device according to Embodiment 1 of the present invention. The CPU 10, acting as a computer, controls each part connected to the bus 60 via the bus 60.
[0014] The input interface 40 acquires input signals from external devices (such as imaging devices, display devices, or operating devices) in a format that the information processing device can process. The output interface 50 outputs output signals to external devices (such as display devices) in a format that the external devices can process.
[0015] The computer programs for realizing the functions of each embodiment are stored in a storage medium such as read-only memory (ROM) 20. The ROM 20 also stores the operating system (OS) and device drivers.
[0016] The RAM (Random Access Memory) 30 temporarily stores these programs. The CPU 10 then executes the computer programs stored in the RAM 30, thereby performing processes according to the flowcharts described later and realizing the functions of each embodiment.
[0017] Furthermore, instead of using software processing with the CPU 10, it is also possible to realize the functions of each embodiment using hardware having an arithmetic unit or circuit corresponding to the processing of each functional unit.
[0018] Embodiment 1 describes an example in which a Mixed Reality (MR) system (hereinafter referred to as the MR system) is used as an external device. Furthermore, the virtual content that the user wants to generate will be described as, for example, a building, but the virtual content is not limited to buildings.
[0019] Figure 2 is an illustrative diagram showing an example of a scene in which the information processing device in Embodiment 1 generates virtual content. The external device, MR System 2, uses a camera mounted on its head-mounted display to capture images of the surrounding environment and generates information about the surrounding environment (hereinafter referred to as "environmental information") based on the captured images.
[0020] The information processing device 1, which is the main component of this embodiment, receives environmental information generated by the MR system 2, generates and updates virtual content 5 that harmonizes with the surrounding environment, and transmits it to the MR system 2. Buildings 3 and 4 are buildings that exist in the real environment.
[0021] Next, the MR system 2 receives virtual content 5 as input, generates an image 6 (hereinafter referred to as the MR image) by combining the virtual content 5 with images of the real environment captured by the camera (for example, buildings 3 and 4), and displays it on the head-mounted display. The user of the MR system 2 can observe image 6 through the head-mounted display to confirm the appearance of the generated virtual content in comparison to the real environment.
[0022] The virtual content generated and verified by MR System 2 is model data that retains three-dimensional shape and texture, and can be used as design data for creating actual buildings and other structures.
[0023] Figure 3 is a functional block diagram showing an example of the logical configuration of the information processing device 1 in Embodiment 1. Note that some of the functional blocks shown in Figure 3 are realized by having the CPU, etc., which acts as a computer within the information processing device 1, execute computer programs stored in memory, which acts as a storage medium.
[0024] However, some or all of these can be implemented in hardware. Hardware options include dedicated circuits (ASICs) and processors (reconfigurable processors, DSPs).
[0025] Furthermore, the functional blocks shown in Figure 3 do not necessarily have to be housed in the same enclosure; they may be composed of separate devices connected to each other via signal paths. The above explanation regarding Figure 3 also applies to Figure 5.
[0026] The information processing device 1 comprises an environmental information acquisition unit 101, a generation unit 102, and an update unit 103. The environmental information acquisition unit 101 and the update unit 103 are connected to an external device 104 (for example, an MR system 2). The environmental information acquisition unit 101 functions as an environmental information acquisition means and acquires environmental information related to the environment held by the external device 104.
[0027] The generation unit 102 functions as a generation means and generates virtual content based on the environmental information acquired by the environmental information acquisition unit 101. In this embodiment, since virtual content is generated using a generation AI, the generation unit 102 holds a generation AI model for virtual content generation.
[0028] The update unit 103 receives and updates the virtual content generated by the generation unit 102. In this embodiment, since the virtual content is updated using a generation AI, the update unit 103 holds a generation AI model for updating the virtual content.
[0029] Furthermore, the update unit 103 functions as an update means that updates the virtual content to harmonize with the environment based on environmental information. The external device 104 transmits environmental information to the information processing device 1. It also retrieves the virtual content from the information processing device 1.
[0030] Figure 4 is a flowchart showing an example of an information processing method using the information processing device in Embodiment 1. The CPU and other components of the information processing device 1 execute a computer program stored in memory, which sequentially performs each step of the flowchart in Figure 4.
[0031] Furthermore, the processing steps in the flowchart described below are not limited to the example shown. Any combination of steps, grouping of multiple processes, or subdivision of processes is possible as long as the results of this embodiment are satisfied. In addition, each process can be individually separated and function as a single functional element, and can be used in combination with processes other than those shown.
[0032] When a user inputs a command to display virtual content via the input means (not shown) of the MR system 2, the processing flow shown in Figure 4 begins processing. Then, in step S1010, initialization processing is performed to make the information processing device 1 operational.
[0033] In this embodiment, virtual content is generated or updated using the generation AI in steps S1040 and S1050. Therefore, in step S1010, the generation unit 102 and the update unit 103 each perform the process of loading the structural data and weight parameters of the generation AI model into memory.
[0034] In step S1020, the environmental information acquisition unit 101 acquires environmental information from the external device 104. Here, the environmental information includes the location and category of objects present in the environment, and in this embodiment, the environmental information is information that is continuously updated in real time by the MR system 2, which is the external device. Step S1020 functions as an environmental information acquisition step to acquire environmental information about the environment.
[0035] In step S1030, the generation unit 102 determines whether or not the virtual content has already been generated. If the virtual content has not yet been generated, the process proceeds to step S1040. If the virtual content has already been generated, the process proceeds to step S1050.
[0036] In step S1040, the generation unit 102 generates virtual content using environmental information acquired by the environmental information acquisition unit 101 from the external device 104. In this embodiment, in order to generate virtual content, a method is used in which text information (hereinafter referred to as prompts) is input and a three-dimensional model is generated using the method described in Non-Patent Literature 1. Here, step S1040 functions as a generation step in which virtual content is generated based on environmental information.
[0037] In this step, the number of buildings included in the environmental information is first aggregated by category, and a prompt is generated instructing the system to create buildings that harmonize with the landscape, particularly those in the category with the largest number of buildings.
[0038] For example, if the category with the largest number of buildings is "Temples," the prompt will be set to the string "Generate a 3D model of buildings that harmonize with a landscape containing many temples."
[0039] Next, the created prompt is input to the generation AI model to generate a three-dimensional model, and the process proceeds to step S1060. Thus, in this embodiment, the generation means generates text information (prompts) for generating virtual content that harmonizes with the environment based on information about the environment, and uses the generation AI model to generate virtual content based on the text information (prompts).
[0040] In this context, the position and orientation of the building as virtual content within the real environment will be automatically estimated by MR System 2. However, the user of MR System 2 may place the virtual content in a desired position and orientation using the user interface while viewing the MR image, or the user may be allowed to adjust the position and orientation automatically estimated by MR System 2.
[0041] In step S1050, the update unit 103 updates the virtual content. Here, step S1050 functions as an update step that updates the virtual content to harmonize with the environment based on environmental information.
[0042] In step S1050, the degree to which the virtual content and the environment are in harmony (hereinafter referred to as the degree of harmony) is first calculated. Next, a prompt is generated based on the calculation result of the degree of harmony, and this prompt is input into the generating AI model to update the three-dimensional model that was generated in step S1040, etc.
[0043] Specifically, in step S1050, the update means calculates the degree of harmony of the virtual content with the environment based on the environmental information, and updates the virtual content to harmonize with the environment based on the degree of harmony. The update of the three-dimensional model may be performed using the method described in Non-Patent Document 1.
[0044] In this embodiment, the external device 104 (MR system 2) acquires virtual content from the update unit 103, combines the virtual content with an image of the real environment (hereinafter referred to as a real image), and generates an MR image.
[0045] Next, the update unit 103 acquires the real image and the MR image from the external device 104 (MR system 2). Furthermore, the update unit 103 calculates the difference in color distribution between the real image and the MR image using the method described in Non-Patent Document 2.
[0046] Finally, a prompt is generated to update the current virtual content based on the difference in color distribution. Specifically, the color family is divided into, for example, 11 categories (red, pink, orange, yellow, green, blue, purple, brown, white, gray, and black), and the difference in color distribution within each category is calculated.
[0047] Next, color families whose difference in color distribution exceeds a threshold are detected. Finally, for each detected color family, prompts are generated to update the virtual content so that the number of colors present more frequently in the real image increases, and the number of colors present more frequently in the MR image decreases.
[0048] For example, if the color families with a difference greater than or equal to a threshold are brown and pink, and brown is more abundant in the real image while pink is more abundant in the MR image, then a prompt would be created saying, "Increase the amount of brown and decrease the amount of pink." Next, the created prompt and virtual content are entered, the virtual content is updated, and the process proceeds to step S1060.
[0049] In step S1060, a decision is made as to whether or not to terminate the processing flow shown in Figure 4. If a termination instruction is received from the user via the input means (not shown) of the MR system 2, the virtual content generated in step S1040 or the virtual content updated in step S1050 is saved, and the processing flow shown in Figure 4 is terminated.
[0050] Otherwise, the process returns to step S1020 and the steps from S1020 to S1060 are repeated. According to the method of this embodiment, it is possible to generate a building design that harmonizes with the surrounding landscape (e.g., color distribution).
[0051] <Variation 1-1> In this embodiment, environmental information was input to the generation unit when generating virtual content, but in addition to this, generation conditions for the virtual content may also be input. Here, the generation conditions for the virtual content are information that determines the appearance (look) of the virtual content, and relate to the color, size, shape, position, orientation, etc. of the virtual content.
[0052] For example, the user specifies the size of the virtual content as a generation condition and inputs it into the generation or update unit. The generation or update unit then generates or updates the virtual content to the specified size. In this case, the size of the virtual content is numerically input as the lengths of the three sides of a rectangular prism that circumscribes the virtual content.
[0053] Alternatively, the size of the virtual content can be set by placing primitives such as cuboids and cylinders in the MR space using the user interface installed in MR system 2. Or, the size of the virtual content can be set by inputting two-dimensional regions from multiple positions and orientations using the method described in Non-Patent Document 7, thereby setting the viewing volume and defining the area in which the virtual content exists.
[0054] Alternatively, the environment information may be modified to include a two-dimensional image containing the appearance characteristics (at least one of the following: color, size, shape, position, orientation, etc.) of the virtual content to be generated. The generation unit may then extract a word or sentence that describes the appearance characteristics of that two-dimensional image and generate a prompt containing that word or sentence.
[0055] Furthermore, a method for extracting words or sentences that describe the visual characteristics (at least one of the following) of a two-dimensional image (such as color, size, shape, position, or orientation) can be used, for example, the method described in Non-Patent Document 4. As mentioned above, by specifying the conditions for generating virtual content, the appearance of the virtual content can be made closer to the user's wishes.
[0056] Thus, the generation means may generate virtual content based on generation conditions relating to at least one of the color, shape, position, and orientation of the virtual content.
[0057] <Variation 1-2> In this embodiment, when generating virtual content, environmental information acquired by the MR system 2 was input to the generation unit, but information regarding the category of the virtual content may also be used.
[0058] For example, a prompt relating to the category of content to be generated is stored in a storage unit (not shown) of the information processing device, and when the generation unit 102 generates a prompt, it generates a prompt that includes the aforementioned category.
[0059] If the retained category is "convenience store", the generation unit 102 generates a prompt such as "Generate a three-dimensional model of a convenience store that harmonizes with a landscape that has many temples", and generates virtual content based on this prompt.
[0060] As mentioned above, by using categories, virtual content can be generated using a prompt that allows the user to specify their desired category.
[0061] <Variation 1-3> In this embodiment, the environmental information acquired by the environmental information acquisition unit 101 was generated by a camera mounted on the external device 104, which is the MR system 2. However, the information may be generated by a device other than the MR system 2, as long as it includes information on the location and category of objects in the environment.
[0062] In other words, for example, information obtained from cameras and position / orientation sensors installed in information terminals such as smartphones and tablets may be used to recognize the position and category of objects around the smartphone or tablet, and the results may be used as environmental information.
[0063] Alternatively, the environmental information may be generated by recognizing the position and category of objects using image recognition technology from an image generated using a technology for generating arbitrary viewpoint images. As a method for generating arbitrary viewpoint images, for example, the methods described in Non-Patent Document 5 and Non-Patent Document 6 may be used.
[0064] Alternatively, environmental information can be obtained by rendering data obtained by three-dimensionally reconstructing the environment using three-dimensional reconstruction techniques such as photogrammetry, or by retrieving images from a database that stores a series of two-dimensional images along with their shooting position and orientation.
[0065] Furthermore, environmental information can be extracted from map data that includes information on the location and category of objects, such as map data for car navigation systems. In this way, environmental information is not limited to its generation method, and environmental information generated by various methods can be used.
[0066] <Variation 1-4> In this embodiment, the virtual content to be generated is described as a building, but this is not limited to this, and it may be another object.
[0067] For example, if the virtual content to be generated is furniture, in step S1040 shown in Figure 4, the prompt will be the string "Generate a three-dimensional model of furniture that will harmonize with a room with many chairs." Furthermore, in step S1050 shown in Figure 4, the prompt will be the string "Increase the amount of yellow."
[0068] Alternatively, if the virtual content to be generated is a tree, in step S1040 shown in Figure 4, the prompt will contain the string "Generate a three-dimensional model of a tree that harmonizes with a landscape full of buildings."
[0069] Furthermore, in step S1050 shown in Figure 4, the prompt may be the string "Please increase the number of trees with red autumn leaves." As mentioned above, virtual content can be used regardless of its category, and virtual content of various categories can be generated and updated.
[0070] <Variation 1-5> In this embodiment, the degree of harmony was calculated using the color distribution of the real environment and virtual content in step S1050 shown in Figure 4, but this is not limited to this. The degree of harmony may also be calculated using at least one element of the real environment information and virtual content, such as color, shape, position, and orientation.
[0071] In other words, the update means may calculate a degree of harmony based on environmental information and at least one of the color, shape, position, and orientation of the virtual content, and update the virtual content to increase the degree of harmony.
[0072] As described in Embodiment 1, one method using color involves using the difference in color distribution between the real image and the MR image, and the degree of harmony is calculated such that the smaller the difference in color distribution, the higher the degree of harmony.
[0073] Alternatively, the degree of harmony may be calculated such that the closer the number of pixels in the MR image whose saturation is within a certain range is to a predetermined number, the higher the degree of harmony. Furthermore, the degree of harmony may be calculated using shape. In that case, for example, the degree of harmony may be calculated based on the difference in height between the virtual content and objects (such as buildings) in the real environment adjacent to the virtual content.
[0074] If the virtual content is a house, the harmony score could be calculated such that the closer the shape and color of the roof of the virtual content are to the real object, the higher the harmony score, from the perspective of the streetscape. Alternatively, if the virtual content is furniture, the harmony score could be calculated such that the smaller the difference in height between the real object and the virtual content, the higher the harmony score.
[0075] One method using location is to utilize the spatial distribution of objects and virtual content existing in the real environment. If the virtual content is a building, the degree of harmony may be calculated such that the greater the number of real buildings within a certain distance from the virtual content, the higher the degree of harmony.
[0076] Alternatively, if the virtual content is furniture, the degree of harmony could be calculated such that the closer the virtual furniture is to real-world furniture or building materials, the higher the degree of harmony.
[0077] One method using posture is to use the difference in posture between objects in the real environment and virtual content. If the virtual content is a building, the degree of harmony can be calculated such that the smaller the difference between the direction of the entrance to the virtual building and the direction of the entrance to a building in the real environment adjacent to the virtual content, the higher the degree of harmony.
[0078] If the virtual content is a furniture shelf, the harmony score may be calculated such that the smaller the difference in orientation between the virtual shelf and the adjacent shelves in the real environment, the higher the harmony score. Alternatively, using a different method that utilizes posture, the harmony score may be calculated such that the harmony score is higher based on the percentage of the virtual content that is visible when the virtual content is placed (hereinafter referred to as the visibility percentage).
[0079] If the virtual content is a building, the harmony score may be calculated such that the higher the visibility of the building's entrance, the higher the harmony score. Alternatively, if the virtual content is furniture, the harmony score may be calculated such that the lower the visibility of the furniture's back, the higher the harmony score.
[0080] Furthermore, the harmony score may be calculated by combining two or more elements from each of the harmony scores described above. For example, the harmony score may be calculated as a weighted sum of the harmony scores calculated using color and the harmony scores calculated using shape.
[0081] Alternatively, the harmony score can be calculated as a weighted sum of the harmony score calculated using position and the harmony score calculated using orientation. As mentioned above, by calculating the harmony score, it is possible to calculate the harmony score using the elements desired by the user and update the virtual content.
[0082] <Variation 1-6> In this embodiment, the categories of objects in the environmental information may be subdivided based on the visual characteristics of the objects. For example, "temple" may be subdivided into five types based on architectural style: "Japanese style," "eclectic style," "Great Buddha style," "Zen style," and "other."
[0083] Furthermore, furniture can be subdivided into four categories: "Renaissance style," "Baroque style," "Rococo style," and "other." Alternatively, it can be subdivided using color, size, etc.
[0084] The subdivided categories are obtained by recognizing objects from surrounding images as environmental information. For example, using the method described in Non-Patent Document 8, a deep learning model is created that has been trained on a dataset corresponding to the subdivided categories, and then objects are recognized from the surrounding images.
[0085] As mentioned above, by subdividing object categories, it becomes possible to generate and update virtual content that harmonizes with the environment at a finer granularity.
[0086] <Variation 1-7> The generation of virtual content may be carried out using an initial three-dimensional model, such as a three-dimensional model that visually resembles the new virtual content to be generated. The initial three-dimensional model may be input to the generation unit 102 by the user using a condition input means (not shown) of the MR system 2. The generation unit 102 then generates virtual content using the above three-dimensional model as input.
[0087] The initial three-dimensional model is one that has been manually created using modeling tools or similar methods. Alternatively, it can be one that has been automatically generated by inputting a two-dimensional image. As a method for generating a three-dimensional model from a two-dimensional image, for example, the method described in Non-Patent Document 3 can be used.
[0088] To generate another different three-dimensional model using the input three-dimensional model as an initial value, for example, the method described in Non-Patent Document 1 can be used. As mentioned above, by using the initial three-dimensional model, it is possible to generate virtual content that is close to the user's desired appearance and harmonizes with the environment.
[0089] <Variation 1-8> In this embodiment, in step S1050 shown in Figure 4, the degree of harmony was calculated using the color distribution of the real environment and the virtual content. However, for example, when virtual content is installed, the degree of harmony may also be calculated based on the amount by which the image of a specific object is obscured by the image of the virtual content (hereinafter referred to as the occlusion amount).
[0090] For example, harmony could be calculated such that the lower the occlusion of a famous castle located outdoors or a painting displayed indoors, the higher the harmony. Alternatively, harmony could be calculated such that the higher the occlusion of equipment on a building's rooftop, the higher the harmony.
[0091] The objects to be included in the occlusion calculation are pre-selected by the user from among the objects included in the environmental information. Alternatively, the two-dimensional region is input from multiple positions and orientations using the method described in Non-Patent Document 7 to set the viewing volume and define the region where the objects to be included in the occlusion calculation are located.
[0092] As mentioned above, by using the occlusion amount to calculate the degree of harmony, virtual content can be generated or updated so that objects that you want to stand out are more prominent. Conversely, virtual content can be generated or updated so that objects that you do not want to stand out are less prominent.
[0093] <Variation 1-9> In Embodiment 1, the range for generating environmental information was the entire environment, but this is not limited to this; it may be limited to a partial region. For example, the entirety of a two-dimensional image of the environment, a point cloud with three-dimensional coordinates generated by environmental measurement, or a three-dimensional mesh model may be targeted.
[0094] Alternatively, a user interface that allows setting a region may be used to specify a subregion of the above-mentioned two-dimensional image, point cloud, or mesh model, and environmental information may be generated targeting objects within the specified range.
[0095] As mentioned above, by generating environmental information, it is possible to create and update virtual content that harmonizes with the user's intended environment.
[0096] <Variation 1-10> In Embodiment 1, the virtual content was updated using a generation AI, but the invention is not limited to this. For example, the parameters of the virtual content's color, shape, position, and orientation may be directly updated based on the results of a harmonicity calculation.
[0097] If calculating the degree of harmony indicates that it would be beneficial to increase the brown and decrease the pink in the virtual content, then some of the non-brown points in the virtual content are selected and changed to brown. Alternatively, some of the pink points are selected and changed to a color other than pink. Or, some of the faces in the virtual content are selected and changed to brown.
[0098] Furthermore, if calculating the degree of harmony indicates that reducing the height of the virtual content would be beneficial, the virtual content will be resized to reduce its height.
[0099] Furthermore, if calculating the degree of harmony indicates that changing the position or orientation of the virtual content would be beneficial, the virtual content will be changed to a different position or orientation. The alternative position or orientation will be one that maximizes the degree of harmony within a certain range of the real-world environment.
[0100] In this case, for position, you can pre-set several candidate positions and select the one that provides the highest degree of harmony. For posture, you can rotate the virtual content horizontally in 90-degree increments and select the posture that provides the highest degree of harmony.
[0101] As described above, by updating the virtual content, it is possible to update it using elements desired by the user, such as the color, shape, position, and orientation of the virtual content, without using a generation AI.
[0102] <Embodiment 2> In Embodiment 1, the generation unit 102 generated virtual content using a generation AI. In this embodiment, one virtual content model is selected and output from a plurality of three-dimensional models of virtual content having pre-generated categories (hereinafter referred to as existing virtual content models) using an object placement characteristics database, which will be described later. As an example of a building, similar to Embodiment 1, an example of a virtual content model to be output will be described.
[0103] Figure 5 is a functional block diagram showing an example of the logical configuration of the information processing device 200 in Embodiment 2. The input and output data of the environmental information acquisition unit 101, the update unit 103, and the external device 104 are the same as in Embodiment 1, so their explanation is omitted here. The holding unit 105 and the generation unit 102 differ from those in Embodiment 1, so the differences from Embodiment 1 will be explained.
[0104] The holding unit 105 holds an object placement characteristics database and multiple existing virtual content models. Here, there is one existing virtual content model for each of the multiple categories.
[0105] Furthermore, the storage unit stores multiple virtual content models, associating them with object categories and location information. Based on environmental information, the generation unit 102 selects and outputs one virtual content from among the multiple existing virtual content models held by the storage unit 105.
[0106] The object placement characteristics database in this embodiment is a database generated by learning the object placement characteristics, consisting of the category and three-dimensional position of objects present in the environment, through deep learning. The construction and use of the object placement characteristics database are carried out using the method described in Patent Document 3. The method for constructing and using the object placement characteristics database is described below.
[0107] The object placement characteristics database in this embodiment is a pre-trained neural network that has been trained to infer unknown object placement characteristics from the surrounding object placement characteristics. Specifically, it is a pre-trained neural network consisting of 24 layers of Ashish et al.'s Transformer ("Attention is All you Need", Ashish.et.el NeuralIPS2017).
[0108] In this embodiment, the Transformer has an input dimension of 512 and an output dimension of 512, meaning that it takes up to 512 pieces of object property information as input and outputs the same number of 512 dimensions.
[0109] As the Transformer, we will use an encoder network, such as the one used in Jacob et al.'s method (Jacob et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv 2018).
[0110] The learning process uses accurate information (hereinafter referred to as "true value information") regarding the categories and three-dimensional positions of objects present in the environment. In this embodiment, the true value information used is information extracted from map data for car navigation systems. Upon completion of this learning process, it becomes possible to infer unknown object placement characteristics based on the object placement characteristics of the environment from which the true value information was obtained.
[0111] The purpose of using the object placement characteristics database is to obtain the most likely category of the virtual content. The environmental information acquired by the environmental information acquisition unit 101 in Figure 5, combined with the virtual content information, is input into the object placement characteristics database to estimate the most likely category of the virtual content. In this embodiment, the virtual content information used is information whose location is known but whose category is unknown.
[0112] Figure 6 is a flowchart showing an example of an information processing method using the information processing device 200 in Embodiment 2. The CPU and other components of the information processing device 1 execute computer programs stored in memory, sequentially performing the operations of each step in the flowchart in Figure 6.
[0113] The flowchart in Figure 6 is similar in structure to the flowchart in Figure 4 of Embodiment 1, but only the processing in step S2040 differs from Embodiment 1; therefore, the processing other than step S2040 will not be explained.
[0114] In step S2040, the generation unit 102 uses the environmental information acquired in step S1020 and the object placement characteristics database held by the holding unit 105 to select and output one virtual content model from among multiple existing virtual content models.
[0115] Figure 7 is a flowchart showing a detailed example of the process in step S2040 of Figure 6. Using Figure 7, an example of how to select a virtual content model using an object placement characteristics database will be explained.
[0116] Furthermore, the CPU and other components of the information processing device 1, acting as a computer, execute the computer program stored in memory, thereby sequentially performing the operations of each step in the flowchart shown in Figure 7.
[0117] In Figure 7, step S7010 calculates the placement of the virtual content. Specifically, first, images of the environment captured using a camera mounted on the head-mounted display of the MR system 2 are processed using semantic segmentation to detect empty areas on the two-dimensional image.
[0118] Next, in step S7020, the average of the three-dimensional positions of the point cloud included in the vacant lot area is used as the coordinates of the vacant lot. The three-dimensional positions of the point cloud included in the vacant lot area can be calculated using, for example, the method described in the literature by Raul et al. (Raul Mur-Artal et.al, ORB-SLAM: A Versatile and Accurate Monocular SLAM System. IEEE Transactions on Robotics).
[0119] Next, in step S7030, the environmental information acquisition unit 101 uses the acquired environmental information to create an object type vector and a position vector. The object type vector is a one-dimensional column vector in which the labels of the categories of objects present in the environment are arranged.
[0120] For example, the first element of the object type vector is the CLS token (a special label that indicates the beginning of the data), followed by the label of the category of the first object included in the environmental information, then the label of the category of the second object, and so on, with the labels of the object categories listed.
[0121] Furthermore, in step S7030, a MASK token (a special label indicating that the object type is unknown) is placed at the end of the object type vector as a label for the virtual content. The position vector is a vector formed by arranging three elements, each consisting of the three-dimensional positions X, Y, Z of each object included in the environment information and the three-dimensional positions X, Y, Z of the virtual content, as column vectors, and then arranging these three elements as a one-dimensional column vector.
[0122] In other words, each column of the position vector stores the X,Y,Z values, which are the position coordinates of the object in the corresponding column included in the object type vector, and the last column stores the three-dimensional position X,Y,Z values of the virtual content. Here, the three-dimensional position X,Y,Z values of the virtual content are the three-dimensional position of the detected vacant lot.
[0123] Next, in step S7040, the object type vector and position vector created in step S7030 are input into the object placement characteristics database, and the object type (category) is predicted by the prediction process in the database. Then, in step S7050, a new object type vector is created in which the MASK token is updated with the predicted object type label.
[0124] Finally, in step S7060, the holding unit 105 selects an existing virtual content model from the existing virtual content model group it holds that corresponds to the object type label at the end of the new object type vector. This allows for the automatic selection of a virtual content model for a plausible category based on pre-trained data. After step S7060, the processing flow shown in Figure 7 is terminated.
[0125] <Variation 2-1> In Embodiment 2, the virtual content's location was known, but its category was unknown. The category was then inferred by inputting this information into the object placement characteristics database. However, this is not limited to this approach. Alternatively, the virtual content's location may be unknown, but its category may be known.
[0126] This allows virtual content to be placed in a position that harmonizes with the environment, based on pre-learned data.
[0127] <Modification 2-2> In Embodiment 2, the location of the virtual content was known, but the category was unknown. This information was entered into the object placement characteristics database, and the category was inferred. However, it is also possible to input information where the location and category of the virtual content are unknown into the object placement characteristics database and infer the location and category.
[0128] This allows for the placement of virtual content in a location that harmonizes with the environment, and the selection of categories that harmonize with the environment, based on pre-learned data.
[0129] <Modified example 2-3> The object placement characteristics database in Embodiment 2 was a database that learned object placement characteristics consisting of the category and three-dimensional position of objects present in the environment. However, a database that simultaneously learns elements related to the appearance of virtual content other than category and three-dimensional position may also be used.
[0130] For example, one could create and use an object placement characteristics database that includes the color and size of virtual content as elements. By using such a database, it becomes possible to infer other unknown information by using the elements related to the appearance of virtual content as known information.
[0131] Furthermore, the category and 3D position of the virtual content, or either of them, may be assumed as known information, while the elements related to the appearance of the virtual content may be assumed as unknown information. Moreover, the category, 3D position, and all elements related to appearance may be assumed as unknown information.
[0132] As mentioned above, by creating and utilizing an object placement characteristics database, it is possible to select and place virtual content that harmonizes with the environment based on elements related to category, three-dimensional position, and the appearance of the virtual content.
[0133] <Modification 2-4> In this embodiment, there was one existing model for each category, but this is not limited to this, and multiple existing models may be used for each category. If there are multiple existing models for a category corresponding to the prediction results for that category obtained using the object placement characteristics database, one may be randomly selected from among them, for example.
[0134] Alternatively, the user interface of MR System 2 could display multiple existing models, allowing the user to select one. As mentioned above, using multiple existing models from each category allows for the selection of a more extensive set of existing models to be used as virtual content.
[0135] <Embodiment 3> In Embodiment 1, the virtual content was updated based on the degree of harmony between the real environment and the virtual content in order to update the virtual content so that it would harmonize with the environment, but calculating the degree of harmony is not essential.
[0136] You can pre-register and use prompts to update virtual content to harmonize with the environment. For example, to harmonize colors, you could use a string like "Make the colors closer to those of adjacent objects."
[0137] Alternatively, to harmonize the heights, you could use the phrase, "Make the heights similar to adjacent objects." Or, if the virtual content is a building, you could use the phrase, "Make it the same architectural style as the surrounding buildings." Using this method, you can update the virtual content to harmonize with the environment, similar to how you would calculate the degree of harmony.
[0138] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible in accordance with the spirit of the present invention, and these are not excluded from the scope of the present invention. Furthermore, some of the above embodiments may be combined as appropriate.
[0139] Furthermore, the present invention includes, for example, a system that implements the functions of the above embodiment using at least one processor such as a CPU, memory, and circuitry (e.g., an ASIC). Alternatively, multiple processors may be used for distributed processing.
[0140] Furthermore, in order to implement some or all of the control in the above embodiment, a computer program that implements the functions of the above embodiment may be supplied to an information processing device, etc., via a network or various storage media.
[0141] Furthermore, the computer (or CPU, MPU, etc.) in the information processing device may read and execute the program. In that case, the program and the storage medium in which the program is stored constitute the present invention. The present invention includes the following combinations.
[0142] (Configuration 1) An information processing apparatus characterized by comprising: an environmental information acquisition means for acquiring environmental information relating to the environment; a generation means for generating virtual content based on the environmental information; and an update means for updating the virtual content to harmonize with the environment based on the environmental information.
[0143] (Configuration 2) The information processing apparatus according to Configuration 1, wherein the update means calculates the degree of harmony of the virtual content with the environment based on the environment information, and updates the virtual content to harmonize with the environment based on the degree of harmony.
[0144] (Configuration 3) The information processing apparatus according to Configuration 2, wherein the updating means calculates the degree of harmony based on at least one of the environmental information and the color, shape, position, and orientation of the virtual content, and updates the virtual content so that the degree of harmony becomes higher.
[0145] (Configuration 4) An information processing device according to any one of Configurations 1 to 3, characterized in that the environmental information includes the location and category of objects present in the environment.
[0146] (Configuration 5) An information processing apparatus according to any one of Configurations 1 to 4, wherein the generation means generates the virtual content based on generation conditions relating to at least one of the color, shape, position, and orientation of the virtual content.
[0147] (Configuration 6) An information processing apparatus according to any one of Configurations 1 to 5, characterized in that the generation means generates text information for generating the virtual content that harmonizes with the environment based on information about the environment, and generates the virtual content based on the text information using a generation AI model.
[0148] (Configuration 7) An information processing apparatus according to any one of Configurations 1 to 6, comprising a holding unit that holds a plurality of virtual content models associated with the category of an object and location information, wherein the generation means selects the virtual content from the plurality of virtual content models based on the environmental information.
[0149] (Method) An information processing method characterized by comprising: an environmental information acquisition step of acquiring environmental information relating to the environment; a generation step of generating virtual content based on the environmental information; and an update step of updating the virtual content to harmonize with the environment based on the environmental information.
[0150] A computer program for controlling each of the means of the information processing device described in any one of configurations 1 to 7 by a computer. [Explanation of Symbols]
[0151] 1, 200: Information Processing Device 2: MR System 3, 4: Buildings 5: Virtual Content
Claims
1. An environmental information acquisition means for acquiring environmental information related to the environment, A generation means for generating virtual content based on the aforementioned environmental information, An update means for updating the virtual content to harmonize with the environment based on the aforementioned environmental information, An information processing device characterized by having the following features.
2. The update means calculates the degree of harmony of the virtual content with the environment based on the environment information, and updates the virtual content to harmonize with the environment based on the degree of harmony. The information processing apparatus according to claim 1, characterized by the following:
3. The update means calculates the degree of harmony based on the environmental information and at least one of the color, shape, position, and orientation of the virtual content. To update the virtual content so that the degree of harmony is higher, The information processing apparatus according to claim 2, characterized in that
4. The aforementioned environmental information includes the location and category of objects present in the environment. The information processing apparatus according to claim 1, characterized by the following:
5. The generation means generates the virtual content based on generation conditions relating to at least one of the color, shape, position, and orientation of the virtual content. The information processing apparatus according to claim 1, characterized by the following:
6. The generation means generates text information for generating virtual content that harmonizes with the environment based on information about the environment, and generates the virtual content based on the text information using a generation AI model. The information processing apparatus according to claim 1, characterized by the following:
7. It has a holding unit that holds multiple virtual content models, associating them with object categories and location information. The generation means selects the virtual content from among the plurality of virtual content models based on the environmental information. The information processing apparatus according to claim 1, characterized by the following:
8. An environmental information acquisition step to obtain environmental information about the environment, A generation step of generating virtual content based on the aforementioned environmental information, An update step of updating the virtual content to harmonize with the environment based on the aforementioned environmental information, An information processing method characterized by having the following features.
9. A computer program for controlling each means of the information processing apparatus described in any one of claims 1 to 7 by a computer.