Method for generating space map including actual image by using image obtained by photographing space, and electronic device for performing same
By integrating actual image data with lidar-based initial maps using a generative model, the method addresses the limitations of lidar-generated spatial maps, resulting in more accurate and realistic representations of spaces, thereby enhancing user experience.
Patent Information
- Application Number
- PCT/KR2024/015112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-10-04
- Publication Date
- 2025-05-30
AI Technical Summary
Spatial maps generated using lidar scan data lack accuracy and realism, failing to accurately represent the structure and appearance of actual spaces due to the inability to distinguish between walls and objects, leading to distorted borderlines and limited user experience.
A method that involves obtaining an initial map from lidar scan data, capturing images of areas of interest, registering actual images with the initial map based on depth information, determining areas requiring image generation, and using a generative model to create images for these areas, thereby generating a more accurate and realistic spatial map.
The proposed method enhances the accuracy and realism of spatial maps by integrating actual image data with lidar-based initial maps, improving user experience by providing a more faithful representation of physical spaces.
Smart Images

Figure KR2024015112_30052025_PF_FP_ABST
Abstract
Description
Method for generating a spatial map composed of real-world images using images taken of a space, and an electronic device for performing the same
[0001] The present disclosure relates to a method for generating a spatial map, and more particularly, to a method for generating a spatial map composed of real-world images using images taken of a space.
[0002] Electronic devices such as robot vacuum cleaners can scan a space using built-in LiDAR sensors as they move through it, and create a map of the space (spatial map) based on the scan data.
[0003] However, when a spatial map is created using only data (LiDAR scan data) scanned using a LiDAR sensor, the created spatial map may differ from the structure of the actual space. This is because various objects (e.g., furniture, home appliances, etc.) may exist in the space, and electronic devices cannot determine whether the point where the signal (light) transmitted from the LiDAR sensor is reflected is a wall of the space or an object existing within the space. The electronic device may actually recognize the point where the object exists as a wall of the space and create a spatial map, and therefore the boundaries of the spatial map may differ from the actual space.
[0004] In addition, although the spatial map generated using lidar scan data expresses the overall structure of the space and the division of areas included in the space, it is different from the actual appearance of the space, and thus has limitations in providing a sense of reality to the user viewing it.
[0005] A method for generating a spatial map according to one embodiment of the present disclosure may include the steps of: obtaining an initial map for a target space; obtaining a captured image of one or more areas of interest in the target space; registering an actual image obtained from the captured image to the initial map based on depth information included in the captured image; determining a generation required area requiring image generation in the initial map to which the actual image is registered; and generating an image for the generation required area using a generative model to generate an actual map.
[0006] An electronic device for generating a spatial map according to one embodiment of the present disclosure includes a memory storing a program for generating a spatial map and at least one processor, wherein the at least one processor executes the program to obtain an initial map for a target space, obtain a captured image of one or more areas of interest in the target space, register an actual image obtained from the captured image to the initial map based on depth information included in the captured image, determine a generation required area requiring image generation in the initial map to which the actual image is registered, and then generate an image for the generation required area using a generative model, thereby generating an actual map.
[0007] A computer-readable recording medium according to one embodiment of the present disclosure may have stored thereon a program for executing at least one of the embodiments of the disclosed method on a computer.
[0008] A computer program according to one embodiment of the present disclosure may be stored in a medium for performing at least one of the embodiments of the disclosed method on a computer.
[0009] FIG. 1 is a diagram illustrating a system for generating a spatial map according to one embodiment of the present disclosure.
[0010] FIG. 2 is a diagram illustrating an initial map generated based on lidar scan data and a real-world map generated using a photographed image according to one embodiment of the present disclosure.
[0011] FIG. 3 is a diagram for explaining the configuration of a server included in a system for generating a spatial map according to one embodiment of the present disclosure.
[0012] FIG. 4 is a drawing for explaining the configuration of a robot cleaner included in a system for generating a spatial map according to one embodiment of the present disclosure.
[0013] FIG. 5 is a diagram illustrating UI screens displayed on a mobile terminal during the process of generating a spatial map according to one embodiment of the present disclosure.
[0014] FIG. 6 is a diagram for explaining a method for recommending to a user an area of interest where an electronic device (server) should perform a photographing, and a photographing location and a photographing direction according to one embodiment of the present disclosure.
[0015] FIG. 7 is a diagram illustrating a method for an electronic device (server) according to one embodiment of the present disclosure to generate an estimated map representing the layout of an area of interest from a captured image of the area of interest.
[0016] FIG. 8 is a diagram illustrating a method for an electronic device (server) according to one embodiment of the present disclosure to match an estimated map to an initial map based on a boundary line extracted from the estimated map.
[0017] FIG. 9 is a diagram illustrating a process in which an electronic device (server) according to one embodiment of the present disclosure generates a real-world map from an initial map in which real-world images are aligned.
[0018] FIGS. 10 to 18 are flowcharts for explaining a method for generating a spatial map according to embodiments of the present disclosure.
[0019] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.
[0020] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.
[0021] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. The disclosed embodiments are provided to ensure that the disclosure of the present disclosure is complete and to fully inform those skilled in the art of the present disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. In addition, when describing an embodiment of the present disclosure, if a detailed description of a related function or configuration is determined to unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.
[0022] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.
[0023] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0024] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.
[0025] Below, the meanings of terms used in this disclosure are explained.
[0026] A "spatial map" may refer to a map representing the structure or shape of a space. According to one embodiment of the present disclosure, a spatial map representing the structure of a house in which a user resides may be generated using an electronic device such as a robot vacuum cleaner or a server. The "spatial map" may be in the form of a general floor plan representing the structure of a house, etc., with various information (e.g., appliances or furniture placed within the space) added. Terms such as "floor plan" or "spatial layout" may be used instead of "spatial map."
[0027] 'LiDAR scan data' may refer to data acquired by scanning a target space using a LiDAR sensor equipped in an electronic device. According to one embodiment of the present disclosure, the electronic device may acquire LiDAR scan data including depth values (distances from the LiDAR sensor to the plurality of points) of a plurality of spots located at a predetermined specific height using the LiDAR sensor. According to one embodiment of the present disclosure, the electronic device may acquire ToF scan data including depth values of a plurality of points located at a specific height using a ToF sensor instead of the LiDAR sensor, and may perform embodiments of the present disclosure using the ToF scan data instead of the LiDAR scan data. Alternatively, according to one embodiment of the present disclosure, the electronic device may acquire stereo vision scan data including depth values of a plurality of points located at a specific height using stereo vision technology using two cameras instead of the LiDAR sensor, and may perform embodiments of the present disclosure using the stereo vision scan data instead of the LiDAR scan data. Instead of 'lidar scan data', the terms '1D scan data', 'lidar data', 'depth scan data' or 'scan data' may also be used.
[0028] The term "initial map" may refer to a spatial map generated based on LiDAR scan data, and the term "actual map" may refer to a spatial map composed of actual images. The actual map may be composed of actual images that express the shape and color of an actual space, and the shape and color of objects (e.g., home appliances, furniture, etc.) placed in the space. According to one embodiment of the present disclosure, an electronic device may generate an actual map by attaching an actual image acquired from a photographed image to an initial map and generating an image in a blank area using a generation model. Instead of the term "initial map," terms such as "initial space map," "LiDAR map," and "LiDAR spatial map" may also be used. Additionally, instead of 'actual map', terms such as 'actual image map', 'actual floor plan', 'actual spatial map', 'user interface space map' or 'realistic map' may be used.
[0029] An "area of interest" may refer to an area that must be photographed to convert an initial map into a real-world map. An electronic device according to one embodiment of the present disclosure may generate a real-world map from the initial map using images of captured areas of interest. Furthermore, the electronic device according to one embodiment of the present disclosure may recommend areas of interest to a user. Terms such as "designated area" or "recommended area" may be used instead of "area of interest."
[0030] The term "captured image" may refer to an image that captures an area of interest. In one embodiment, the captured image may be an RGB image, and may be converted into a depth image through depth estimation as described in the embodiments of the present disclosure.
[0031] The term "actual image" may refer to an image that represents the actual appearance of a space. An actual image may represent the shape and color of an actual space, as well as the shape and color of objects (e.g., home appliances, furniture, etc.) placed in the space. According to one embodiment of the present disclosure, an electronic device may acquire an actual image (at a different point in time from the captured image) from a photographed image. Terms such as "realistic image" may also be used instead of "actual image."
[0032] A "depth image" can refer to an image containing depth information for a specific space. Multiple points within a depth image can each have corresponding depth values. As previously explained, a depth image can be obtained by performing depth estimation on a captured image. Terms such as "depth map" or "depth estimation image" may also be used instead of "depth image."
[0033] A 'depth value' can mean the distance from a reference location for measurement (e.g., the location of an electronic device performing a photograph or scan) to a specific point in space.
[0034] "Registration" may refer to a processing technique for representing different images in a single coordinate system. In the embodiments of the present disclosure, "registering a real-world image to an initial map" may refer to an operation of estimating a location corresponding to a photographed image in the initial map and attaching a real-world image obtained from the photographed image to the estimated location. Terms such as "matching" or "attaching" may also be used instead of "registration."
[0035] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0036] Embodiments of the present disclosure relate to a method for generating a real-world map representing an actual space using images captured of the space from an initial map generated based on data scanned by a LiDAR sensor, and an electronic device for performing the same.
[0037] Electronic devices implementing the embodiments described in this disclosure may be implemented in various forms. According to one embodiment, the electronic device may be a robotic mobile device, such as a robot vacuum cleaner, or a server connected to a robotic mobile device. The roles and operations performed by the electronic devices (robot vacuum cleaner, server, mobile terminal, etc.) included in the system according to the embodiments of the present disclosure are described in detail below with reference to the drawings.
[0038] FIG. 1 is a diagram illustrating a system for generating a spatial map according to one embodiment of the present disclosure. Referring to FIG. 1, the system for generating a spatial map according to one embodiment of the present disclosure may include a server (100), a robot vacuum cleaner (200), and a mobile terminal (300).
[0039] As described above, the spatial map according to the embodiments of the present disclosure may include an initial map and a real map, and the initial map and the real map will be described separately below.
[0040] A system according to one embodiment of the present disclosure can create an initial map using lidar scan data (1D scan data) acquired by a robot cleaner (200), and can create a real-world map from the initial map using a photographed image captured using a mobile terminal (300).
[0041] The operations of generating an initial map and generating a real-world map may be performed by any one of the devices included in the system (server (100), robot cleaner (200), mobile terminal (300)), or by a combination of two or more. For example, if the robot cleaner (200) transmits lidar scan data to the server (100) and the mobile terminal (300) transmits a photographed image to the server (100), the server (100) may generate both the initial map and the real-world map. Alternatively, for example, if the robot cleaner (200) generates an initial map based on lidar scan data and then transmits the generated initial map to the server (100), and the mobile terminal (300) transmits a photographed image to the server (100), the server (100) may also generate a real-world map from the received initial map.
[0042] In this disclosure, embodiments are described assuming that the server (100) generates both an initial map and a real-world map. While the embodiments of this disclosure are described as the server (100) performing all processes for generating the real-world map (10), some or all of the processes may be performed on another device (e.g., a robot vacuum cleaner (200) or a mobile terminal (300)).
[0043] The robot cleaner (200) can acquire LiDAR scan data by measuring the depth of multiple points located at specific heights in a target space using a LiDAR sensor (250). According to one embodiment of the present disclosure, other types of electronic devices (e.g., robotic mobile devices - devices that move automatically or according to user commands and perform various operations) other than the robot cleaner (200) may be used. For example, electronic devices such as a butler robot or a pet robot may be used.
[0044] When the robot cleaner (200) transmits the acquired LiDAR scan data to the server (100), the server (100) can generate an initial map based on the LiDAR scan data. FIG. 2 illustrates an initial map (20) generated based on the LiDAR scan data. Referring to FIG. 2, the initial map (20) shows the layout of the target space as seen from the top view, and displays multiple areas (living room, bedroom) separately.
[0045] The initial map (20) only expresses the structure of the target space (boundary and division of areas), whereas the real-world map (10) expresses the actual appearance of the target space (such as the shape and color of the target space and appliances / furniture placed in the target space). In addition, a comparison of the initial map (20) and the real-world map (10) reveals that the border of the initial map (20) does not accurately reflect the structure of the actual space.
[0046] In this way, if the server (100) generates the initial map (20) using only the lidar scan data, distortion of the space may occur. Various objects (e.g., furniture, home appliances, etc.) may exist in the space, and the server (100) cannot distinguish whether the point where the laser pulse is reflected is a wall of the space or an object using only the lidar scan data, and thus recognizes the object as a wall of the space and generates the initial map (20). In addition, according to one embodiment of the present disclosure, the robot cleaner (200) performs a lidar scan at a low height to acquire lidar scan data, and there is a high possibility that an object exists at the height at which the lidar scan is performed (e.g., a height close to the floor). Therefore, the initial map (20) generated based only on the lidar scan data may differ from the actual structure (floor plan) of the space.
[0047] When a mobile terminal (300) transmits images of multiple areas (areas of interest) within a target space to a server (100), the server (100) can generate a real-world map (10) from an initial map (20) using the received captured images. In this process, a user must capture an area of interest using a mobile terminal (300), and the server (100) can recommend an area of interest, a capturing location, and a capturing direction to the user through the mobile terminal (300). When the server (100) transmits the real-world map (10) to the mobile terminal (300), the user can check the real-world map (10) through the screen of the mobile terminal (300).
[0048] According to one embodiment of the present disclosure, the server (100) may generate a real-world map (10) using images taken by a robot cleaner (200) of multiple areas (areas of interest) within a target space. To this end, the server (100) may transmit information about the location of the area of interest, the location and direction in which the robot cleaner (200) should perform the shooting, etc. to the robot cleaner (200).
[0049] According to one embodiment of the present disclosure, an application for using and controlling a robot cleaner (200) may be installed on a mobile terminal (300). The server (100) may guide the user through the process of creating a real-world map (10) through the application, or request actions (e.g., capturing an area of interest) necessary for creating the real-world map (10). When the application is executed, the user can check the real-world map (10) through a UI screen displayed on the mobile terminal (300).
[0050] Below, with reference to drawings, a detailed description will be given of a specific process for generating a real-world map (10) from an initial map (20) using a photographed image received by a server (100) from a mobile terminal (300).
[0051] First, the components included in the server (100) and the robot cleaner (200) will be described with reference to FIGS. 3 and 4.
[0052] FIG. 3 is a diagram illustrating the configuration of a server included in a system for generating a spatial map according to one embodiment of the present disclosure. Referring to FIG. 3, a server (100) according to one embodiment of the present disclosure may include a communication interface (110), a processor (120), and a memory (130). However, the components of the server (100) are not limited to the above-described examples, and the server (100) may include more or fewer components than the above-described components. Some or all of the communication interface (110), the processor (120), and the memory (130) may be implemented in the form of a single chip. According to one embodiment of the present disclosure, the server (100) may be a home IoT server that controls IoT devices in the home, such as a robot vacuum cleaner (200).
[0053] The communication interface (110) is a configuration for transmitting and receiving signals (control commands, data, etc.) with an external device via wire or wirelessly, and may be implemented to include a communication chipset that supports various communication protocols. The communication interface (110) may receive a signal from the outside and output it to the processor (120), or transmit a signal output from the processor (120) to the outside. The server (100) may communicate with the robot cleaner (200) or the mobile terminal (300) via the communication interface (110).
[0054] The processor (120) controls a series of processes to enable the server (100) to operate according to the embodiments described below, and may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU or a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. For example, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0055] The processor (120) can record data in the memory (130) or read data stored in the memory (130), and in particular, process data according to predefined operating rules or artificial intelligence models by executing a program stored in the memory (130). Accordingly, the processor (120) can perform the operations described in the following embodiments. Unless otherwise specified, operations described as being performed by the server (100) in the following embodiments can be considered to be performed by the processor (120).
[0056] The memory (130) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (130) may not exist separately and may be configured to be included in the processor (120). The memory (130) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program for performing operations according to embodiments described below may be stored in the memory (130). The memory (130) may also provide stored data to the processor (120) at the request of the processor (120).
[0057] The server (100) receives lidar scan data from the robot cleaner (200) through the communication interface (110), and the processor (120) can create an initial map (20) based on the received lidar scan data and then store the initial map (20) in the memory (130). In addition, the server (100) receives a photographed image from the mobile terminal (300) through the communication interface (110), and the processor (120) can use the received photographed image to create a real-world map (10) from the initial map (20) stored in the memory (130), and then store the real-world map (10) in the memory (130) again. When there is a request from the mobile terminal (300), the server (100) can transmit the real-world map (10) stored in the memory (130) to the mobile terminal (300).
[0058] FIG. 4 is a diagram illustrating a configuration of a robot cleaner included in a system for generating a spatial map according to an embodiment of the present disclosure. Referring to FIG. 4, a robot cleaner (200) according to an embodiment of the present disclosure may include a communication interface (210), an input / output interface (220), a memory (230), a processor (240), a lidar sensor (250), a camera (260), and a driving unit (270). However, the components of the robot cleaner (200) are not limited to the above-described examples, and the robot cleaner (200) may include more or fewer components than the above-described components. Some or all of the communication interface (210), the input / output interface (220), the memory (230), and the processor (240) may be implemented in the form of a single chip.
[0059] The communication interface (210) is a configuration for transmitting and receiving signals (control commands and data, etc.) with an external device via wire or wirelessly, and may be implemented to include a communication chipset that supports various communication protocols. The communication interface (210) may receive a signal from the outside and output it to the processor (240), or transmit a signal output from the processor (240) to the outside. The robot cleaner (200) may communicate with the server (100) through the communication interface (210). According to one embodiment of the present disclosure, the robot cleaner (200) may transmit the lidar scan data acquired through the lidar sensor (250) described below, or the photographed images acquired through the camera (260) described below, to the server (100) or the mobile terminal (300) through the communication interface (210).
[0060] The input / output interface (220) may include an input interface (e.g., touch screen, hard button, microphone, etc.) for receiving control commands or information from a user, and an output interface (e.g., display panel, speaker, etc.) for displaying the results of execution of an operation according to the user's control or the status of the robot cleaner (200).
[0061] The memory (230) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (230) may not exist separately and may be configured to be included in the processor (240). The memory (230) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program for performing operations according to the embodiments described below may be stored in the memory (230). The memory (230) may also provide stored data to the processor (240) upon request of the processor (240).
[0062] The processor (240) controls a series of processes to operate the robot cleaner (200) according to the embodiments described below, and may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU, a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. For example, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0063] The processor (240) can record data in the memory (230) or read data stored in the memory (230), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program stored in the memory (230). Accordingly, the processor (240) can perform operations described in the following embodiments. Unless otherwise specified, operations described as being performed by the robot cleaner (200) in the following embodiments can be considered to be performed by the processor (240).
[0064] The lidar sensor (250) is a configuration for the robot cleaner (200) to scan the distance (depth) to a wall or object in the surrounding space. The robot cleaner (200) can measure the depth of multiple areas included in the target space using the lidar sensor (250), and transmit the measured depth values to the server (100) or create a space map (initial map) based on the measured depth values.
[0065] The camera (260) is a component for the robot cleaner (200) to capture images of the surrounding space. The robot cleaner (200) may be equipped with a camera (260) to recognize objects (e.g., obstacles) in front. According to one embodiment of the present disclosure, the robot cleaner (200) may capture an image of an area of interest through the camera (260) and transmit it to the server (100).
[0066] The driving unit (270) is a component that provides the power necessary for the robot cleaner (200) to perform cleaning operations. The robot cleaner (200) can move within a space and perform cleaning operations (e.g., suction, etc.) due to the driving force provided by the driving unit (270). According to one embodiment, the driving unit (270) may include a motor, a battery, etc.
[0067] Below, the process of generating a spatial map (initial map, real map) by a server (100) according to one embodiment of the present disclosure is described in detail.
[0068] FIG. 5 is a diagram illustrating UI screens displayed on a mobile terminal during the process of generating a spatial map according to one embodiment of the present disclosure. The UI screens (510, 520, 530) illustrated in FIG. 5 may be execution screens of an application installed on a mobile terminal (300), and the server (100) may transmit information (e.g., initial map, real-world map, shooting location / shooting direction, etc.) necessary for displaying the UI screens (510, 520, 530) to the mobile terminal (300).
[0069] The first UI screen (510) displays an initial map (511) generated based on lidar scan data. An icon (512) recommending a shooting location and shooting direction may be displayed on the initial map (511) of the first UI screen (510). The meaning of the icon (512) is explained as follows: from the location where the vertex of the icon (512) is located (shooting location), it means to shoot a space in the direction in which the icon (512) is spread (shooting direction). The first UI screen (510) may also display a message prompting the user to select an area to scan (shoot).
[0070] When a user selects an icon (512) and takes a picture using a mobile terminal (300) according to the shooting position and shooting direction recommended by the icon (512), the mobile terminal (300) can obtain a captured image of a region of interest recommended by the server (100). The method by which the server (100) recommends a region of interest, a shooting position, and a shooting direction is described in detail below with reference to FIG. 6. Meanwhile, as described above, instead of the user taking a picture of the regions of interest using the mobile terminal (300), the robot cleaner (200) can take a picture of the regions of interest using a camera (260) equipped thereon and transmit the captured image to the server (100).
[0071] When a user takes a picture of a target space (area of interest) using a mobile terminal (300) according to icons (512) on an initial map (511) displayed on a first UI screen (510), an initial map (521) with attached real-world images of the areas of interest may be displayed on a second UI screen (520). According to one embodiment of the present disclosure, when the mobile terminal (300) transmits the captured images of the areas of interest to the server (100), the server (100) may register the real-world images obtained from the captured images to the initial map (511) and transmit the result to the mobile terminal (300). The mobile terminal (300) may display the initial map (521) with the registered real-world images on the second UI screen (520), and may also display a message on the second UI screen (520) stating that scanning (capturing) of the areas of interest has been completed and that the floor plan (real-world map) is being restored.
[0072] Then, after a certain period of time, image generation (e.g., outpainting) using a generation model is performed on the initial map (521) to which a real-world image is attached, and the real-world map (531) may be displayed on the third UI screen (530). According to one embodiment of the present disclosure, the server (100) may perform outpainting on an area of the initial map (521) that does not include a real-world image, and the present invention is not limited thereto, and a robot cleaner (200) or a mobile terminal (300) may also perform outpainting.
[0073] 1. How to recommend areas of interest, shooting locations, and shooting directions (Fig. 6)
[0074] FIG. 6 is a diagram for explaining a method for recommending to a user an area of interest where an electronic device (server) should perform a photographing, and a photographing location and a photographing direction according to one embodiment of the present disclosure.
[0075] FIG. 6 illustrates a first initial map (610) and a second initial map (620). When the server (100) generates a spatial map based on the lidar scan data received from the robot cleaner (200), the spatial map has a form similar to the first initial map (610). The server (100) may perform contour approximation on the first initial map (610) to determine areas of interest requiring imaging. When the server (100) performs contour approximation on the first initial map (610), a border organized in a straight line form, such as the second initial map (620), may be displayed.
[0076] The server (100) may determine at least one vertex (position corresponding to the vertex) that satisfies a preset condition among a plurality of vertices included in the boundary of the second initial map (620) as a region of interest. According to one embodiment of the present disclosure, the server (100) may determine a concave vertex among a plurality of vertices included in the boundary of the second initial map (620) as a region of interest. According to one embodiment of the present disclosure, the server (100) may determine the region of interest by comparing the size of the internal angle of the vertex with a preset reference value, and may determine, for example, a vertex whose internal angle is 180 degrees or more as a region of interest. In addition, the server (100) according to one embodiment of the present disclosure may select a vertex according to various other conditions and determine the selected vertex as a region of interest.
[0077] Once the server (100) has determined the area of interest, it can also determine the location (shooting location) and direction (shooting direction) for capturing the area of interest and recommend them to the user. How the server (100) sets the shooting location and shooting direction for the area of interest (621) determined in the second initial map (620) of FIG. 6 is described in detail below.
[0078] Referring to the second initial map (620), it is assumed that the vertex determined as the area of interest (621) is placed at the origin (O), and the two vertices adjacent to the vertex located at the origin (O) are located at A and B, respectively. The server (100) creates a vector ( ) and the vector () from the origin (O) to B. ) into the following mathematical expression 1, and the vector can be obtained.
[0079]
[0080] The server (100) determines the location corresponding to C as the shooting location, and the vector The direction of the shooting can be determined as the shooting direction. The server (100) can display an icon on the second initial map (620) according to the determined shooting location and shooting direction.
[0081] As described above, the server (100) determines a vertex selected from among the vertices included in the boundary line of the initial map according to preset conditions as an area of interest, and can recommend a shooting location and shooting direction to the user based on the distance from the selected vertex to neighboring vertices.
[0082] 2. Method for aligning a real image obtained from a photographed image to an initial map (Figs. 7 and 8)
[0083] Since the user captures the area of interest according to the location and direction recommended by the server (100), the server (100) can roughly know where the captured image corresponds on the initial map. However, in order for the server (100) to attach a real-world image acquired from the captured image to the initial map, it is necessary to determine the exact location corresponding to the captured image (or the real-world image). According to one embodiment of the present disclosure, the server (100) can register the real-world image acquired from the captured image to the initial map based on depth information included in the captured image.
[0084] Hereinafter, with reference to FIGS. 7 and 8, a detailed description will be given of a specific process by which the server (100) aligns a real-world image obtained from a photographed image to an initial map.
[0085] First, a process for generating an estimated map is described with reference to FIG. 7. FIG. 7 is a diagram illustrating a method by which an electronic device (server) according to one embodiment of the present disclosure generates an estimated map representing the layout of an area of interest from a captured image of the area of interest.
[0086] The server (100) can receive a captured image (710) of one area of interest (71) included in the initial map (70) from the mobile terminal (300). The server (100) can obtain a depth image (720) by performing depth estimation on the captured image (710) and convert the depth image (720) into a point cloud (730).
[0087] Next, the server (100) can change the viewpoint of the point cloud to a top view. A point cloud (hereinafter, “top view point cloud”) (740) with the viewpoint changed to a top view is illustrated in FIG. 7.
[0088] The server (100) can obtain an estimated map (750) for an area corresponding to the captured image (710) based on the top-view point cloud (740). The estimated map (750) can represent a layout at a height corresponding to the initial map (70). At this time, the 'height corresponding to the initial map (70)' can mean a height at which 1D scan data (lidar scan data) is collected from the target space to generate the initial map (70). That is, if the initial map (70) is generated based on the lidar scan data collected by the robot cleaner (200), the height (scan height) of the lidar sensor (250) installed in the robot cleaner (200) can be the height corresponding to the initial map (70).
[0089] The server (100) can extract the density of the top-view point cloud (740) for the height corresponding to the initial map (70) and estimate the layout for the height corresponding to the initial map (70) based on the extracted density. Subsequently, the server (100) can generate an estimated map (750) by expressing the estimated layout as a map. The density extracted by the server (100) from the top-view point cloud (740) may correspond to the density on a cross-section of the height corresponding to the initial map (70). In other words, the density extracted by the server (100) from the top-view point cloud (740) may correspond to the density of voxels located at the height corresponding to the initial map (70).
[0090] Hereinafter, a specific method for estimating a layout for a height corresponding to an initial map (70) based on a density extracted from a top-view point cloud (740) by the server (100) is described. According to one embodiment of the present disclosure, the server (100) may compare the density extracted from the top-view point cloud (740) with a preset reference value, binarize a cross-section of a height corresponding to the initial map (70) based on the comparison result, and estimate a layout based on the binarization result. For example, the server (100) may binarize by assigning 1 to a voxel whose extracted density is lower than the reference value and assigning 0 to a voxel whose extracted density is higher than the reference value, and estimate a layout based on the binarization result. At this time, the reference value may be set to an appropriate value so that the binarization result can most closely represent the layout of an actual space.
[0091] The server (100) can estimate the layout based on the binarization result and represent the estimated layout as an estimation map (750). For example, the server (100) can generate the estimation map (750) by displaying voxels assigned 1 in white and voxels assigned 0 in black.
[0092] The reason why the server (100) generates the estimated map (750) by estimating the layout for the height corresponding to the initial map (70) is as follows. The server (100) needs to identify the boundary of the space in order to find the exact location corresponding to the captured image (710). The server (100) can extract the boundary of the space (the boundary between the wall and the floor) (741) from the top-view point cloud (740), but it is difficult to identify the exact boundary in the top view because part of the boundary may be obscured by objects. Therefore, the server (100) estimates the layout at the height at which the information (lidar scan data) used to generate the initial map (70) was collected, and identifies the boundary in the estimated layout.
[0093] Next, with reference to FIG. 8, a process of matching the generated estimated map to the initial map will be described. FIG. 8 is a diagram illustrating a method by which an electronic device (server) according to one embodiment of the present disclosure matches an estimated map to the initial map based on a boundary line extracted from the estimated map.
[0094] Referring to FIGS. 7 and 8, the server (100) can extract a boundary line (751) of a target space from an estimated map (750). For example, the server (100) can extract a boundary line (751) between a wall and a floor from the estimated map (750). In addition, the server (100) can extract a boundary line (72) from an area of interest (area of interest corresponding to the estimated map (750)) (71) of the initial map (70). By comparing the boundary line (751) extracted from the estimated map (750) with the boundary line (72) extracted from the initial map (70), the server (100) can confirm a location where the estimated map (750) matches on the initial map (70).
[0095] Meanwhile, according to one embodiment of the present disclosure, the server (100) may slice (Z-slicing) the point cloud (730) at a predetermined height interval from the ground surface, and estimate a location corresponding to the photographed image (710) in the initial map (70) based on the sliced cross-sections. For example, when the server (100) inputs the cross-sections obtained by slicing the point cloud (730) and the initial map (70) into a pre-trained neural network model, the neural network model may output a result of estimating the location corresponding to the point cloud (730), i.e., the location corresponding to the photographed image (710). To this end, the neural network model may be trained using cross-sections obtained by slicing point clouds corresponding to a plurality of areas in space at predetermined height intervals, and the shape of the boundary line of the spatial map corresponding to the areas, as training data.
[0096] When the server (100) confirms a location where the estimated map (750) matches the initial map (70), it can attach a real-world image obtained from the photographed image (710) to the confirmed location. A specific method by which the server (100) obtains a real-world image from the photographed image (710) is as follows.
[0097] As described above, the server (100) can obtain a point cloud (730) from a photographed image (710) and change the viewpoint of the point cloud (730) to a top view. The server (100) can obtain a real-world image from the top-view point cloud (740).
[0098] If the captured image (710) is an RGB image, the captured image (710) may be composed of pixels to which RGB values are assigned. RGB values may be assigned to voxels constituting the point cloud (730) based on the RGB values assigned to the pixels of the captured image (710). The server (100) may generate a real-world image based on the RGB values assigned to the voxels of the top-view point cloud (740).
[0099] Since the top-view point cloud (740) has a different viewpoint from the captured image (710), some voxels included in the top-view point cloud (740) may not have RGB values assigned. According to one embodiment of the present disclosure, the server (100) may assign RGB values to voxels to which RGB values are not assigned using interpolation. Alternatively, the server (100) may assign arbitrary RGB values to voxels to which RGB values are not assigned, or may not assign any value at all.
[0100] In this way, due to a lack of information, some pixels included in the real image obtained from the photographed image (710) may be assigned RGB values that do not accurately reflect the appearance of the actual space. However, since the real image is reduced and attached to the initial map (70), the user may not be aware of this error.
[0101] The real image obtained according to the method described above may correspond to an image in which the viewpoint of the photographed image (710) is changed to a top view.
[0102] The server (100) can acquire photographed images corresponding to multiple regions of interest included in the target space and align all real-world images acquired from the photographed images to an initial map. Once the real-world images corresponding to the regions of interest are aligned to the initial map, preparations for image generation using the generation model are complete.
[0103] 3. A method for generating an image by determining the area that needs to be generated and generating a prompt based on metadata (Fig. 9).
[0104] Once the real images are aligned to the initial map, the server (100) can determine a generation required area where image generation is required and generate an image (e.g., perform outpainting) for the generation required area using a generative model.
[0105] FIG. 9 is a diagram illustrating a process in which an electronic device (server) according to one embodiment of the present disclosure generates a real-world map from an initial map in which real-world images are aligned.
[0106] The first spatial map (910) illustrated in FIG. 9 is an initial map in which real-world images are aligned. The hatched areas in the first spatial map (910) correspond to areas where images must be generated (areas requiring generation). However, as previously described, there is a high possibility that some areas of the actual space will be omitted in the first spatial map (910) due to the problem of some object surfaces being mistaken for walls during initial map generation. Therefore, the server (100) according to one embodiment of the present disclosure can expand the areas requiring generation to reflect the structure of the actual space as accurately as possible.
[0107] The server (100) may determine boundaries by performing outline approximation on the first spatial map (910), and as a result, a second spatial map (920) may be generated. According to one embodiment of the present disclosure, the server (100) may perform outline approximation by enclosing an area so that the inner angle of the portion where the boundary line is bent becomes a preset constant angle (e.g., 90 degrees or 270 degrees). The second spatial map (920) displays a boundary line, and there are empty areas (areas where no real-world image is attached and no hatching is indicated) within the boundary line. There is a high possibility that objects exist in the empty areas. Therefore, the server (100) may expand the area requiring generation to include the empty areas.
[0108] The result of expanding the area requiring generation in the second spatial map (920) is shown in the third spatial map (930). Referring to the third spatial map (930), all areas within the boundary line, except for the area with the attached real-world image, were determined to be areas requiring generation.
[0109] Once the area requiring generation is determined, the server (100) can generate a prompt for generating an image using a generation model. According to one embodiment of the present disclosure, the server (100) can generate the prompt using metadata about the target space and metadata about objects included in the target space (objects included in the captured image).
[0110] According to one embodiment of the present disclosure, 'metadata for the target space' may include information about the area for which an image is to be generated (e.g., the type of area if the target space is divided into multiple areas such as a bedroom, living room, etc.). The initial map may include metadata for the target space, or the server (100) may obtain metadata for the target space by performing semantic segmentation on the initial map. Alternatively, the user may directly input metadata for the target space. Since the overall atmosphere or layout, etc. differ depending on the type of area included in the space, the server (100) may generate a real-world map (940) that is closer to the actual appearance by reflecting metadata for the target space when generating a prompt.
[0111] According to one embodiment of the present disclosure, 'metadata about an object' may include information about an object included in a captured image (e.g., the type of object). The server (100) may acquire metadata about an object by performing object recognition on the captured image.
[0112] Fig. 9 illustrates an example of a prompt for performing outpainting on a third spatial map (930). The server (100) can generate a first prompt (931) such as "living room layout with bed" using "living room" as metadata for the target space and "bed" as metadata for the object for the bedroom area of the third spatial map (930). Similarly, the server (100) can generate a second prompt (932) such as "living room layout with sofa, TV stand, dining table, and island dining table" using "living room" as metadata for the target space and "sofa," "TV stand," "dining table," and "island dining table" as metadata for the objects for the living room area of the third spatial map (930).
[0113] When the server (100) inputs the first prompt (931) and the second prompt (932) together with the third spatial map (930) into the generation model, the generation model can output a real-world map (940). The generation model may be a neural network model implemented by the processor (120) of the server (100) executing a program stored in the memory (130).
[0114] According to one embodiment of the present disclosure, the server (100) may recommend a new placement of home appliances or furniture to the user by generating a real-world map (940) to include objects (e.g., home appliances or furniture) that do not exist in the actual space.
[0115] For example, if a user inputs information about an appliance or furniture that he or she wishes to additionally place in a target space via a mobile terminal (300), the server (100) may generate a prompt reflecting the information entered by the user. For example, if the user inputs information about the additional placement of an air purifier in a bedroom via the mobile terminal (300), the server (100) may generate a prompt such as "a bedroom layout with an air purifier next to the bed." In this way, the server (100) may generate a prompt to recommend the most appropriate location (e.g., next to the bed) for the object to be added (e.g., an air purifier).
[0116] Or, for example, if a user inputs a request through a mobile terminal (300) to recommend appliances or furniture that can be additionally placed in a target space, the server (100) may determine objects that fit each area of the space and generate a prompt to include the determined objects. If a user inputs a request through a mobile terminal (300) to recommend appliances or furniture that can be additionally placed in a living room, the server (100) may determine to add an audio device to the living room and generate a prompt such as "a living room layout with an audio device next to a TV stand." In this way, the server (100) may determine an object (e.g., audio device) that can be added to a specific area of a space (e.g., living room), determine a suitable location for the determined object to be placed (e.g., next to a TV stand), and then generate a prompt to recommend the object and its placement location.
[0117] Hereinafter, a method for generating a spatial map according to embodiments of the present disclosure will be described with reference to the flowcharts of FIGS. 10 to 18. The steps included in the flowcharts of FIGS. 10 to 18 are performed by at least one of the server (100), the robot cleaner (200), and the mobile terminal (300) of FIG. 1, and therefore, even if the contents described above with reference to FIGS. 1 to 9 are omitted below, they can be equally applied to FIGS. 10 to 18.
[0118] Hereinafter, all steps included in the flowcharts of FIGS. 10 to 18 are described as being performed by the server (100), but some or all of the above steps may be performed by a robot cleaner (200) or a mobile terminal (300).
[0119] Referring to FIG. 10, in step 1001, the server (100) may obtain an initial map of the target space. According to one embodiment of the present disclosure, the server (100) may generate the initial map based on lidar scan data received from the robot cleaner (200). Alternatively, the server (100) may also receive the initial map from the robot cleaner (200).
[0120] In step 1002, the server (100) can acquire a captured image capturing one or more regions of interest in the target space. Once the server (100) determines the region of interest in the target space and recommends a shooting location and shooting direction to the user, the user can capture the region of interest using a mobile terminal (300), and the mobile terminal (300) can transmit the captured image to the server (100). Step 1002 will be described in detail with reference to FIG. 16.
[0121] FIG. 16 is a flowchart illustrating detailed steps included in step 1002 of FIG. 10. Referring to FIG. 16, in step 1601, the server (100) may select at least one vertex that satisfies a preset condition from among a plurality of vertices included in the boundary line of the initial map. For example, the server (100) may select a concave vertex (e.g., a vertex with an internal angle of 180 degrees or greater) from among the vertices included in the boundary line of the initial map.
[0122] At step 1602, the server (100) can determine at least one selected vertex as a region of interest.
[0123] At step 1603, the server (100) can recommend a shooting location and shooting direction to the user based on the distances from at least one selected vertex to neighboring vertices. The server (100) can recommend the shooting location and shooting direction to the user according to the method described above with reference to FIG. 6. The server (100) can recommend the shooting location and shooting direction to the user in various ways, and according to one embodiment, the server (100) can display icons indicating the shooting location and shooting direction on the initial map displayed on the screen of the mobile terminal (300).
[0124] At step 1604, the server (100) can acquire a captured image of the area of interest captured at the recommended shooting location. When the user confirms the recommended shooting location and shooting direction through the mobile terminal (300) and captures the area of interest accordingly, the mobile terminal (300) can transmit the captured image to the server (100).
[0125] Returning to FIG. 10, in step 1003, the server (100) can align the real-world image obtained from the captured image to the initial map based on the depth information contained in the captured image. The detailed steps included in step 1003 will be described in detail with reference to FIGS. 11 to 15.
[0126] FIG. 11 is a flowchart illustrating detailed steps included in step 1003 of FIG. 10. Referring to FIG. 11, in step 1101, the server (100) can estimate a location corresponding to a captured image on the initial map. In other words, the server (100) can determine a location on the initial map where the captured image accurately matches. Since the user captures an area of interest according to the location and direction recommended by the server (100), the server (100) can roughly know where the captured image roughly corresponds to on the initial map. However, in order for the server (100) to attach a real-world image acquired from the captured image to the initial map, it is necessary to determine the exact location corresponding to the captured image (or real-world image). A method for the server (100) to determine the exact location corresponding to the captured image will be described in detail with reference to FIG. 12.
[0127] FIG. 12 is a flowchart illustrating detailed steps included in step 1101 of FIG. 11. Referring to FIG. 12, in step 1201, the server (100) can acquire a depth image corresponding to the captured image. The server (100) can acquire the depth image by performing depth estimation on the captured image.
[0128] At step 1202, the server (100) can acquire a point cloud corresponding to the depth image. The server (100) can generate the point cloud using depth information included in the depth image.
[0129] At step 1203, the server (100) can obtain an estimated map for the area corresponding to the captured image based on the point cloud. The estimated map represents the layout of the area corresponding to the captured image, and in particular, can represent the layout as seen from a height corresponding to the initial map. A specific method for obtaining the estimated map is described in detail with reference to FIG. 13.
[0130] Fig. 13 is a flowchart illustrating detailed steps included in step 1203 of Fig. 12. Referring to Fig. 13, in step 1301, the server (100) can change the viewpoint of the point cloud to a top view.
[0131] In step 1302, the server (100) can extract the density of the point cloud for the height corresponding to the initial map. As described above, the 'height corresponding to the initial map' may refer to the height at which 1D scan data (lidar scan data) is collected from the target space to generate the initial map. The density extracted from the point cloud by the server (100) may correspond to the density on a cross-section at the height corresponding to the initial map. In other words, the density extracted from the point cloud by the server (100) may correspond to the density of voxels located at the height corresponding to the initial map.
[0132] At step 1303, the server (100) can estimate a layout for a height corresponding to the initial map based on the extracted density. A specific method for estimating the layout is described in detail with reference to FIG. 14.
[0133] FIG. 14 is a flowchart illustrating detailed steps included in step 1303 of FIG. 13. Referring to FIG. 14, in step 1401, the server (100) may binarize a cross-section of a height corresponding to an initial map based on a result of comparing a density extracted from a point cloud with a reference value. According to one embodiment of the present disclosure, the server (100) may perform binarization by assigning 1 to a voxel whose extracted density is lower than the reference value and assigning 0 to a voxel whose extracted density is higher than the reference value. At this time, the reference value may be set to an appropriate value so that the binarization result can most closely represent the layout of an actual space.
[0134] At step 1402, the server (100) can estimate the layout based on the binarization result.
[0135] Returning to FIG. 13, at step 1304, the server (100) can obtain an estimated map by representing the estimated layout as a map. According to one embodiment of the present disclosure, the server (100) can generate the estimated map by displaying voxels assigned 1 in white and voxels assigned 0 in black.
[0136] Returning to FIG. 12, at step 1204, the server (100) can match the estimated map to the initial map. According to one embodiment of the present disclosure, the server (100) can match the estimated map to the initial map based on the boundary of the target space identified from the estimated map. A detailed description of the specific method by which the server (100) matches the estimated map to the initial map will be provided with reference to FIG. 15.
[0137] Figure 15 is a flowchart illustrating the detailed steps included in step 1204 of Figure 12. Referring to Figure 15, in step 1501, the server (100) can extract the boundary line of the target space from the estimated map. For example, the server (100) can extract the boundary line between a wall and a floor from the estimated map.
[0138] At step 1502, the server (100) can identify a location on the initial map where the estimated map matches based on the boundary line extracted from the estimated map. According to one embodiment of the present disclosure, the server (100) can extract a boundary line from an area of interest (an area of interest corresponding to the estimated map) of the initial map. By comparing the boundary line extracted from the estimated map with the boundary line extracted from the initial map, the server (100) can identify a location on the initial map where the estimated map matches.
[0139] Returning to Figure 12 again, at step 1205, the server (100) can estimate the location matched by the estimated map as the location corresponding to the captured image.
[0140] Returning to FIG. 11, at step 1102, the server (100) can acquire a real-world image by changing the viewpoint of the captured image to a top view. According to one embodiment of the present disclosure, the server (100) can acquire a point cloud from the captured image, change the viewpoint of the point cloud to a top view, and then acquire a real-world image from the point cloud with the changed viewpoint.
[0141] A point cloud may store information about RGB values corresponding to each location in space (e.g., a location expressed in xyz coordinates). For example, if a captured image is an RGB image, the captured image is composed of pixels to which RGB values are assigned, and RGB values can be assigned to voxels constituting the point cloud based on the RGB values assigned to the pixels of the captured image.
[0142] Accordingly, the server (100) can generate a real-world image based on RGB values assigned to voxels of the point cloud whose viewpoint has been changed to a top view.
[0143] In step 1103, the server (100) can attach a real-world image to an estimated location. Specifically, the server (100) can attach a real-world image (a real-world image representing a top-view layout) acquired from a photographed image to a location estimated to correspond to the photographed image on the initial map.
[0144] Returning to FIG. 10, the server (100) can determine a generation required area where an image needs to be generated from the initial map where the real-world image is aligned in step 1004. Here, the "generation required area" may refer to an area where an image needs to be generated using a generation model. A detailed description of the specific method by which the server (100) determines a generation required area from the initial map will be provided with reference to FIG. 17.
[0145] FIG. 17 is a flowchart illustrating detailed steps included in step 1004 of FIG. 10. Referring to FIG. 17, in step 1701, the server (100) may determine the boundary of the initial map by performing contour approximation on the initial map to which the real-world image is aligned. According to one embodiment of the present disclosure, the server (100) may perform contour approximation by enclosing an area so that the inner angle of the bent portion of the boundary line becomes a preset constant angle (e.g., 90 degrees or 270 degrees).
[0146] At step 1702, the server (100) may determine areas within the initial map's boundary that do not have attached real-world images as areas requiring generation. As previously described, the initial map is likely to omit some areas of the actual space due to the problem of some object surfaces being mistaken for walls during initial map generation. Therefore, the server (100) can expand the area requiring generation to all areas within the initial map's boundary that do not have attached real-world images, thereby reflecting the structure of the actual space as accurately as possible.
[0147] Returning to FIG. 10, at step 1005, the server (100) can generate a real-world map by generating an image for the area requiring generation using the generation model. The server (100) can generate an image by generating a prompt using information about space and objects and inputting the generated prompt into the generation model, which will be described in detail with reference to FIG. 18.
[0148] FIG. 18 is a flowchart illustrating detailed steps included in step 1005 of FIG. 10. Referring to FIG. 18, in step 1801, the server (100) may generate a prompt using metadata about a target space and metadata about an object included in a captured image. The method by which the server (100) generates a prompt using metadata about the target space and the object is as described above with reference to FIG. 9. In addition, as described above, according to one embodiment of the present disclosure, the server (100) may recommend a new arrangement of home appliances or furniture to a user by generating a prompt to include information about an object (e.g., home appliances or furniture) that does not exist in an actual space.
[0149] At step 1802, the server (100) can generate a real-world map by inputting the generated prompt into the generation model together with the initial map in which the real-world image is aligned.
[0150] According to the embodiments described above, it is expected that the user experience will be improved by providing the user with a spatial map (realistic map) expressed similarly to the appearance of an actual space.
[0151] A method for generating a spatial map according to one embodiment of the present disclosure may include the steps of: obtaining an initial map for a target space; obtaining a captured image of one or more areas of interest in the target space; registering an actual image obtained from the captured image to the initial map based on depth information included in the captured image; determining a generation required area requiring image generation in the initial map to which the actual image is registered; and generating an image for the generation required area using a generative model to generate an actual map.
[0152] According to one embodiment, the step of aligning with the initial map may include the step of estimating a location corresponding to the photographed image in the initial map, the step of changing the viewpoint of the photographed image to a top view to acquire the real-world image, and the step of attaching the real-world image to the estimated location.
[0153] According to one embodiment, the step of estimating a location corresponding to the captured image may include the steps of obtaining a depth image corresponding to the captured image, obtaining a point cloud corresponding to the depth image, obtaining an estimated map for an area corresponding to the captured image based on the point cloud, matching the estimated map to the initial map, and estimating a location to which the estimated map is matched as a location corresponding to the captured image.
[0154] According to one embodiment, the step of obtaining the estimated map may include the step of changing the viewpoint of the point cloud to a top view, the step of extracting the density of the point cloud for a height corresponding to the initial map, the step of estimating the layout for the height corresponding to the initial map based on the extracted density, and the step of expressing the estimated layout as a map.
[0155] According to one embodiment, the step of estimating the layout may include the step of binarizing a cross-section of a height corresponding to the initial map based on a result of comparing the extracted density with a preset reference value, and the step of estimating the layout based on the binarization result.
[0156] According to one embodiment, the height corresponding to the initial map may be a height at which 1D scan data is collected from the target space to generate the initial map.
[0157] According to one embodiment, the step of matching the estimated map to the initial map may include the step of extracting a border of the target space from the estimated map and the step of confirming a location on the initial map where the estimated map matches based on the border extracted from the estimated map.
[0158] According to one embodiment, the step of obtaining the photographed image may include the step of selecting at least one vertex that satisfies a preset condition from among a plurality of vertices included in a boundary line of the initial map, the step of determining the selected at least one vertex as the region of interest, the step of recommending a photographing location to a user based on a distance from the selected at least one vertex to neighboring vertices, and the step of obtaining the photographed image obtained by photographing the region of interest at the recommended photographing location.
[0159] According to one embodiment, the step of determining the region requiring generation may include the step of determining a boundary line of the initial map by performing contour approximation on the initial map to which the real-world image is aligned, and the step of determining an area within the boundary line to which the real-world image is not attached as the region requiring generation.
[0160] According to one embodiment, the step of generating the real-world map may include the step of generating a prompt using metadata for the target space and metadata for an object included in the photographed image, and the step of generating the real-world map by inputting the generated prompt into the generation model together with an initial map to which the real-world image is aligned.
[0161] In one embodiment, the prompt may include information about an object to be additionally placed in the target space, and the real-world map may include an image of the object to be additionally placed.
[0162] An electronic device for generating a spatial map according to one embodiment of the present disclosure includes a memory storing a program for generating a spatial map and at least one processor, wherein the at least one processor executes the program to obtain an initial map for a target space, obtain a captured image of one or more areas of interest in the target space, register an actual image obtained from the captured image to the initial map based on depth information included in the captured image, determine a generation required area requiring image generation in the initial map to which the actual image is registered, and then generate an image for the generation required area using a generative model, thereby generating an actual map.
[0163] According to one embodiment, the at least one processor may estimate a location corresponding to the photographed image in the initial map when aligning the real-world image with the initial map, change the viewpoint of the photographed image to a top view to obtain the real-world image, and then attach the real-world image to the estimated location.
[0164] According to one embodiment, the at least one processor may, in estimating a location corresponding to the captured image, obtain a depth image corresponding to the captured image, obtain a point cloud corresponding to the depth image, obtain an estimated map for an area corresponding to the captured image based on the point cloud, match the estimated map to the initial map, and then estimate a location to which the estimated map is matched as a location corresponding to the captured image.
[0165] According to one embodiment, the at least one processor may change the viewpoint of the point cloud to a top view when obtaining the estimated map, extract the density of the point cloud for a height corresponding to the initial map, estimate a layout for the height corresponding to the initial map based on the extracted density, and then express the estimated layout as a map.
[0166] According to one embodiment, the at least one processor may estimate the layout by binarizing a cross-section of a height corresponding to the initial map based on a result of comparing the extracted density with a preset reference value, and then estimating the layout based on the binarization result.
[0167] According to one embodiment, the at least one processor may, when acquiring the captured image, select at least one vertex that satisfies a preset condition from among a plurality of vertices included in a boundary line of the initial map, determine the selected at least one vertex as the region of interest, recommend a shooting location to the user based on a distance from the selected at least one vertex to neighboring vertices, and then acquire the captured image in which the region of interest is captured at the recommended shooting location.
[0168] According to one embodiment, in determining the region requiring generation, the at least one processor may determine a boundary of the initial map by performing contour approximation on the initial map to which the real-world image is aligned, and then determine an area within the boundary to which the real-world image is not attached as the region requiring generation.
[0169] According to one embodiment, the at least one processor may generate a prompt using metadata about the target space and metadata about an object included in the photographed image to generate the real-world map, and then input the generated prompt into the generation model together with an initial map to which the real-world image is aligned, thereby generating the real-world map.
[0170] Various embodiments of the present disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.
[0171] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.
[0172] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0173] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.
[0174] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A step of obtaining an initial map for the target space; A step of obtaining a captured image of one or more areas of interest among the above target spaces; A step of registering an actual image obtained from the photographed image to the initial map based on depth information included in the photographed image; A step of determining a generation required area where image generation is required in the initial map to which the above real-world image is aligned; and A method comprising the step of generating an actual map by generating an image for the region requiring generation using a generative model.
2. In paragraph 1, The step of matching the above initial map is, A step of estimating a location corresponding to the photographed image in the initial map; A step of obtaining the real image by changing the viewpoint of the above shooting image to a top view; and A method characterized by comprising a step of attaching the above-mentioned real image to the above-mentioned estimated location.
3. In either of paragraphs 1 and 2, The step of estimating the location corresponding to the above photographed image is: A step of obtaining a depth image corresponding to the above photographed image; A step of obtaining a point cloud corresponding to the above depth image; A step of obtaining an estimated map for an area corresponding to the photographed image based on the point cloud; a step of matching the above estimated map to the initial map; and A method characterized by including a step of estimating a location matched by the above estimated map as a location corresponding to the photographed image.
4. In any one of paragraphs 1 to 3, The step of obtaining the above estimated map is: A step of changing the viewpoint of the above point cloud to a top view; A step of extracting the density of the point cloud for a height corresponding to the initial map; A step of estimating a layout for a height corresponding to the initial map based on the extracted density; and A method characterized by comprising a step of expressing the estimated layout as a map.
5. In any one of paragraphs 1 to 4, The steps for estimating the above layout are: A step of binarizing a cross-section of a height corresponding to the initial map based on the result of comparing the extracted density with a preset reference value; and A method characterized by comprising a step of estimating the layout based on the binarization result.
6. In any one of paragraphs 1 to 5, The step of obtaining the above photographed image is: A step of selecting at least one vertex that satisfies a preset condition among a plurality of vertices included in the boundary line of the initial map; A step of determining at least one of the selected vertices as the region of interest; A step of recommending a shooting location to the user based on the distance to at least one selected vertex and neighboring vertices; and A method characterized by comprising the step of acquiring the photographed image in which the region of interest is photographed at the recommended photographing location.
7. In any one of paragraphs 1 to 6, The step of determining the area that needs to be created above is: A step of determining the boundary of the initial map by performing contour approximation on the initial map to which the above real-world image is aligned; and A method characterized by comprising a step of determining an area within the above boundary line to which the real image is not attached as the area requiring generation.
8. In any one of paragraphs 1 to 7, The steps for generating the above real-world map are: A step of generating a prompt using metadata for the target space and metadata for an object included in the photographed image; and A method characterized by comprising the step of generating the real-world map by inputting the generated prompt together with the initial map to which the real-world image is aligned into the generation model.
9. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 8 on a computer.
10. In an electronic device (100) for generating a spatial map, A memory (130) in which a program for generating a spatial map is stored; and comprising at least one processor (120), The above at least one processor (120) executes the above program, Obtain an initial map of the target space, Obtain a captured image of one or more areas of interest among the above target spaces, Based on the depth information included in the above-mentioned captured image, an actual image obtained from the above-mentioned captured image is registered to the initial map, After determining the generation required area where the image needs to be generated from the initial map where the above real-world image is aligned, An electronic device that creates an actual map by generating an image for the region requiring generation using a generative model.
11. In paragraph 10, The at least one processor (120) aligns the real image to the initial map, Estimate the location corresponding to the photographed image in the above initial map, After changing the viewpoint of the above shooting image to top view and obtaining the above real image, An electronic device characterized by attaching the above-mentioned real image to the above-mentioned estimated location.
12. In either of paragraphs 10 and 11, The at least one processor (120) estimates a location corresponding to the photographed image, Obtain a depth image corresponding to the above photographed image, Obtain a point cloud corresponding to the above depth image, Based on the above point cloud, an estimated map for the area corresponding to the captured image is obtained, After matching the above estimated map to the initial map, An electronic device characterized in that the location matched by the above estimated map is estimated as a location corresponding to the photographed image.
13. In any one of paragraphs 10 to 12, The at least one processor (120) obtains the estimated map, Change the viewpoint of the above point cloud to top view, Extract the density of the point cloud for the height corresponding to the initial map, After estimating the layout for the height corresponding to the initial map based on the extracted density, An electronic device characterized by representing the above estimated layout as a map.
14. In any one of paragraphs 10 to 13, The at least one processor (120) estimates the layout, Based on the result of comparing the extracted density with a preset reference value, the cross-section of the height corresponding to the initial map is binarized, An electronic device characterized in that the layout is estimated based on the binarization result.
15. In any one of paragraphs 10 to 14, The above at least one processor (120) acquires the photographed image, Select at least one vertex that satisfies a preset condition among multiple vertices included in the boundary of the initial map, determining at least one of the above selected vertices as the region of interest, After recommending a shooting location to the user based on the distance to at least one selected vertex and neighboring vertices, An electronic device characterized by obtaining the photographed image by photographing the region of interest at the recommended photographing location.
Citation Information
Patent Citations
Map generation method and device, electronic equipment and storage medium
CN115239903A
Method and Apparatus for displaying the video on 3D map
KR101932537B1
Method for Processing Registration of Professional Counseling Media
KR1020220007221A
Method and apparatus for brokering transaction based on specification
KR1020220041742A
Apparatus and method for updating space map
KR102339625B1