Method for updating recognition model of robot-type mobile device and electronic device for performing same
By generating training data for robotic mobile devices using virtual environments, the recognition model is optimized for specific spaces, enhancing object detection and reducing operational failures.
Patent Information
- Application Number
- PCT/KR2024/021316
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Recognition models in robotic mobile devices, such as robot vacuum cleaners, exhibit varying performance due to the type of object being recognized and the surrounding environment, leading to inefficiencies and potential operational failures.
A method for updating the recognition model by generating training data using a virtual space similar to the actual environment, incorporating spatial and illuminance characteristics, and simulating failure scenarios to enhance recognition performance.
Improves recognition accuracy and reduces operational failures by optimizing the model for specific spaces, ensuring effective object detection and navigation.
Smart Images

Figure KR2024021316_03072025_PF_FP_ABST
Abstract
Description
Method for updating a recognition model of a robotic mobile device and an electronic device for performing the same
[0001] The present disclosure relates to a method for updating a recognition model of a robotic mobile device and an electronic device for performing the same.
[0002] Neural network models for object recognition (hereinafter referred to as "recognition models") are widely used in electronic devices commonly found around us. For example, electronic devices like robot vacuum cleaners, which move within a specific space (e.g., a house) and perform actions, can use recognition models to recognize objects in their surroundings and perform actions based on the recognition results.
[0003] The recognition performance of a recognition model can be affected by the type of object being recognized and the surrounding environment (e.g., spatial structure, spatial brightness, surrounding objects, etc.). Therefore, the recognition performance of a recognition model can vary depending on the space in which it is used.
[0004] According to one aspect of the present disclosure, a method for updating a recognition model of a robotic mobile device may include a step in which an electronic device obtains spatial scan data for a target space from the robotic mobile device, a step in which the electronic device obtains spatial information including information on a structure of the target space and an object within the target space based on the spatial scan data, a step in which the electronic device inputs the spatial information into a generative model to obtain virtual object data including information on a type of a virtual object and a location of the virtual object, a step in which the electronic device obtains training data using the spatial information and the virtual object data, and a step in which the electronic device updates the recognition model of the robotic mobile device using the training data.
[0005] According to one aspect of the present disclosure, an electronic device for performing a method for updating a recognition model of a robotic mobile device includes a memory in which a program or at least one instruction is stored, and at least one processor operably connected to the memory, whereby the at least one processor executes the program or the at least one instruction, such that the electronic device obtains spatial scan data for a target space from the robotic mobile device, and based on the spatial scan data, the electronic device obtains spatial information including information on a structure of the target space and an object within the target space, and inputs the spatial information into a generative model to obtain virtual object data including information on a type of a virtual object and a position of the virtual object, and after the electronic device obtains training data using the spatial information and the virtual object data, the electronic device can update the recognition model of the robotic mobile device using the training data.
[0006] The above and other aspects and features of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0007] FIG. 1 is a diagram illustrating a system environment for updating a recognition model of a robotic mobile device according to one embodiment of the present disclosure.
[0008] FIG. 2 is a drawing for explaining hardware configurations included in a server according to one embodiment of the present disclosure.
[0009] FIG. 3 is a drawing for explaining components included in a robot vacuum cleaner according to one embodiment of the present disclosure.
[0010] FIG. 4 is a drawing for explaining detailed configurations classified based on functions or roles included in a server according to one embodiment of the present disclosure.
[0011] FIG. 5 is a diagram illustrating an operation of a robotic mobile device according to one embodiment of the present disclosure transmitting failure log data to a server.
[0012] FIG. 6 is a diagram for explaining a process in which learning and inference are performed in a virtual object creation model included in a server according to one embodiment of the present disclosure.
[0013] FIG. 7 is a diagram illustrating an example of training data for learning a virtual object creation model included in a server according to one embodiment of the present disclosure.
[0014] FIG. 8 is a drawing for explaining an operation of a robotic mobile device according to one embodiment of the present disclosure to acquire spatial scan data for a target space and transmit it to a server.
[0015] FIG. 9 is a diagram illustrating a method for a server to create a virtual space according to one embodiment of the present disclosure.
[0016] FIG. 10 is a diagram for explaining a process of creating a virtual object by a virtual object creation model according to one embodiment of the present disclosure.
[0017] FIG. 11 is a diagram illustrating an operation of a server according to one embodiment of the present disclosure to augment a virtual object in a virtual space and generate training data through simulation.
[0018] FIG. 12 is a diagram illustrating an operation in which a server according to one embodiment of the present disclosure performs fine tuning on a recognition model using training data and updates a recognition model of a robotic mobile device based on the result.
[0019] FIGS. 13, 14, 15, 16, 17 and 18 are flowcharts illustrating a method for updating a recognition model of a robotic mobile device according to one embodiment of the present disclosure.
[0020] In the present disclosure, the expression “at least one of a, b or c” may refer to any one of “a”, “b”, “c”, “a and b”, “a and c”, “b and c” or “all of a, b and c”.
[0021] In describing this disclosure, descriptions of technical details that are well-known in the technical field to which this disclosure pertains and are not directly related to this disclosure will be omitted. This is to avoid obscuring the gist of this disclosure by omitting unnecessary explanations and to convey it more clearly. Furthermore, the terms described below are defined based on their functions in this disclosure and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification as a whole.
[0022] For the same reason, some components in the attached drawings are exaggerated, omitted, or schematically depicted. Furthermore, the dimensions of each component do not entirely reflect its actual size. Identical or corresponding components in each drawing are assigned the same reference numbers.
[0023] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to one or more embodiments described in detail below with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. One or more embodiments disclosed are provided to ensure that the disclosure is complete and to fully inform those skilled in the art of the disclosure of the scope of the disclosure. An embodiment of the present disclosure may be defined according to the claims. Like reference numerals denote like elements throughout the specification. Furthermore, when describing an embodiment of the present disclosure, a detailed description of a related function or configuration will be omitted if it is determined that it may unnecessarily obscure the gist of the present disclosure. In addition, the terms described below are terms defined in consideration of the functions of the present disclosure and may vary depending on the intentions or customs of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification.
[0024] In one embodiment, each block of the flowchart diagrams and combinations of the flowchart diagrams can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, and the instructions, when executed by the processor of the computer or other programmable data processing apparatus, can create means for performing the functions described in the flowchart block(s). The computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing apparatus to implement the functions in a particular manner, and the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). The computer program instructions can also be installed on a computer or other programmable data processing apparatus.
[0025] Additionally, each block in the flowchart diagram may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). In one embodiment, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may be executed substantially simultaneously or, depending on the function, may be executed in reverse order.
[0026] The term '~ unit' used in one embodiment of the present disclosure may represent software or a hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC), and the '~ unit' may perform a specific role. Meanwhile, the '~ unit' is not limited to software or hardware. The '~ unit' may be configured to be on an addressable storage medium and may be configured to play one or more processors. In one embodiment, the '~ unit' may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided through a specific component or a specific '~ unit' may be combined to reduce the number of components or separated into additional components. In addition, in one embodiment, the '~ unit' may include one or more processors.
[0027] Below, the meanings of terms used in this disclosure are explained.
[0028] A "robotic mobile device" may refer to any type of device that moves automatically or under user control and performs various actions. The robotic mobile device described in the present disclosure can photograph the surrounding environment to perform actions, recognize objects included in the photographed images, and perform actions based on the object recognition results. Therefore, the robotic mobile device may be equipped with a neural network model (recognition model) for object recognition. A representative example of a robotic mobile device is a robot vacuum cleaner. Although the robotic mobile device is described assuming that the robotic mobile device is a robot vacuum cleaner, the learning method of the neural network model (recognition model) according to one or more embodiments described in the present disclosure and the neural network model learned thereby are not limited thereto and may be used in various types of robotic mobile devices. In addition, the learning method of the neural network model according to one or more embodiments described in the present disclosure and the neural network model learned thereby may be used by other types of devices that need to perform object recognition, not just robotic mobile devices. Instead of 'robotic mobile device', terms such as 'robotic apparatus' or 'automatic driving apparatus' may be used.
[0029] The term "recognition model" may refer to a neural network model for recognizing objects included in images or videos. In the present disclosure, a robotic mobile device may recognize an object in front using the recognition model and control driving based on the recognition result. In one or more embodiments of the present disclosure, the recognition model may be fine-tuned to enhance recognition performance in a specific space (a specific user's space). Terms such as "object recognition model" or "object detection model" may be used instead of "recognition model."
[0030] 'Spatial scan data' may refer to data acquired by a robotic mobile device by scanning a space using sensors. According to one embodiment of the present disclosure, the robotic mobile device may scan a space using a LiDAR sensor and a camera, and the resulting spatial scan data may include map data (e.g., LiDAR scan data) and captured images. Depending on the type of sensor installed in the robotic mobile device, the spatial scan data may include other types of data. According to one embodiment of the present disclosure, the server may acquire spatial information and illuminance characteristic information, which will be described later, by analyzing the spatial scan data received from the robotic mobile device. Terms such as 'scan data' may be used instead of 'spatial scan data'.
[0031] 'Spatial information' refers to information about the layout of a space, and may include information about the structure of the space (e.g., the location of walls, etc.) and information about objects (e.g., home appliances, furniture, etc.) placed in the space (e.g., the type and location of the objects). In addition, the spatial information may additionally include information about the background of the space (e.g., the pattern, texture, feeling of color, color distribution, and tone (white tone, wood tone, etc.) of the floor or wallpaper, etc.). According to one embodiment of the present disclosure, a robotic mobile device or a server may generate a 'spatial map' through SLAM (Simultaneous Localization And Mapping), and the spatial map may include spatial information. Instead of 'spatial information', terms such as 'spatial layout' or 'layout information' may be used. Additionally, terms such as 'map data' or 'LiDAR map' may be used instead of 'spatial map'.
[0032] "Illuminance characteristic information" may refer to information regarding the characteristics of illuminance generated by lighting or natural light within a space. Specifically, illuminance characteristic information may include various types of information related to the brightness and color temperature of multiple areas within a space. For example, illuminance characteristic information may include information regarding how brightness and color temperature vary across areas within a space and over time. According to one embodiment, images captured of multiple areas within a specific space may include illuminance characteristic information. Each captured image may indicate the brightness level at the time of capture (how bright the captured location is). Furthermore, each captured image may further include metadata corresponding to the illuminance characteristic information. According to one embodiment, a spatial map for a specific space may also include illuminance characteristic information. For example, information such as average illuminance may be recorded for each of the multiple areas within the spatial map. The illuminance characteristics of a space may be influenced by the structure of the space (e.g., the location of windows and lights, the arrangement of objects, etc.). According to one embodiment of the present disclosure, a server can obtain illumination characteristic information for a space by analyzing a captured image of the space. Instead of "illuminance characteristic information," terms such as "brightness information" or "illuminance information" may also be used.
[0033] 'Failure log data' may refer to log data regarding a failure event that occurs while a robotic mobile device is moving and operating within a space. In this case, a 'failure event' or 'failure' may refer to a situation in which a problem occurs in the movement or operation of the robotic mobile device due to a failure in object recognition of the robotic mobile device. In addition, a 'failure event' or 'failure' may refer to a situation in which the robotic mobile device fails to recognize an object. The types of failure events may vary, such as a cable snag or a failure in avoiding feces. The failure log data may include information regarding the type and location of the failure event, and additionally, may include spatial information (e.g., a spatial map acquired through SLAM) and an image capturing the failure event (e.g., an image captured 5 seconds before the failure event occurred, an image captured at the time the failure event occurred, etc.). Therefore, the failure log data may include information regarding where and what type of failure (failure) occurred within the space. Instead of 'failure log data', the terms 'failure log', 'failure event log data', 'fault log', or 'failure situation information' may also be used.
[0034] "Generative AI" can refer to AI technology capable of generating new text, images, etc. in response to input data (e.g., text, images, etc.). Representative examples of generative AI are described in the "Generative Models" section below.
[0035] A "generative model" can refer to a neural network model that implements generative AI technology. By learning the patterns and structures of training data, a generative model can generate new data with similar characteristics to the input data or new data corresponding to the input data. For example, if the input data is text containing a question, the generative model can generate and output an answer to the question. Or, for example, if the input data is text containing a request, the generative model can output text or images generated according to the request. The "virtual object generation model" described below is a generative model. Terms such as "generative AI model" may also be used instead of "generative model."
[0036] "Virtual space data" may refer to data representing a virtual space (e.g., 3D modeling data). According to one embodiment of the present disclosure, a server may generate a virtual space by reflecting the characteristics of an actual space (e.g., the layout and illumination characteristics of the space, etc.). That is, the server may generate virtual space data based on spatial information and illumination characteristic information about a specific space. The virtual space may be a space that serves as a basis for generating training data used for learning a recognition model. According to one embodiment of the present disclosure, the server may provide realistic and diverse variations to the space by generating the virtual space as a 3D space, thereby obtaining a large amount of training data. Terms such as "simulation space data" may also be used instead of "virtual space data."
[0037] A 'virtual object generation model' may refer to a generative AI model that generates virtual objects that are likely to be placed in a space based on the characteristics of the space. When spatial information about a specific space is input, the virtual object generation model may determine a virtual object that is likely to be placed in the space based on the spatial information, and generate and output data about the determined virtual object. Alternatively, when spatial information about a specific space is input, the virtual object generation model may determine a virtual object that is likely to cause a failure event in the space based on the spatial information, and generate and output data about the determined virtual object. According to one embodiment of the present disclosure, the virtual object generation model may generate and output data about virtual objects that are likely to be placed in a specific space (e.g., type and location of the object) based on context, such as the structure of the specific space and the arrangement of objects within the specific space. The virtual object generated by the virtual object generation model may be augmented in the virtual space described above.Instead of 'virtual object generation model', terms such as 'failure environment generation model', 'failure case generation model', 'dangerous object generation model', 'failure-causing object generation model', 'failure likelihood generation model', 'difficulty generation model', 'generative artificial intelligence model', 'generative model', or 'neural network model' could also be used.
[0038] 'Virtual object data' may refer to data including information about the type and location of a virtual object. According to one embodiment of the present disclosure, a virtual object may refer to an object that a robotic mobile device must avoid or be careful of while driving or operating. Alternatively, according to one embodiment of the present disclosure, a virtual object may refer to an object that is likely to cause a failure in the driving or operation of a robotic mobile device. According to one embodiment of the present disclosure, virtual object data may further include context information, which is information about objects that are likely to be located around the virtual object (objects related to the virtual object). For example, if the virtual object is a 'wire,' the context information of the virtual object may be 'home appliances' and 'computers,' which means that the virtual object of 'wires' is likely to be located around 'home appliances' or 'computers.' According to one embodiment of the present disclosure, a server can create a virtual space environment for acquiring training data by augmenting virtual objects in a virtual space based on virtual space data and virtual object data. Instead of "virtual object data," terms such as "dangerous object data," "failure-causing object data," "failure environment data," or "difficulty information / obstacle information" may also be used.
[0039] "Training data" may refer to data for training a recognition model. According to one embodiment of the present disclosure, the training data is data for performing supervised learning on a recognition model, and may be composed of pairs of images containing objects and labels indicating the types of objects contained in the images. According to one embodiment of the present disclosure, the training data may include "real data" and "synthetic data." Synthetic data may refer to virtual data that has statistical characteristics similar to real data and is generated by reproducing results similar to the results of analyzing real data. The synthetic data according to one embodiment of the present disclosure may include data acquired in the virtual space environment described above (e.g., an image virtually photographing a virtual space in which virtual objects are placed). Instead of "synthetic data," terms such as "synthetic image," "virtual data," or "virtual recognition data" may also be used. Additionally, terms such as 'learning data' or 'recognition data' may be used instead of 'training data'.
[0040] Hereinafter, one or more embodiments of the present disclosure will be described in detail with reference to the drawings.
[0041] The present disclosure relates to a method for updating a recognition model installed in a robotic mobile device, and more particularly, to a method for updating the recognition model so that it can exhibit high recognition performance in a specific space. To this end, one or more embodiments of the present disclosure generate training data using a virtual space similar to a real space and virtual objects likely to cause recognition failure in the real space, and the generated training data is used to train the recognition model.
[0042] FIG. 1 is a diagram illustrating a system environment for updating a recognition model of a robotic mobile device according to one embodiment of the present disclosure. Referring to FIG. 1, a system according to one embodiment of the present disclosure may include a server (100) and a robot vacuum cleaner (200). As described above, one or more embodiments of the present disclosure may be applied to various types of robotic mobile devices other than the robot vacuum cleaner (200).
[0043] As illustrated in FIG. 1, it is assumed that the robot cleaner (200) operates in space A. For example, space A may be the home of a user using the robot cleaner (200). Space A may be referred to as a target space, and the server (100) may update the recognition model of the robot cleaner (200) to be optimized for the target space.
[0044] The server (100) may be any type of electronic device having a computing function. According to one embodiment of the present disclosure, the server (100) may be an IoT server for controlling a robot cleaner (200) or providing services related to the robot cleaner (200).
[0045] When the robot cleaner (200) requests the server (100) to update the recognition model, the server (100) may update the recognition model and transmit the updated recognition model to the robot cleaner (200). Specifically, the server (100) stores a recognition model (e.g., a recognition model identical to the recognition model installed in the robot cleaner (200) or a basic recognition model), and the server (100) may update the recognition model by performing fine tuning on the stored recognition model, and then transmit parameter information of the updated recognition model to the robot cleaner (200). Alternatively, according to one embodiment of the present disclosure, the server (100) may transmit training data to the robot cleaner (200), and the robot cleaner (200) may update the recognition model using the training data received.
[0046] The server (100) can update the recognition model of the robot cleaner (200) so that the robot cleaner (200) can exhibit high recognition performance when used in space A. To this end, the server (100) can acquire training data using a virtual space similar to space A, and update the recognition model using the acquired training data. The server (100) can use a generative model to generate virtual objects that the recognition model is likely to fail to recognize, and can use the generated virtual objects when acquiring training data. The purpose and reason for the server (100) updating the recognition model of the robot cleaner (200) will be described in detail as follows.
[0047] The robot cleaner (200) can identify an object located in front by performing object recognition on an image captured by a camera equipped in the robot cleaner (200), and can determine whether to clean up the object or perform evasive maneuvers based on the type of the identified object. The robot cleaner (200) can include a neural network model for object recognition, i.e., a recognition model, to recognize an object included in the captured image. However, if the recognition performance (recognition accuracy) of the recognition model equipped in the robot cleaner (200) is low, problems such as incorrectly recognizing an object and not performing evasive maneuvers when necessary may occur. For example, the robot cleaner (200) may not recognize a wire laid on the floor and may get caught on the wire while driving, or the robot cleaner (200) may not recognize pet feces and fail to perform evasive maneuvers. In this way, a situation in which a problem occurs in the driving or operation of the robot cleaner (200) due to the recognition model of the robot cleaner (200) failing to recognize an object can be referred to as a 'failure event'. In addition, a 'failure event' can also mean a situation in which the recognition model of the robot cleaner (200) fails to recognize an object.
[0048] The recognition performance of the recognition model installed in the robot cleaner (200) may be affected by the type of object to be recognized or the surrounding environment (e.g., space structure, space brightness, surrounding objects, space background, etc.), and therefore, the recognition performance of the recognition model may vary depending on the space in which the robot cleaner (200) is used.
[0049] Typically, the recognition model installed in a robot vacuum cleaner (200) upon shipment is a model commonly installed in all units, and is likely not optimized for a specific space. Therefore, when the robot vacuum cleaner (200) is used in a specific space, the expected recognition performance may not be achieved.
[0050] To solve this problem, a recognition model must be additionally trained to be specialized for the space in which the robot cleaner (200) is used (e.g., the user's home). However, if training data is collected by photographing the space while placing objects in the space, it is difficult to secure a sufficient amount of data for learning.
[0051] In one or more embodiments of the present disclosure, the server (100) can secure a sufficient amount of training data by acquiring training data using a virtual space. Specifically, the server (100) can generate a virtual space (e.g., a 3D space) reflecting the characteristics of space A in which the robot cleaner (200) is used, generate virtual objects that the recognition model is likely to fail to recognize, and then acquire training data using the generated virtual space and virtual objects. The specific process by which the server (100) acquires training data and updates the recognition model is described in detail below with reference to other drawings.
[0052] FIG. 2 is a diagram illustrating hardware components included in a server according to one embodiment of the present disclosure. Referring to FIG. 2, a server (100) according to one embodiment of the present disclosure may include a communication interface (110), a processor (120), and a memory (130). However, the server (100) may include more or fewer components than the aforementioned components. Some or all of the communication interface (110), the processor (120), and the memory (130) may be implemented in the form of a single chip.
[0053] The communication interface (110) is a configuration for transmitting and receiving signals (such as control commands and data) with an external device via wire or wirelessly, and may be implemented to include a communication chipset that supports various communication protocols. The communication interface (110) may receive signals from the outside and output them to the processor (120), or transmit signals output from the processor (120) to the outside. The server (100) may communicate with the robot cleaner (200) via the communication interface (110).
[0054] The server (100) can receive a request for updating a recognition model from a robot cleaner (200) through a communication interface (110) and transmit parameter information of the updated recognition model to the robot cleaner (200).
[0055] The processor (120) controls a series of processes so that the server (100) operates according to one or more embodiments described below, and may be composed of one or more processors. The one or more processors included in the processor (120) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. The one or more processors included in the processor (120) may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP). When the one or more processors included in the processor (120) are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0056] The processor (120) can write data to the memory (130) or read data stored in the memory (130), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program or at least one instruction stored in the memory (130). Accordingly, the processor (120) can perform operations described in one or more of the following embodiments. Operations described as being performed by the server (100) or detailed components (410 to 490 of FIG. 4) included in the server (100) in one or more of the following embodiments can be regarded as being performed by the processor (120) unless otherwise specifically described.
[0057] The memory (130) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (130) may not exist separately and may be configured to be included in the processor (120). The memory (130) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program or at least one instruction for performing operations according to one or more embodiments described below may be stored in the memory (130). The memory (130) may also provide stored data to the processor (120) at the request of the processor (120).
[0058] FIG. 3 is a diagram for explaining components included in a robot cleaner (200) according to one embodiment of the present disclosure. Referring to FIG. 3, the robot cleaner (200) according to one embodiment of the present disclosure may include a communication interface (210), an input / output interface (220), a memory (230), a processor (240), a camera (250), a lidar sensor (260), and a driving module (270). However, the components of the robot cleaner (200) are not limited to the above-described examples, and the robot cleaner (200) may include more or fewer components than the above-described components. In one embodiment, some or all of the communication interface (210), the input / output interface (220), the memory (230), and the processor (240) may be implemented in the form of a single chip, and the processor (240) may include one or more processors.
[0059] The communication interface (210) is a configuration for transmitting and receiving signals (control commands and data, etc.) with an external device via wire or wirelessly, and may be configured to include a communication chipset that supports various communication protocols. The communication interface (210) may receive a signal from the outside and output it to the processor (240), or transmit a signal output from the processor (240) to the outside. According to one embodiment, the robot cleaner (200) may transmit spatial scan data acquired using a camera (250) and a lidar sensor (260) described below to the server (100) via the communication interface (210).
[0060] The input / output interface (220) may include an input interface (e.g., touch screen, hard button, microphone, etc.) for receiving control commands or information from a user, and an output interface (e.g., display panel, speaker, etc.) for displaying the results of execution of an operation according to the user's control or the status of the robot cleaner (200).
[0061] The memory (230) is a configuration for storing various programs or data, and may be configured as a storage medium such as a ROM, a RAM, a hard disk, a CD-ROM, and a DVD, or a combination of storage media. The memory (230) may not exist separately and may be configured to be included in the processor (240). The memory (230) may be configured as a volatile memory, a non-volatile memory, or a combination of volatile memory and non-volatile memory. A program for performing operations according to one or more embodiments described below may be stored in the memory (230). The memory (230) may also provide stored data to the processor (240) at the request of the processor (240).
[0062] The processor (240) controls a series of processes to operate the robot cleaner (200) according to one or more embodiments described below, and may be composed of one or more processors. In this case, the one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU, a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. For example, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0063] The processor (240) can record data in the memory (230) or read data stored in the memory (230), and in particular, process data according to predefined operation rules or artificial intelligence models by executing a program stored in the memory (230). Accordingly, the processor (240) can perform operations described in one or more of the following embodiments, and operations described as being performed by the robot cleaner (200) in one or more of the following embodiments can be regarded as being performed by the processor (240) unless otherwise specifically described.
[0064] In FIG. 3, the processor (240) is illustrated as including a recognition model (241). The recognition model (241) may refer to a neural network model implemented by the processor (240) executing a program stored in the memory (230). The recognition model (241) may recognize an object included in an image captured by the camera (250). The processor (240) may update the recognition model (241) according to parameter information received from the server (100).
[0065] The camera (250) is configured to capture images of the surroundings of the robot cleaner (200). The robot cleaner (200) can capture images of the front using the camera (250) and recognize objects included in the captured images using a recognition model (241).
[0066] The lidar sensor (260) is configured to scan the distance (depth) to a wall or object in the surrounding space. The robot cleaner (200) can measure the depth of multiple areas included in the target space using the lidar sensor (260) and transmit the measured depth values (lidar scan data) to the server (100).
[0067] The drive module (270) is a module for driving or performing cleaning operations of the robot cleaner (200), and may include a motor and a battery. The motor is a component for providing the power necessary for the robot cleaner (200) to perform cleaning operations. According to one embodiment, the robot cleaner (200) can move within a space and perform cleaning operations (e.g., suction, etc.) due to the driving force provided by the motor. The battery can provide power to components included in the robot cleaner (200).
[0068] FIG. 4 is a diagram for explaining detailed configurations classified based on function or role, included in a server according to one embodiment of the present disclosure. The detailed configurations (410 to 490) of the server (100) illustrated in FIG. 4 may be software configurations implemented by the processor (120) of the server (100) executing a program stored in the memory (130), or may be virtual configurations for which no matching hardware device actually exists. In other words, the operations performed by the processor (120) of the server (100) by executing the program stored in the memory (130) may be classified into a plurality of groups based on function or purpose, and the entities performing the operations included in each classified group may be expressed as the detailed configurations (410 to 490) of FIG. 4. Accordingly, the operations described as being performed by the detailed configurations (410 to 490) illustrated in FIG. 4 may be viewed as actually being performed by the processor (120) of the server (100) executing the program stored in the memory (130). Hereinafter, one or more embodiments in which a server (100) updates a recognition model (241) of a robot cleaner (200) will be described in detail with reference to FIGS. 4 to 11.
[0069] 1. Collection and transmission of failure log data
[0070] When a failure event occurs while the robot cleaner (200) is performing cleaning in space A (target space), the robot cleaner (200) can store information related to the failure event as failure log data and transmit it to the server (100). At this time, the failure event can mean both a situation in which the recognition model (241) of the robot cleaner (200) fails to recognize an object, as described above, or a situation in which the cleaning operation of the robot cleaner (200) fails due to this.
[0071] Failure log data may include the type and location of the failure event, spatial information (e.g., spatial maps acquired via SLAM), and captured images of the failure situation (e.g., images captured seconds before the failure event, images captured at the time the failure event occurred, etc.). In other words, failure log data may include information about where and what type of failure occurred, the structure of the space where the failure occurred, and what the failure looked like when it occurred.
[0072] According to one embodiment of the present disclosure, the robot cleaner (200) may store failure log data in the memory (230) and transmit the failure log data to the server (100) when a certain condition is met. For example, the robot cleaner (200) may transmit the failure log data to the server (100) when a preset number of failure events occur or when a preset percentage of the total number of recently performed cleaning events occur. Alternatively, for example, the robot cleaner (200) may periodically transmit the failure log data to the server (100).
[0073] Referring to FIG. 4, when the robot cleaner (200) transmits failure log data to the communication module (410) of the server (100), the communication module (410) can transmit the received failure log data to the log data analysis module (430). The log data analysis module (430) can analyze the failure log data and output information included in the failure log data.
[0074] 2. Learning a virtual object creation model using failure log data
[0075] Among the information included in the failure log data, the type and location of the failure event and spatial information can be used to train the virtual object generation model (450). As described above, the virtual object generation model (450) is a generative AI model for generating virtual object data based on spatial information. The virtual object generation model (450) can be said to learn about the relationship between spatial characteristics (particularly, the types and locations of objects placed in the space) and objects (particularly, objects likely to cause failure events). By training the virtual object generation model (450) using information about where and what type of failure event actually occurred, the virtual object generation model (450) can generate data (types and locations of virtual objects) about virtual objects likely to cause failure events according to the characteristics of the space. In detail, when spatial information about a target space (structure of the space, types and locations of objects placed in the space) and information about a failure event (type and location of an object that caused a failure event) are input into a virtual object creation model (450), the virtual object creation model (450) can add a new object that is likely to be placed in the target space (in particular, an object that is likely to cause a failure event), and the new object added at this time corresponds to a virtual object.
[0076] For example, in an environment where there are electronic devices around, a wire is laid on the floor, but the recognition model (241) of the robot cleaner (200) does not recognize the wire, and a failure event occurs in which the robot cleaner (200) gets caught on the wire, and it is assumed that a virtual object creation model (450) is trained using information about this failure event. In this case, when spatial information representing a similar environment (an environment where electronic devices are present) is input, the virtual object creation model (450) can create the wire laid around the electronic device as a virtual object.
[0077] The virtual object data output by the virtual object generation model (450) may include information regarding the type and location of the virtual object. Furthermore, the virtual object data may further include context information, which is information regarding surrounding objects related to the virtual object (objects likely to be located around the virtual object).
[0078] A specific method for training a virtual object creation model (450) using failure log data is described in detail below with reference to FIG. 6.
[0079] FIG. 5 is a diagram illustrating an operation of a robotic mobile device transmitting failure log data to a server according to one embodiment of the present disclosure. Referring to FIG. 5 , when a robot cleaner (200) transmits failure log data to a server (100), the server (100) can train a virtual object creation model using the failure log data.
[0080] FIG. 5 illustrates a first screen (510) indicating a failure event and a configuration (520) of failure log data. According to one embodiment of the present disclosure, when a failure event occurs in a robot cleaner (200), the first screen (510) is displayed on a user device (e.g., a smartphone) for controlling the robot cleaner (200), thereby allowing the user to check information about the currently occurring failure event. The first screen (510) may display a space map (511) for a target space (space A), and the location where the failure event occurred may be displayed on the space map (511). In addition, the first screen (510) may display an image (512) capturing the situation in which the failure event occurred and information (513) about the type of object that caused the failure event.
[0081] Referring to the structure of failure log data (520), the failure log data may include an ID of a robot cleaner for identifying a device where a failure event occurred, a location where the failure event occurred (location on a spatial map), a type of failure event (e.g., cable snagging, failure to avoid feces, etc.), spatial information (e.g., spatial map acquired through SLAM), a sample image capturing a failure event (e.g., an image captured 5 seconds before the failure event occurred, an image captured at the time the failure event occurred, etc.), etc. According to one embodiment of the present disclosure, the server (100) may receive failure log data not only from the robot cleaner (200) operating in space A, but also from various external sources (e.g., robot cleaners) used in other spaces. Accordingly, the server (100) may train a virtual object generation model (450) using information on failure events that occurred in various spaces. As a result, the virtual object generation model (450) may sufficiently learn about the relationship between the characteristics of a space (e.g., structure of a space, arrangement of objects in a space, etc.) and an object that causes a failure.
[0082] Meanwhile, as described below, the captured images of failure situations included in the failure log data may be used when generating training data for a recognition model. For example, after a failure event occurs, the robot cleaner (200) can accurately identify the type of object that caused the failure event through additional recognition attempts, and the server (100) can generate training data by labeling the type of object in the captured images included in the failure log data.
[0083] Hereinafter, with reference to FIGS. 6 and 7, a specific method for training a virtual object creation model (450) using failure log data will be described in detail.
[0084] FIG. 6 is a diagram for explaining a process in which learning and inference are performed in a virtual object generation model included in a server (100) according to one embodiment of the present disclosure. Referring to FIG. 6, the virtual object generation model (450) may include a layout encoder (451), an item encoder (452), an obstacle encoder (difficulties encoder) (453), a transformer encoder (454), and an obstacle extractor (difficulties extractor) (455).
[0085] The virtual object generation model (450) can be trained using spatial information about the target space, information about the type and location of the failure event. According to one embodiment of the present disclosure, when spatial information about the target space, information about the type and location of the failure event are input into the virtual object generation model (450), the virtual object generation model (450) can output information about a virtual object that is likely to cause a failure event (e.g., information about the type and location of the virtual object, and objects around the virtual object).
[0086] Hereinafter, the input and output of the detailed components (451 to 455) included in the virtual object generation model (450) will be described in detail. The virtual object generation model (450) according to one embodiment of the present disclosure is implemented as a transformer encoder, and thus can be trained by an unsupervised learning method. That is, the virtual object generation model (450) can be trained through a loss function that compares the outputs of the layout encoder (451), the object encoder (452), and the obstacle encoder (453) with the outputs of the transformer encoder (454), and a detailed description thereof will be omitted.
[0087] The processor (120) of the server (100) can obtain an image (61) representing the structure of the target space from spatial information and input the obtained image (61) to the layout encoder (451). According to one embodiment of the present disclosure, as illustrated in FIG. 6, a 2D top view image (61) of the target space can be input to the layout encoder (451). At this time, a convolutional neural network (CNN) suitable for extracting features from the image can be used as the layout encoder (451). The layout encoder (451) can extract features from the image (61) representing the structure of the target space and output a feature vector.
[0088] The processor (120) of the server (100) determines the types (c) of objects (62) placed in the target space from spatial information. j ), location (t j ) and size (s j ) can be extracted and input to the object encoder (452). At this time, the position (t) of the objects (62) j ) may be coordinate information indicating the location on the 2D top view image (61). In addition, at this time, the size of the objects (s j) may be information indicating a size (e.g. 1D size) measured based on one direction. The object encoder (452) may output a feature vector for each object (62).
[0089] The processor (120) of the server (100) determines the type (c) of objects (63) that caused the failure from the information about the type and location of the failure event. j ) and location (t j ) can be extracted and input to the obstacle encoder (453). At this time, the positions (t) of the objects (63) that caused the failure j ) may also be coordinate information indicating a location on a 2D top view image (61). The obstacle encoder (453) may output a feature vector for each of the objects (63) that caused a failure. According to one embodiment of the present disclosure, the processor (120) may enable the virtual object generation model (450) to be trained autoregressively by masking some of the vectors extracted from the objects (63) that caused a failure.
[0090] An example of input (training data) for learning a virtual object generation model (450) is illustrated in Fig. 7. The first image (710) of Fig. 7 is an image visualizing data (training data) input to the virtual object generation model (450), and the second image (720) displays vectors input to the virtual object generation model (450).
[0091] Looking at the first image (710), objects and objects that triggered failure events are displayed on a 2D top-view image representing the structure of the target space. An image from which the objects and objects that triggered failure events are removed from the first image (710) can be input to the layout encoder (451).
[0092] Vectors indicating the types and locations of objects displayed in the first image (710) and objects that caused failure events are displayed in the second image (720). The vectors displayed in the second image (720) can be input to the object encoder (452) and the obstacle encoder (453).
[0093] When feature vectors extracted from the layout encoder (451), the object encoder (452), and the obstacle encoder (453) are input to the transformer encoder (454), the transformer encoder (454) can output feature vectors similar to the input feature vectors. In particular, the transformer encoder (454) can learn the placement probability of obstacles (objects that caused failure) according to the placement of objects (62) by outputting feature vectors of surrounding objects together with objects (63) that caused failure, as can be seen in the first region (610). As a result, the query vector (64), which is the output of the transformer encoder (454), can include feature vectors for surrounding objects together with the feature vector for obstacles, similar to the vectors included in the first region (610). When the query vector (64) is passed through the obstacle extraction unit (455), information on the type and location of a new virtual object (obstacle) (65) is output, and as can be confirmed in the second area (620), feature vectors for the new virtual object (65) and feature vectors for surrounding objects can be output. That is, the output of the virtual object creation model (450) can include the type and location of the new virtual object (65), and context information for the new virtual object (65).
[0094] 3. Start updating the robot vacuum cleaner's recognition model.
[0095] Methods for starting an update of a recognition model (241) of a robot vacuum cleaner (200) are described.
[0096] According to one embodiment of the present disclosure, the robot cleaner (200) may ask the user whether to update the recognition model (241) upon initial operation or reset, and if the user agrees, may initiate an update of the recognition model (241). Alternatively, the robot cleaner (200) may automatically initiate an update of the recognition model (241) upon initial operation or reset.
[0097] According to one embodiment of the present disclosure, the robot cleaner (200) can start updating the recognition model (241) when a user requests an update of the recognition model (241) through the input / output interface (220).
[0098] According to one embodiment of the present disclosure, when a failure event occurs, the robot cleaner (200) may ask the user whether to update the recognition model (241), and if the user agrees, the robot cleaner (200) may start updating the recognition model (241). Alternatively, the robot cleaner (200) may automatically start updating the recognition model (241) when a failure event occurs. Specifically, when a preset number of failure events occur or a preset ratio of failure events occurs among the total number of recently performed cleanings, the robot cleaner (200) may transmit failure log data to the server (100), and then start updating the recognition model (241) automatically or with the user's consent.
[0099] In this way, the robot cleaner (200) can start updating the recognition model (241) automatically or with the user's consent when a preset condition is met.
[0100] 4. Collection and transmission of spatial scan data
[0101] When the recognition model (241) of the robot cleaner (200) is updated, the robot cleaner (200) can collect space scan data by scanning the target space (space A) using the equipped sensors (camera (250) and lidar sensor (260)) and transmit the space scan data to the server (100). The space scan data can include map data (e.g. lidar scan data) collected through the lidar sensor (260) of the robot cleaner (200) and captured images acquired through the camera (250) (e.g. images captured in various environments (illuminances) of multiple areas within the target space).
[0102] A method for a robot vacuum cleaner (200) to acquire space scan data for a target space (space A) is described in detail with reference to FIG. 8.
[0103] If there is a pre-generated space map (800), the robot cleaner (200) can set an optimal driving path on the space map (800) (e.g., set a path using a single brush drawing method) and collect space scan data while driving along the set driving path. For example, the robot cleaner (200) can scan the depth of the space using a lidar sensor (260) or capture images of the space using a camera (250) at regular intervals or regular periods while driving along the driving path. At this time, the captured images can be used later when the server (100) identifies the types and locations of objects placed in the target space, analyzes the illuminance characteristics of the target space, or analyzes the background of the target space (e.g., pattern, texture, color, color distribution, and color tone (white tone, wood tone, etc.) of the floor or wallpaper, etc.).
[0104] Even if there is no pre-generated space map (800), the robot cleaner (200) can collect space scan data while driving in the target space. For example, the robot cleaner (200) can drive and scan to create a space map (800) for the target space when first started, and the robot cleaner (200) can also acquire space scan data during this process. In this case, since the robot cleaner (200) cannot set a driving path in advance, it can collect space scan data while setting a driving path in real time based on information sensed while driving.
[0105] According to one embodiment of the present disclosure, the robot cleaner (200) may collect additional space scan data by automatically or at the user's request, changing the environment of the target space (e.g., turning lights on and off, artificially placing objects that may cause failure, etc.).
[0106] When the robot cleaner (200) transmits the spatial scan data collected according to the method described above to the server (100), the server (100) can obtain spatial information and illuminance characteristic information by analyzing the spatial scan data. When the communication module (410) of the server (100) transmits the received spatial scan data to the spatial information analysis module (420), the spatial information analysis module (420) can obtain spatial information and illuminance characteristic information for the target space (space A) by analyzing the spatial scan data. In detail, the spatial information analysis module (420) can obtain spatial information by analyzing the structure of the target space based on at least one of map data and a photographed image, analyzing the type and location of an object placed in the target space, and analyzing the background of the target space. In addition, the spatial information analysis module (420) can identify the illuminance characteristic of the target space based on the photographed image.
[0107] The method by which the spatial information analysis module (420) obtains illumination characteristic information for the target space is described in detail as follows.
[0108] The spatial information analysis module (420) can determine the color temperature distribution and brightness distribution in the target space by analyzing the captured images included in the spatial scan data. According to one embodiment of the present disclosure, the spatial information analysis module (420) can analyze the color according to the RGB distribution of the captured image, convert the color space of the captured image (e.g., convert from the RGB color space to the CIE LAB color space), and then analyze the brightness of the captured image. As a result, the spatial information analysis module (420) can obtain data on the color temperature and brightness for each area of the target space. In addition, the spatial information analysis module (420) can obtain illuminance characteristic information for the target space by analyzing the captured images included in the spatial scan data in various ways.
[0109] 5. Generating training data (for learning the recognition model)
[0110] In order to train the recognition model (241) of the robot cleaner (200) to be optimized for the target space (space A), images of objects photographed while various types of objects (especially, objects with a high probability of causing a failure) are placed in multiple areas within space A are required. However, if the recognition model is trained using only images of objects photographed in an actual space (e.g., images photographing failure events), i.e., real data, it is difficult to secure a sufficient amount of real data. It takes a lot of time to secure a large amount of real data, and there is a practical limit to the amount of data that can be acquired from an actual user's space. Therefore, in one or more embodiments of the present disclosure, the server (100) may create a virtual space similar to the target space (space A), create virtual objects with a high probability of causing a failure event in the target space, and then use these to generate synthetic data to be used for training the recognition model (241).
[0111] Hereinafter, with reference to FIGS. 9 and 10, a detailed description will be given of a specific process in which a server (100) creates a virtual space and a virtual object, and creates training data through simulation of the virtual space.
[0112] (1) Creation of virtual space
[0113] As previously explained, the recognition performance of the recognition model (241) can be greatly affected by the structure of the space, the arrangement of objects within the space, the background of the space, the color of the space according to the light source, the illuminance characteristics of the space, etc. Therefore, in order to ensure that the major factors affecting the recognition performance are reflected in the training data similarly to the user's actual space (space A), the server (100) can create a virtual space based on spatial information about the user's space and illuminance characteristic information of the space.
[0114] FIG. 9 is a diagram illustrating a method for generating a virtual space by a server (100) according to one embodiment of the present disclosure. Referring to FIG. 9, the virtual space generation module (440) of the server (100) can generate a virtual space based on spatial information and illumination characteristic information received from the spatial information analysis module (420).
[0115] The first image (910) of FIG. 9 may be an image visualizing spatial information inputted into the virtual space creation module (440). That is, the first image (910) may be a result of visualizing spatial information acquired by the spatial information analysis module (420) by analyzing map data included in spatial scan data. The first image (910) expresses the overall structure of the target space (space A) and the types and positions of objects placed in the target space. That is, the spatial information inputted into the virtual space creation module (440) may include information on the structure of the target space and information on objects placed in the target space. In addition, the spatial information inputted into the virtual space creation module (440) may also include information on the background of the space (e.g., pattern, texture, color, color distribution, and color tone (white tone, wood tone, etc.) of floor or wall materials, etc.).
[0116] The spatial information input to the virtual space creation module (440) may be in the form of an image, such as the first image (910), or may be in the form of text containing information about the structure of the target space and the arrangement of objects. In addition, the spatial information input to the virtual space creation module (440) may also include information about the background of the space expressed in the form of an image (e.g., an image taken of a floor or wallpaper).
[0117] The illuminance characteristic information input to the virtual space creation module (440) is a result of the spatial information analysis module (420) analyzing the captured images included in the spatial scan data, and may be various forms of data (e.g., images or text) representing the color temperature distribution and brightness distribution in the target space. That is, the illuminance characteristic information input to the virtual space creation module (440) may be data representing the color temperature and brightness of multiple areas included in the target space.
[0118] The virtual space generation module (440) can generate a virtual space based on input spatial information and illumination characteristic information, and output virtual space data representing the generated virtual space. The second image (920) and the third image (930) of FIG. 9 are images visualizing the process of the virtual space generation module (440) generating a virtual space.
[0119] Referring to the second image (920), the virtual space generation module (440) can generate a structure of a virtual space based on the input spatial information. As described above, the spatial information may include information on the structure of the target space. Accordingly, the virtual space generation module (440) can generate a structure of a virtual space by reflecting the structure of the target space. For example, the virtual space generation module (440) can generate a structure of a virtual space identical to the structure of the target space. Alternatively, the virtual space generation module (440) can generate a structure of a virtual space by maintaining the structural characteristics of the target space (e.g., the shape of the overall spatial structure, etc.) while making changes in detailed parts.
[0120] According to one embodiment of the present disclosure, the virtual space generation module (440) can also generate the structure of the virtual space using a realistic 3D model. In this way, when a virtual space is generated using a realistic 3D model, it has the advantage of allowing for various changes to the characteristics of the space, the shooting time, etc. during the process of acquiring training data (synthetic data) through simulation.
[0121] Referring to the third image (930), the virtual space creation module (440) can place objects in a virtual space whose structure is determined. As described above, the spatial information may include information on the types and locations of objects placed in the target space. Accordingly, the virtual space creation module (440) can place objects in the virtual space by reflecting the types and locations of objects placed in the target space. For example, the virtual space creation module (440) can place objects in the virtual space in the same manner as in the target space. Alternatively, the virtual space creation module (440) can place objects by changing detailed parts while maintaining the characteristics of the object placement within the target space (e.g., at least one room has a desk and a desktop is placed on the desk).
[0122] When the structure of the virtual space is created and the arrangement of objects in the virtual space is completed, the virtual space creation module (440) can adjust the illumination of the virtual space based on the illumination characteristic information. The illumination adjustment of the virtual space may also be performed in the simulation process for generating training data described below. According to one embodiment of the present disclosure, the virtual space creation module (440) determines an area of a corresponding target space for each of a plurality of areas included in the virtual space, and applies the color temperature and brightness of the area of the corresponding target space as is to the virtual space, or adjusts the color temperature and brightness of the area of the target space by a certain ratio (e.g., 20% to 50%) and applies them to the virtual space.
[0123] The virtual space generation module (440) can also adjust the background of the virtual space by reflecting information about the background of the target space included in the spatial information. For example, the virtual space generation module (440) can generate a pattern and texture of the floor or wallpaper of the virtual space that is identical or similar to that of the target space.
[0124] As a result, the virtual space generation module (440) can generate a virtual space by applying the same layout of the target space that can be identified from the spatial information, or can generate a virtual space with slightly modified arrangement of structures or objects while maintaining the characteristics of the layout of the target space (e.g., many pieces of furniture placed in the living room, wood patterns on the floor, etc.). In addition, the virtual space generation module (440) can provide various illumination changes to the virtual space based on illumination characteristic information. In addition, the virtual space generation module (440) can adjust the background of the virtual space in various ways based on information about the background of the target space.
[0125] As described above, the server (100) creates a virtual space based on spatial information of an actual target space, so realistic synthetic data can be created and used as training data.
[0126] (2) Creation of virtual objects
[0127] As described above, the virtual object generation model (450) can be said to learn about the relationship between the characteristics of a space (e.g., the structure of the space, the arrangement of objects within the space, etc.) and the objects that cause failure. Therefore, when spatial information about the target space (the structure of the space, the types and locations of objects arranged within the space) and information about a failure event (the type and location of the object that caused the failure event) are input, the virtual object generation model (450) can generate a new object (virtual object) that is likely to cause failure in the target space. In other words, the virtual object generation model (450) can generate and output data about a virtual object that is likely to cause failure according to context such as the structure of the target space and the arrangement of objects within the target space. According to one embodiment of the present disclosure, the virtual object may be an object that is likely to fail to be recognized by the recognition model (241) of the robot cleaner (200) in the target space above a preset standard.
[0128] Referring to FIG. 10, the process of generating a virtual object by the virtual object generation model (450) will be described in detail. The first image (1010) of FIG. 10 is an image visualizing the input of the virtual object generation model (450). The first image (1010) shows the structure of the target space (space A) and the types and locations of objects placed in the target space. In addition, the first image (1010) shows the types and locations of objects (11, 12, 13) that caused a failure event. As illustrated in FIG. 4, the virtual object generation model (450) may receive spatial information of the target space output by the spatial information analysis module (420) and the type and location of the failure event output by the log data analysis module (430).
[0129] As previously described with reference to FIG. 6, information included in the first image (1010) (spatial information of the target space and the type and location of the failure event) can be input into the virtual object creation model (450) in image form or vectorized. Specifically, the following three types of data can be input into the virtual object creation model (450).
[0130] 1) 2D top-view image of the target space (image with objects and objects that caused failure events (11, 12, 13) removed from the first image (1010))
[0131] 2) A vector representing the type, location (2D coordinate information), and size of objects (appliances, furniture) shown in the first image (1010).
[0132] 3) A vector indicating the type and location (2D coordinate information) of objects (11, 12, 13) that caused the failure event shown in the first image (1010).
[0133] The second image (1020) of Fig. 10 is an image visualizing the output of the virtual object generation model (450). Compared to the first image (1010), a new object (14) has been added to the second image (1020). The virtual object generation model (450) learns the relationship between the objects (11, 12, 13) that caused the failure and the structure of the space and objects arranged in the space (e.g., home appliances, furniture, etc.), and based on the results, can generate a new object (14) according to the context of the target space.
[0134] According to one embodiment of the present disclosure, virtual object data actually output from the virtual object generation model (450) may include the type (excrement) and location (2D coordinate information) of the newly created virtual object (14). In addition, the virtual object data may further include context information, which is information on surrounding objects related to the newly created virtual object (14) (e.g., the excrement is located around the TV). The second image (1020) of FIG. 10 corresponds to the result of adding the virtual object (14) to the 2D top-view image of the target space according to the virtual object data as described above.
[0135] According to one embodiment of the present disclosure, the location of the virtual object (14) included in the virtual object data may be determined based on context information included in the virtual object data. For example, in the second image (1020) of FIG. 10, the location of the virtual object (14) may be determined to be around the TV based on context information indicating that "the feces are located around the TV."
[0136] In summary, the virtual object creation model (450) can create virtual object data (type and location of virtual object, context information) when spatial information about the target space and information about a failure event are input.
[0137] That is, the virtual object creation model (450) can understand in which situations failure is likely to occur and create similar situations (situations in which a specific type of virtual object is placed around specific objects in space).
[0138] In other words, the virtual object generation model (450) can learn to understand where in the target space certain objects are likely to be placed, and further, where in the target space certain objects are likely to fail to be recognized by the recognition model, based on the structure of the target space and the arrangement of objects (appliances, furniture) within the target space, and can generate virtual objects accordingly.
[0139] (3) Augmentation of virtual objects in virtual space
[0140] The augmentation module (460) can augment a virtual object in a certain area within a virtual space based on virtual space data and virtual object data.
[0141] As described above, the virtual object data may include information on the location of the virtual object, and the location of the virtual object may be expressed as coordinate information on a 2D top-view image of the target space. In addition, as described above, the virtual space generation module (440) may generate a structure of the virtual space that is identical or similar to the structure of the target space. Therefore, according to one embodiment of the present disclosure, the augmentation module (460) may identify a location in the virtual space corresponding to a location on the 2D top-view image of the target space, and may place a virtual object at the identified location. For example, the augmentation module (460) may identify a location in the virtual space corresponding to the location of the virtual object (14) on the second image (1020) of FIG. 10, and place the virtual object (feces) at the identified location.
[0142] As described above, virtual object data may include context information, which is information about objects that may be located around the virtual object. Furthermore, as described above, the virtual space generation module (440) may place objects in the virtual space by reflecting the types and locations of objects placed in the target space. Therefore, according to one embodiment of the present disclosure, the augmentation module (460) may determine a location where a virtual object is to be placed in the virtual space based on objects placed in the virtual space and context information included in the virtual object data. In other words, the augmentation module (460) may determine objects to be located around the virtual object among objects placed in the virtual space according to the context information, and place the virtual object around the determined objects. For example, if the type of the generated virtual object is 'electrical wire' and the context information is 'desk and computer', the augmentation module (460) may place a virtual object (electrical wire) around the locations where the desk and the computer are placed in the virtual space.
[0143] In summary, according to one embodiment of the present disclosure, the augmentation module (460) can determine a location where a virtual object is to be placed in a virtual space based on location information of the virtual object included in the virtual object data, and can place the virtual object at the determined location. Furthermore, according to one embodiment of the present disclosure, the augmentation module (460) can determine a location where a virtual object is to be placed in a virtual space based on context information of objects and virtual objects placed in the virtual space, and can place the virtual object at the determined location.
[0144] (4) Generation of training data through simulation
[0145] When a virtual object is augmented in a virtual space, the training data generation module (470) can generate training data through a simulation using a virtual robot cleaner. According to one embodiment of the present disclosure, the training data generation module (470) can execute a simulation so that the virtual robot cleaner performs a cleaning operation or a scanning operation (e.g., taking pictures with a camera while driving) in a virtual space where a virtual object is placed. The training data generation module (470) can generate training data for training the recognition model (241) using an image taken by the virtual robot cleaner of a virtual object in a virtual space (hereinafter, referred to as a 'virtual shooting image' or 'synthetic image').
[0146] Below, the process of generating training data by the training data generation module (470) is described in detail with reference to FIG. 11.
[0147] The first image (1110) of FIG. 11 is an image showing a situation in which a virtual robot cleaner (1111) is placed in a virtual space created by the virtual space creation module (440) of the server (100). The training data creation module (470) can place the virtual robot cleaner (1111) in the virtual space for simulation. Referring to the first image (1110) of FIG. 11, the virtual space creation module (440) can change the lighting or texture rendering of the virtual space in various ways before executing the simulation. As described above, the virtual space creation module (440) can adjust the illumination of the virtual space based on the illumination characteristic information of the target space. In addition, the virtual space creation module (440) can adjust the background of the virtual space (e.g., the pattern, texture, color, color distribution, and color tone (white tone, wood tone, etc.) of the floor or wallpaper, etc.) based on the information about the background of the space included in the spatial information. By changing the lighting or texture rendering of the virtual space using the virtual space generation module (440), various environments that can appear in an actual target space can be created in the virtual space.
[0148] The second image (1120) of Fig. 11 is an image for explaining the process of augmenting a virtual object (14) in a virtual space by the augmentation module (460). As described above, the virtual object data output by the virtual object creation model (450) may include information on the location of the virtual object (14) expressed as coordinates on a 2D top-view image of the target space, and thus the augmentation module (460) may find and place the location of the virtual object (14) in a virtual space implemented identically or similarly to the target space.
[0149] Alternatively, the augmentation module (460) can augment a virtual object in a virtual space according to various methods described above.
[0150] The third image (1130) of FIG. 11 is an image illustrating a situation in which a virtual robot cleaner (1111) acquires a virtual photographed image (synthetic data) in a state in which a virtual object (14) is augmented in a virtual space. The training data generation module (470) can execute a simulation so that the virtual robot cleaner (1111) performs a cleaning operation or a scanning operation in the virtual space in which the virtual object (14) is placed. When the simulation is executed, the virtual robot cleaner (1111) can acquire a virtual photographed image (1132) by performing photographing while driving in the virtual space. The training data generation module (470) can generate training data by labeling the type (feces) of the virtual object (14) in the virtual photographed image (1132).
[0151] In this way, the training data generation module (470) can efficiently obtain training data for various situations by generating synthetic data (virtual captured images) using virtual space and generating training data using synthetic data.
[0152] According to one embodiment of the present disclosure, the training data generation module (470) may select some of the synthetic images collected through simulation to generate training data. For example, the training data generation module (470) may perform a test using the recognition model (490) on the synthetic image and determine whether to use the synthetic image as training data based on the test result. The recognition model (490) used for the test may be the same model as the recognition model (241) installed in the robot cleaner (200), or may be a basic recognition model commonly installed in all units (robot cleaners) upon shipment.
[0153] In detail, the training data generation module (470) can attempt object recognition using the recognition model (490) for synthetic images acquired through simulation and determine whether recognition is successful. The training data generation module (470) can generate training data by adjusting the ratio of synthetic images in which the recognition model (490) succeeds in object recognition and synthetic images in which the recognition model (490) fails to recognize the object to a preset value. For example, the training data generation module (470) can perform an object recognition test using the recognition model (490) on the initially collected synthetic images and select synthetic images such that the number of synthetic images in which the test result is 'success' and the number of synthetic images in which the test result is 'failure' are in a preset constant ratio (e.g., 1:1 or 10:1). Subsequently, the training data generation module (470) can generate training data by labeling the types of virtual objects included in the selected synthetic images.
[0154] In other words, the training data generation module (470) determines whether to use each synthetic image as training data through an object recognition test using a recognition model (490), and if it is decided to use the synthetic image as training data, the training data can be generated by labeling the type of virtual object in the synthetic image.
[0155] If only synthetic images in which the recognition model (490) succeeds in object recognition are included in the training data, or conversely, if only synthetic images in which the recognition model (490) fails in object recognition are included in the training data, the problem of overfitting or underfitting may occur. The training data generation module (470) can secure data balance by adjusting the number of synthetic images with a test result of 'success' and the number of synthetic images with a test result of 'failure' at a certain ratio, thereby increasing the learning effect of the recognition model (490).
[0156] According to one embodiment of the present disclosure, when a failure event occurs while a virtual robot vacuum cleaner (1111) performs a cleaning operation in a virtual space, the training data generation module (470) may generate training data using a synthetic image captured at the failure event.
[0157] Meanwhile, the training data generation module (470) may also generate training data using captured images included in the failure log data. Such real-data-based training data and synthetic-data-based training data may be used together to train the recognition model (490).
[0158] 6. Fine tuning of the recognition model using the generated training data.
[0159] FIG. 12 is a diagram illustrating an operation in which a server according to one embodiment of the present disclosure performs fine tuning on a recognition model using training data and updates a recognition model of a robotic mobile device based on the result.
[0160] Referring to FIG. 12, the fine tuning module (480) of the server (100) can fine tune the recognition model (490) using the training data (1200) output from the training data generation module (470). At this time, the recognition model (490) may be the same recognition model as the recognition model (241) installed in the robot vacuum cleaner (200) or may be a recognition model commonly used for all objects.
[0161] According to one embodiment of the present disclosure, the fine tuning module (480) can update the parameters of the recognition model (490) by further training the recognition model (490) in a supervised learning manner using training data labeled with the correct answer (type of object).
[0162] 7. Update the robot vacuum cleaner's recognition model based on the fine-tuned recognition model.
[0163] Once fine tuning (additional learning) for the recognition model (490) is completed, the server (100) can request an update of the recognition model (241) by transmitting updated parameter information of the recognition model (490) to the robot cleaner (200) via the communication module (410). The robot cleaner (200) can update the recognition model (241) according to the received parameter information.
[0164] Hereinafter, with reference to FIGS. 13 to 18, a method for updating a recognition model of a robotic mobile device according to one or more embodiments of the present disclosure will be described. The steps included in the flowcharts of FIGS. 13 to 18 are performed by the server (100) of FIGS. 1 and 2 , and therefore, the contents previously described with reference to FIGS. 1 to 12 may be equally applied to FIGS. 13 to 18 even if omitted below.
[0165] Referring to FIG. 13, in step 1301, the server (100) can obtain spatial scan data for a target space from a robotic mobile device. According to one embodiment of the present disclosure, the robotic mobile device can ask the user whether to update the recognition model when first started or when reset, and if the user agrees, can start updating the recognition model. According to one embodiment of the present disclosure, when a failure event of the robotic mobile device occurs, the recognition model update can also start automatically or with the user's consent. When the update of the recognition model of the robotic mobile device starts, the robotic mobile device can collect spatial scan data by scanning the target space using the provided sensors (such as a camera and a lidar) and transmit the spatial scan data to the server (100). The spatial scan data can include map data (e.g., lidar scan data) collected through the lidar sensor of the robotic mobile device and captured images acquired through the camera.
[0166] In step 1302, the server (100) can acquire spatial information including information about the structure of the target space and objects placed in the target space based on the spatial scan data. Detailed steps included in step 1302 are illustrated in FIG. 14.
[0167] Referring to FIG. 14, in step 1401, the server (100) can analyze the structure of the target space based on the map data included in the space scan data.
[0168] At step 1402, the server (100) can analyze the type and location of objects placed in the target space based on map data and captured images included in the spatial scan data.
[0169] According to one embodiment of the present disclosure, the server (100) may analyze the background of the target space based on the captured image included in the space scan data.
[0170] According to one embodiment of the present disclosure, the server (100) may obtain illumination characteristic information for a target space based on spatial scan data. For example, the server (100) may determine color temperature distribution and brightness distribution in the target space by analyzing captured images included in the spatial scan data.
[0171] Returning to FIG. 13 again, in step 1303, the server (100) may input spatial information into a generation model (virtual object generation model) to obtain virtual object data including information on the type and location of the virtual object. According to one embodiment of the present disclosure, the generation model is a neural network model trained using the type of failure event that occurred in the target space, the location where the failure event occurred, and spatial information of the target space, and the failure event may be an event in which the recognition model of the robotic mobile device fails to recognize an object. In addition, the virtual object may be an object in the target space in which the recognition model of the robotic mobile device has a high probability of failing to recognize above a preset standard.
[0172] According to one embodiment of the present disclosure, the generative model learns about the relationship between spatial features (e.g., spatial structure, arrangement of objects within the space, etc.) and objects that cause failures, so that when spatial information is input, data about virtual objects that are likely to cause failures can be generated and output. For details on the learning and inference (generation of virtual objects) of the generative model, refer to the parts described above with reference to FIGS. 6 and 10. Meanwhile, according to one embodiment of the present disclosure, the virtual object data may further include context information, which is information about surrounding objects related to the virtual object.
[0173] In step 1304, the server (100) can acquire training data using spatial information and virtual object data. The detailed steps included in step 1304 are illustrated in FIG. 15.
[0174] Referring to FIG. 15, in step 1501, the server (100) can create a virtual space based on spatial information. Detailed steps included in step 1501 are illustrated in FIG. 16.
[0175] Referring to FIG. 16, in step 1601, the server (100) can determine the structure of the virtual space and the types and positions of objects placed in the virtual space based on spatial information. The spatial information may include information about the structure of the target space, and the server (100) can generate the structure of the virtual space by reflecting the structure of the target space. For example, the server (100) can generate the structure of the virtual space in the same manner as the structure of the target space, or can generate the structure of the virtual space by maintaining the structural characteristics of the target space while making changes in the details. The spatial information may include information about the types and positions of objects placed in the target space, and the server (100) can place objects in the virtual space by reflecting the types and positions of objects placed in the target space. For example, the server (100) can place objects in the same manner as the target space, or can place objects by making changes in the details while maintaining the characteristics of the objects placed in the target space.
[0176] In step 1602, the server (100) can determine the illuminance for each of multiple areas of the virtual space based on the illuminance characteristic information for the target space. According to one embodiment of the present disclosure, the server (100) determines the corresponding area of the target space for each of the multiple areas included in the virtual space, and applies the color temperature and brightness of the corresponding area of the target space as is to the virtual space, or adjusts the color temperature and brightness of the area of the target space by a certain ratio and applies them to the virtual space.
[0177] According to one embodiment of the present disclosure, the server (100) may adjust the background of the virtual space by reflecting information about the background of the target space included in the spatial information.
[0178] Returning to FIG. 15, in step 1502, the server (100) can augment a virtual object in the virtual space. The detailed steps included in step 1502 are illustrated in FIG. 17.
[0179] Referring to FIG. 17, in step 1701, the server (100) determines a location where a virtual object is to be placed based on objects and context information placed in a virtual space, and then in step 1702, the server (100) can place the virtual object at the determined location.
[0180] The virtual object data may include information regarding the location of the virtual object, and the location of the virtual object may be expressed as coordinate information on a 2D top-view image of the target space. Furthermore, the structure of the virtual space may be created to be identical or similar to the structure of the target space. Therefore, according to one embodiment of the present disclosure, the server (100) may identify a location in the virtual space corresponding to a location on the 2D top-view image of the target space, and may place a virtual object at the identified location.
[0181] Virtual object data may also include contextual information, which is information about objects likely to be located around the virtual object. Accordingly, the server (100) may determine the location where the virtual object will be placed in the virtual space based on the objects placed in the virtual space and the contextual information included in the virtual object data.
[0182] Returning to FIG. 15 again, at step 1503, the server (100) can obtain a composite image of a virtual object captured within the virtual space by performing a simulation of the virtual space. According to one embodiment of the present disclosure, the server (100) can obtain an image (composite image) of a virtual object captured by the virtual robotic mobile device by executing a simulation in which the virtual robotic mobile device performs a scanning operation (e.g., taking pictures with a camera while moving) in the virtual space.
[0183] At step 1504, the server (100) can generate training data using the synthetic image. According to one embodiment of the present disclosure, the server (100) can generate training data by labeling the type of virtual object for the synthetic image.
[0184] Furthermore, according to one embodiment of the present disclosure, the server (100) may perform object recognition on a synthetic image using a recognition model installed in the server (100), determine whether object recognition is successful, and then determine whether to utilize the synthetic image as training data based on the determination result. If it is determined to utilize the synthetic image as training data, the server (100) may generate training data by labeling the type of virtual object in the synthetic image.
[0185] Returning to Figure 13, at step 1305, the server (100) can update the recognition model of the robotic mobile device using the training data. The detailed steps included in step 1305 are illustrated in Figure 18.
[0186] Referring to FIG. 18, in step 1801, the server (100) can perform fine tuning on a recognition model included in an electronic device (server (100)) using training data.
[0187] At step 1802, the server (100) may request an update of the recognition model by transmitting parameter information of the fine-tuned recognition model to the robotic mobile device. The robotic mobile device may update the recognition model installed in the robotic mobile device based on the received parameter information.
[0188] According to one or more embodiments of the present disclosure described above, by updating the recognition model of a robotic mobile device to optimize it for the target space in which it is used, high recognition performance can be expected in that space. Furthermore, according to one or more embodiments of the present disclosure, by acquiring training data using virtual objects generated to reflect actual failure events, the effect of efficiently acquiring training data for various situations can be expected.
[0189] A method for updating a recognition model of a robotic mobile device according to one embodiment of the present disclosure may include a step in which an electronic device obtains spatial scan data for a target space from the robotic mobile device, a step in which the electronic device obtains spatial information including information on a structure of the target space and an object placed in the target space based on the spatial scan data, a step in which the electronic device inputs the spatial information into a generative model to obtain virtual object data including information on a type and a location of a virtual object, a step in which the electronic device obtains training data using the spatial information and the virtual object data, and a step in which the electronic device updates the recognition model of the robotic mobile device using the training data.
[0190] According to one embodiment, the space scan data may include map data obtained by scanning the target space using a LiDAR sensor and a photographed image taken of the target space using a camera.
[0191] According to one embodiment, the step of obtaining the spatial information may include a step in which the electronic device analyzes the structure of the target space based on the map data, and a step in which the electronic device analyzes the type and location of an object placed in the target space based on the map data and the photographed image.
[0192] According to one embodiment, the generation model is a neural network model learned using the type of failure event that occurred in the target space, the location where the failure event occurred, and spatial information of the target space, and the failure event may be an event in which the recognition model of the robotic mobile device fails to recognize an object.
[0193] According to one embodiment, the virtual object may be an object that has a higher probability of failure in recognition by the recognition model of the robotic mobile device in the target space than a preset standard.
[0194] According to one embodiment, the step of obtaining the training data may include a step in which the electronic device creates a virtual space based on the spatial information, a step in which the electronic device augments the virtual object in the virtual space, a step in which the electronic device performs a simulation of the virtual space to obtain a synthetic image of the virtual object in the virtual space, and a step in which the electronic device generates training data using the synthetic image.
[0195] According to one embodiment, the step of obtaining the spatial information may include an operation in which the electronic device obtains illuminance characteristic information for the target space based on the spatial scan data, and the step of generating the virtual space may include a step in which the electronic device determines a structure of the virtual space and types and positions of objects placed in the virtual space based on the spatial information, and a step in which the electronic device determines illuminance for each of a plurality of areas of the virtual space based on the illuminance characteristic information.
[0196] According to one embodiment, the virtual object data further includes context information, which is information about surrounding objects related to the virtual object, and the step of augmenting the virtual object may include a step in which the electronic device determines a location at which the virtual object is to be placed based on an object placed in the virtual space and the context information, and a step in which the virtual object is to be placed at the determined location.
[0197] According to one embodiment, the step of generating training data using the synthetic image may generate the training data by labeling the type of the virtual object for the synthetic image.
[0198] According to one embodiment, the step of generating training data using the synthetic image may include the step of performing object recognition on the synthetic image using the recognition model, the step of determining whether the object recognition is successful, the step of determining whether to use the synthetic image as the training data based on the determination result, and the step of generating the training data by labeling the type of the virtual object in the synthetic image if it is determined to use the synthetic image as the training data.
[0199] According to one embodiment, the step of updating the recognition model of the robotic mobile device may include a step in which the electronic device performs fine tuning on the recognition model included in the electronic device using the training data, and a step in which parameter information of the recognition model on which the fine tuning has been performed is transmitted to the robotic mobile device while requesting an update of the recognition model.
[0200] An electronic device for updating a recognition model of a robotic mobile device according to one embodiment of the present disclosure includes a memory storing a program for updating the recognition model and at least one processor, and the at least one processor executes the program stored in the memory or at least one instruction, whereby the electronic device obtains spatial scan data for a target space from the robotic mobile device, and the electronic device obtains spatial information including information on a structure of the target space and an object arranged in the target space based on the spatial scan data, and the electronic device inputs the spatial information into a generative model to obtain virtual object data including information on a type and a position of a virtual object, and after the electronic device obtains training data using the spatial information and the virtual object data, the electronic device can update the recognition model of the robotic mobile device using the training data.
[0201] According to one embodiment, the space scan data may include map data obtained by scanning the target space using a LiDAR sensor and a photographed image taken of the target space using a camera.
[0202] According to one embodiment, when the electronic device obtains the spatial information, the electronic device may analyze the structure of the target space based on the map data, and then analyze the type and location of an object placed in the target space based on the map data and the photographed image.
[0203] According to one embodiment, the generation model is a neural network model learned using the type of failure event that occurred in the target space, the location where the failure event occurred, and spatial information of the target space, and the failure event may be an event in which the recognition model of the robotic mobile device fails to recognize an object.
[0204] According to one embodiment, the virtual object may be an object that has a higher probability of failure in recognition by the recognition model of the robotic mobile device in the target space than a preset standard.
[0205] According to one embodiment, when the electronic device acquires the training data, the electronic device creates a virtual space based on the spatial information, augments the virtual object in the virtual space, and performs a simulation of the virtual space, thereby acquiring a synthetic image of the virtual object within the virtual space, and then the electronic device can generate training data using the synthetic image.
[0206] According to one embodiment, the electronic device obtains illuminance characteristic information for the target space based on the spatial scan data, and when generating the virtual space, the electronic device determines the structure of the virtual space and the types and positions of objects placed in the virtual space based on the spatial information, and then the electronic device determines illuminance for each of a plurality of areas of the virtual space based on the illuminance characteristic information.
[0207] According to one embodiment, the virtual object data further includes context information, which is information about surrounding objects related to the virtual object, and when the electronic device augments the virtual object, the electronic device can determine a location where the virtual object is to be placed based on an object placed in the virtual space and the context information, and then place the virtual object at the determined location.
[0208] One or more embodiments of the present disclosure may be implemented or supported by one or more computer programs, which may be formed from computer-readable program code and embodied in a computer-readable medium. In the present disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive (HDD), a compact disc (CD), a digital video disc (DVD), or various types of memory.
[0209] Additionally, a device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a 'non-transitory storage medium' is a tangible device and may exclude wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Meanwhile, this 'non-transitory storage medium' does not distinguish between cases where data is permanently stored in the storage medium and cases where it is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer where data is temporarily stored. A computer-readable medium may be any available medium that can be accessed by a computer, and may include both volatile and non-volatile media, and removable and non-removable media. A computer-readable medium includes a medium on which data can be permanently stored and a medium on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable memory device.
[0210] According to one embodiment, a method according to one or more embodiments disclosed herein may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0211] The above description of the present disclosure is for illustrative purposes only, and those skilled in the art will appreciate that the present disclosure can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present disclosure. For example, suitable results can be achieved even if the described techniques are performed in a different order than the described method, and / or components of the systems, structures, devices, circuits, etc. described are combined or combined in a different form than the described method, or are replaced or substituted by other components or equivalents. Therefore, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. For example, each component described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined form.
[0212] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A method for updating a recognition model of a robotic mobile device, A step in which an electronic device (100) obtains spatial scan data for a target space from a robotic mobile device (200); A step in which the electronic device (100) acquires spatial information including information on the structure of the target space and objects within the target space based on the spatial scan data; A step in which the electronic device (100) inputs the spatial information into a generative model to obtain virtual object data including information on the type of virtual object and the location of the virtual object; A step in which the electronic device (100) obtains training data using the spatial information and the virtual object data; and A method, comprising the step of updating a recognition model of the robotic mobile device (200) using the training data by the electronic device (100).
2. In paragraph 1, The above space scan data is, A method characterized by including map data obtained by scanning the target space using a LiDAR sensor (260) and a photographed image taken of the target space using a camera (250).
3. In either of paragraphs 1 and 2, The step of obtaining the above spatial information is: A step in which the electronic device (100) analyzes the structure of the target space based on the map data; and A method characterized in that the electronic device (100) includes a step of analyzing the type and location of an object within the target space based on the map data and the photographed image.
4. In any one of paragraphs 1 to 3, The above generation model includes a neural network model learned using the type of failure event that occurred in the target space, the location where the failure event occurred, and spatial information of the target space. A method characterized in that the above failure event includes an event in which the recognition model fails to recognize an object.
5. In any one of paragraphs 1 to 4, A method characterized in that the virtual object is an object that the recognition model has a higher probability of failing to recognize than a preset standard in the target space.
6. In any one of paragraphs 1 to 5, The step of obtaining the above training data is: A step in which the electronic device (100) creates a virtual space based on the spatial information; A step of augmenting the virtual object in the virtual space by the electronic device (100); A step of obtaining a synthetic image of the virtual object in the virtual space by having the electronic device (100) perform a simulation of the virtual space; and A method characterized in that the electronic device (100) includes a step of generating training data using the synthetic image.
7. In any one of paragraphs 1 to 6, The step of obtaining the above spatial information is: The electronic device (100) includes an operation of obtaining illuminance characteristic information for the target space based on the space scan data, The steps for creating the above virtual space are: A step in which the electronic device (100) determines the structure of the virtual space and the type and location of objects placed in the virtual space based on the spatial information; and A method characterized in that the electronic device (100) includes a step of determining illuminance for each of a plurality of areas of the virtual space based on the illuminance characteristic information.
8. In any one of paragraphs 1 to 7, The above virtual object data further includes context information, which is information about surrounding objects related to the virtual object. The step of augmenting the above virtual object is: A step for determining a location where the virtual object is to be placed in the virtual space based on an object placed in the virtual space and the context information of the electronic device (100); and A method characterized by comprising a step of placing the virtual object at the determined location in the virtual space.
9. In any one of paragraphs 1 to 8, The step of generating training data using the above synthetic image is as follows: A method characterized by labeling the type of the virtual object for the synthetic image.
10. In any one of paragraphs 1 to 9, The step of generating training data using the above synthetic image is as follows: A step of performing object recognition on the synthetic image using the above recognition model; A step for determining whether the above object recognition is successful; A step of determining whether to use the synthetic image as training data based on the result of determining whether the object recognition is successful; and A method characterized by comprising a step of generating the training data by labeling the type of the virtual object in the synthetic image based on a decision to utilize the synthetic image as the training data.
11. In an electronic device (100) for updating a recognition model of a robotic mobile device (200), A memory (130) in which a program or at least one instruction is stored; and At least one processor (120) operably connected to the above memory (130), The electronic device (100) executes the program or the at least one instruction by the at least one processor (120). Obtain spatial scan data for a target space from a robot-type mobile device (200), Based on the above spatial scan data, spatial information including information about the structure of the target space and objects within the target space is obtained, By inputting the above spatial information into a generative model, virtual object data including information on the type of virtual object and the location of the virtual object is obtained, After obtaining training data using the above spatial information and the above virtual object data, An electronic device that updates the recognition model of the robotic mobile device (200) using the training data.
12. In paragraph 11, The above space scan data is, An electronic device characterized by including map data acquired by scanning the target space using a LiDAR sensor (260) and a photographed image captured by using a camera (250) of the target space.
13. In either of paragraphs 11 and 12, The electronic device (100) obtains the spatial information by executing the program or the at least one instruction by the at least one processor (120). After analyzing the structure of the target space based on the above map data, An electronic device characterized by analyzing the type and location of an object within the target space based on the map data and the photographed image.
14. In any one of paragraphs 11 to 13, The above generation model includes a type of failure event that occurred in the target space, a location where the failure event occurred, and a neural network model learned using the spatial information. An electronic device, characterized in that the above failure event includes an event in which the recognition model fails to recognize an object.
15. In any one of paragraphs 11 to 14, The electronic device (100) obtains the training data by executing the program or the at least one instruction by the at least one processor (120). Create a virtual space based on the above spatial information, Augmenting the virtual object in the virtual space, By performing a simulation of the above virtual space, a synthetic image of the virtual object captured within the above virtual space is obtained, An electronic device characterized by generating training data using the above synthetic image.
Citation Information
Patent Citations
Method for preparing of cyclic amidine compounds using borane catalyst and cyclic amidine compounds prepared therefrom
KR1020210120242A
Method of encoding and decoding signal, and device of encoding and decoding signal perporming the methods
KR1020230146860A
Environment Learning Device and Method in Charging Time for Autonomous Mobile Robot
KR102301759B1
An artificial intelligence robot for cleaning using zoned pollution information and method for the same
KR102331563B1
TURNTABLE SYSTEM for FLOW CYTOMETRY of MICROALGAE and METHOD for FLOW CYTOMETRY of MICROALGAE USING the same
KR102543482B1