A method for updating the recognition model of a robotic mobile device and an electronic device for performing the method.

By generating virtual spaces and virtual objects, and using generative AI models to train data to update the recognition model, the problem of poor recognition performance of robotic mobile devices in specific spaces is solved, thereby improving recognition accuracy and operational stability.

CN122497569APending Publication Date: 2026-07-31SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2024-12-27
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

The recognition performance of a robot's mobile device recognition model in a specific space is affected by the object type and environment, leading to recognition failures, and it is difficult to optimize the model by collecting enough training data in that space.

Method used

By generating virtual spaces and virtual objects, generative AI models are used to generate training data and update the recognition model to improve recognition performance in specific spaces.

Benefits of technology

It improves the recognition accuracy of robot mobile devices in specific spaces, reduces the occurrence of recognition failures, and enhances operational stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497569A_ABST
    Figure CN122497569A_ABST
Patent Text Reader

Abstract

A method for updating a recognition model of a robotic mobile device may include the following steps: an electronic device obtains spatial scan data of a target space from the robotic mobile device; the electronic device obtains spatial information based on the spatial scan data, wherein the spatial information includes information about the structure of the target space and objects in the target space; the electronic device inputs the spatial information into a generative model to obtain virtual object data, wherein the virtual object data includes information about the type and location of the virtual objects; the electronic device obtains training data by using the spatial information and the virtual object data; and the electronic device updates the recognition model of the robotic mobile device by using the training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method for updating a recognition model of a robot mobile device and an electronic device for performing the method. Background Technology

[0002] Neural network models for object recognition (hereinafter referred to as "recognition models") are widely used in electronic devices commonly used in everyday life. For example, electronic devices that move within a specific space (e.g., a house) and perform operations (such as robotic vacuum cleaners) can use recognition models to identify surrounding objects and perform operations based on the recognition results.

[0003] The recognition performance of a recognition model may be affected by the type or category of the object being recognized or the surrounding environment (e.g., the structure of the space, the brightness of the space, nearby objects, etc.). Therefore, the recognition performance of a recognition model may vary depending on the space in which the recognition model is used. Summary of the Invention

[0004] Technical solution According to one aspect of this disclosure, a method for updating a recognition model of a robot mobile device includes: obtaining spatial scanning data about a target space from the robot mobile device by an electronic device; obtaining spatial information based on the spatial scanning data by the electronic device, wherein the spatial information includes information about the structure of the target space and objects in the target space; obtaining virtual object data by the electronic device by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category and location of the virtual object; obtaining training data by the electronic device using the spatial information and the virtual object data; and updating the recognition model of the robot mobile device by the electronic device using the training data.

[0005] According to one aspect of this disclosure, an electronic device for updating a recognition model of a robotic mobile device includes: a memory, a stored program or at least one instruction; and at least one processor configured to run the program or at least one instruction, wherein the program or at least one instruction, when run by the at least one processor, causes the electronic device to perform the following operations: obtaining spatial scan data about a target space from the robotic mobile device; obtaining spatial information based on the spatial scan data, wherein the spatial information includes information about the structure of the target space and objects in the target space; obtaining virtual object data by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category and location of the virtual object; obtaining training data by using the spatial information and the virtual object data; and updating the recognition model of the robotic mobile device using the training data. Attached Figure Description

[0006] The above and other aspects and features of specific embodiments of this disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which: Figure 1 This is a diagram illustrating a system environment for updating the recognition model of a robot mobile device according to an embodiment of the present disclosure; Figure 2 This is a diagram illustrating hardware components included in a server according to embodiments of the present disclosure; Figure 3 This is a diagram illustrating components included in a robotic vacuum cleaner according to an embodiment of the present disclosure; Figure 4 This is a diagram illustrating detailed components of a server according to embodiments of the present disclosure, differentiated based on their function or role; Figure 5 This is a diagram illustrating the operation of a robot mobile device sending failure log data to a server according to an embodiment of the present disclosure; Figure 6 This is a diagram illustrating the process of performing training and inference in a virtual object generation model on a server according to an embodiment of the present disclosure: Figure 7 This is a diagram illustrating an example of training data for training a virtual object generation model included in a server, according to an embodiment of the present disclosure. Figure 8 This is a diagram illustrating the operation of a robotic mobile device according to an embodiment of the present disclosure, which obtains spatial scan data about a target space and sends the spatial scan data to a server; Figure 9 This is a diagram illustrating a method for generating a virtual space performed by a server according to an embodiment of the present disclosure; Figure 10 This is a diagram illustrating the process of generating virtual objects according to an embodiment of the present disclosure using a virtual object generation model; Figure 11 This is a diagram illustrating the operation of a server adding virtual objects in a virtual space and generating training data through simulation according to an embodiment of the present disclosure; Figure 12 This is a diagram illustrating the operation of a server according to an embodiment of the present disclosure, which performs fine-tuning on a recognition model using training data and updates the recognition model of a robot mobile device based on the fine-tuning results; and Figure 13 , Figure 14 , Figure 15 , Figure 16 , Figure 17 and Figure 18 This is a flowchart illustrating a method for updating the recognition model of a robotic mobile device according to an embodiment of the present disclosure. Detailed Implementation

[0007] Throughout this disclosure, the expression "at least one of a, b, or c" means any one of the following: only a, only b, only c, both a and b, both a and c, both b and c, or all of a, b, and c.

[0008] In describing this disclosure, descriptions of technical concepts well-known in the art to which this disclosure pertains and not directly related to this disclosure will be omitted. This is to convey the essence of this disclosure more clearly without obscuring it by omitting unnecessary descriptions. Furthermore, the terminology used below is defined by consideration of the functions described in this disclosure and may be changed according to the intent, practice, etc., of the user or operator. Therefore, the terminology should be defined based on the overall description of this disclosure.

[0009] For the same reason, some components are shown enlarged, omitted, or schematically in the accompanying drawings. Furthermore, the dimensions of each component do not perfectly reflect the actual dimensions. Throughout the drawings, the same reference numerals refer to the same or corresponding elements.

[0010] The advantages and features of this disclosure, as well as methods of implementing it, will be more readily understood by referring to the following description and accompanying drawings of one or more embodiments of this disclosure. However, this disclosure may be implemented in many different forms and should not be construed as limited to the embodiments of this disclosure set forth below. Rather, one or more embodiments of this disclosure are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art to which this disclosure pertains. Embodiments of this disclosure may be defined by the appended claims. Throughout the specification, the same reference numerals refer to the same elements. Furthermore, in the following description of this disclosure, relevant functions or configurations will not be described in detail where it is determined that such details would obscure the essence of this disclosure. Additionally, the terminology used below is defined by consideration of the functions described in this disclosure and may be varied according to the intentions, practices, etc., of the user or operator. Therefore, the terminology should be defined based on the overall description of this disclosure.

[0011] In embodiments of this disclosure, each block in the flowchart and combinations of blocks in the flowchart can be executed by computer program instructions. These computer program instructions can be loaded into a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, and the instructions, executed by the processor of the computer or other programmable data processing apparatus, can generate units for performing the functions specified in the flowchart blocks. The computer program instructions can also be stored in a computer-executable memory or computer-readable memory capable of directing the computer or other programmable data processing apparatus to perform its functions in a particular manner, and the instructions stored in the computer-executable memory or computer-readable memory can produce an article of art including instructions for performing the functions specified in the flowchart blocks. The computer program instructions can also be loaded into a computer or other programmable data processing apparatus.

[0012] Furthermore, each box in the flowchart may represent a module, fragment, or portion of code, which includes one or more executable instructions for performing a specified logical function. In embodiments of this disclosure, the functions mentioned in the boxes may occur out of order. For example, two boxes shown consecutively may execute substantially simultaneously, or boxes may sometimes execute in reverse order according to their corresponding functions.

[0013] As used in embodiments of this disclosure, the term "unit" refers to a software element or hardware element (such as a field-programmable gate array (FPGA) or application-specific integrated circuit (ASIC)) that performs a predetermined function. However, the term "unit" is not limited to software or hardware. A "unit" may be configured to reside in an addressable storage medium or to operate one or more processors. In embodiments of this disclosure, the term "unit" may include elements (such as software elements, object-oriented software elements, class elements, and task elements), processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and parameters. The functionality provided by a particular element or unit may be combined to reduce the number of elements or may be further divided into additional elements. Furthermore, in embodiments of this disclosure, a "unit" may include one or more processors.

[0014] The meanings of the terms used in this article are described below.

[0015] The term "robot mobile device" can refer to any type of device that moves automatically or under user control to perform various operations. As described in this disclosure, a robot mobile device can capture images of its surrounding environment to perform operations, identify objects included in the captured images, and perform operations based on the results of object recognition. Therefore, a robot mobile device can be equipped with a neural network model (recognition model) for object recognition. When this disclosure is described with reference to a robotic vacuum cleaner as a representative example of a robot mobile device, the method of training a neural network model (recognition model) according to one or more embodiments of this disclosure, and the neural network model trained using that method, are not limited to robotic vacuum cleaners and can be used with various other types of robot mobile devices. Furthermore, the method of training a neural network model according to one or more embodiments of this disclosure, and the neural network model trained using that method, can be used by specific types of devices other than robot mobile devices. Terms such as "robot device" and "automatically driven device" may be used instead of "robot mobile device."

[0016] As used herein, "recognition model" can refer to a neural network model used to recognize objects included in an image or video. In this disclosure, a robotic mobile device can use a recognition model to recognize objects in front of it and control its actuation based on the recognition results. In one or more embodiments of this disclosure, the recognition model can be fine-tuned to improve recognition performance in a specific space (the space of a particular user). Terms such as "object recognition model" and "object detection model" may also be used instead of "recognition model".

[0017] "Spatial scan data" can refer to data obtained by a robotic mobile device scanning space using sensors. According to embodiments of this disclosure, the robotic mobile device can scan space using a light detection and ranging (LiDAR) sensor and a camera, and therefore, spatial scan data can include map data (e.g., LiDAR scan data) and captured images. Depending on the type of sensors mounted on the robotic mobile device, the spatial scan data can include different types of data. According to embodiments of this disclosure, a server can obtain spatial information and illumination characteristics information, as described below, by analyzing the spatial scan data received from the robotic mobile device. Terms such as "scan data" can be used instead of "spatial scan data".

[0018] As used herein, "spatial information" refers to information about the layout of a space and may include information about the structure of the space (e.g., the location of walls, etc.) and information about objects arranged in the space (e.g., household appliances, furniture, etc.) (e.g., the category and location of objects). Furthermore, spatial information may additionally include information about the background of the space (e.g., the pattern, texture, hue, color distribution, and tone (white tone, wood tone, etc.) of the floor or wallpaper). According to embodiments of this disclosure, a robotic mobile device or server may generate a "spatial map" via Simultaneous Localization and Mapping (SLAM), and the spatial map may include spatial information. Terms such as "spatial layout" and "layout information" may be used instead of "spatial information." Additionally, terms such as "map data" and "LiDAR map" may be used instead of "spatial map."

[0019] As used herein, "illuminance characteristic information" can refer to information about the characteristics of illuminance generated by lighting or natural light in a space. That is, illuminance characteristic information can include various types of information related to the brightness and color temperature of multiple regions included in the space. For example, illuminance characteristic information can include information about variations in brightness and color temperature by region within the space and by time. According to embodiments of this disclosure, captured images including multiple regions in a particular space can each include illuminance characteristic information. Each captured image can indicate a brightness level (how bright the captured area was at the time of capture). Furthermore, each captured image can also include metadata corresponding to the illuminance characteristic information. According to embodiments of this disclosure, a spatial map of a particular space can also include illuminance characteristic information. For example, information such as average illuminance can be recorded for each of the multiple regions included in the spatial map. The illuminance characteristics of a space may be influenced by the structure of the space (e.g., the location of windows and lighting, the arrangement of objects, etc.). According to embodiments of this disclosure, a server can obtain illuminance characteristic information about a space by analyzing captured images of the space. Terms such as "brightness information" and "illuminance information" can also be used instead of "illuminance characteristic information."

[0020] As used herein, "failure log data" can refer to log data about failure events that occur while a robot mobile device is moving and operating within a space. In this context, a "failure event" or "failure" can refer to a situation where the robot mobile device encounters a problem in its movement or operation due to a failure to recognize an object. Furthermore, a "failure event" or "failure" can refer to a situation where the robot mobile device fails to recognize an object. Failure events can be of various types, such as getting stuck in a cable or failing to avoid feces. Failure log data can include information about the type and location of the failure event, and may also include spatial information about the failure event (e.g., a spatial map obtained via SLAM) and captured images (e.g., an image taken 5 seconds before the failure event occurred, an image taken at the time of the failure event, etc.). Therefore, failure log data can include information about where the failure (error) occurred in space and what type of failure (error) occurred. Terms such as "failure log," "failure event log data," "error log," or "failure situation information" can also be used instead of "failure log data."

[0021] "Generative artificial intelligence (AI)" can refer to a type of AI technology that can generate new text, images, etc. in response to input data (e.g., text, images, etc.). Representative examples of generative AI are described in the "Generative Models" section below.

[0022] A "generative model" can refer to a neural network model that implements generative AI technology. Generative models can learn patterns and structures in training data to generate new data with similar characteristics to the input data, or new data corresponding to the input data. For example, when the input data is text containing a question, a generative model can generate and output the answer to the question. Alternatively, for example, when the input data is text containing a request, a generative model can output text or images generated in response to the request. The "virtual object generation model" described below corresponds to a generative model. Terms such as "generative AI model" can be used instead of "generative model."

[0023] As used herein, "virtual space data" can refer to data representing a virtual space (e.g., three-dimensional (3D) modeling data). According to embodiments of this disclosure, a server can generate a virtual space by reflecting characteristics of real-world space (e.g., the layout and illumination characteristics of the space). That is, the server can generate virtual space data based on spatial information about a specific space and illumination characteristics information about that specific space. The virtual space can be the base space used to generate training data for training a recognition model. According to embodiments of this disclosure, the server can provide realistic and diverse variations to the space by generating the virtual space as a 3D space, and thus obtain a large amount of training data. Terms such as "simulated space data" can also be used instead of "virtual space data."

[0024] As used herein, a "virtual object generation model" can refer to a generative AI model that generates virtual objects that can be placed in a space based on the characteristics of that space. When receiving spatial information about a specific space as input, the virtual object generation model can determine virtual objects that may be placed in the space based on the spatial information, and generate and output data about the determined virtual objects. Optionally, when receiving spatial information about a specific space as input, the virtual object generation model can determine virtual objects that may cause failure events in the space based on the spatial information, and generate and output data about the determined virtual objects. According to embodiments of this disclosure, the virtual object generation model can generate and output data about virtual objects that may be placed in a specific space (e.g., the category and location of the virtual objects) based on context (such as the structure of a specific space and the arrangement of objects in the specific space). Virtual objects generated by the virtual object generation model can be added to the aforementioned virtual space. Terms such as "failure environment generation model," "failure situation generation model," "dangerous object generation model," "object causing failure generation model," "failure probability generation model," "obstacle generation model," "generative AI model," "generative model," and "neural network model" can also be used instead of "virtual object generation model."

[0025] As used herein, "virtual object data" can refer to data that includes information about the category and location of virtual objects. According to embodiments of this disclosure, a virtual object can be an object that a robotic mobile device needs to avoid or pay attention to while moving or performing an operation. Optionally, according to embodiments of this disclosure, a virtual object can refer to an object that may cause the robotic mobile device to fail in its movement or operation. According to embodiments of this disclosure, virtual object data may also include contextual information, which is information about objects more likely to be located near the virtual object (objects associated with the virtual object). For example, when the virtual object is a "cable," the contextual information of the virtual object could be "household appliance" and "computer," meaning that the virtual object as a "cable" is highly likely to be located near a "household appliance" or a "computer." According to embodiments of this disclosure, a server can create a virtual space environment to obtain training data by adding virtual objects in the virtual space based on virtual space data and virtual object data. Terms such as "dangerous object data," "object data leading to failure," "failure environment data," or "obstacle information / barrier information" may also be used instead of "virtual object data."

[0026] "Training data" can refer to data used to train a recognition model. According to embodiments of this disclosure, training data is data used to perform supervised learning on the recognition model and may consist of pairs of images containing objects and labels indicating the categories of the objects contained in the images. According to embodiments of this disclosure, training data may include "real data" and "synthetic data." Synthetic data can refer to virtual data that has similar statistical properties to real data and is generated by reproducing results similar to those obtained from analyzing real data. According to embodiments of this disclosure, synthetic data may include data obtained in the aforementioned virtual space environment (e.g., a virtual captured image of a virtual space where virtual objects are placed). Terms such as "synthetic image," "virtual data," or "virtual recognition data" may also be used instead of "synthetic data." Furthermore, terms such as "learning data" or "recognition data" may be used instead of "training data."

[0027] In the following, one or more embodiments of the present disclosure are described in detail with reference to the accompanying drawings.

[0028] This disclosure relates to a method for updating a recognition model embedded in a robotic mobile device, and more particularly, to a method for updating the recognition model to achieve high recognition performance in a specific space. To this end, in one or more embodiments of this disclosure, training data can be generated using a virtual space similar to real-world space and virtual objects that are highly likely to cause recognition failure in real-world space, and the generated training data can be used to train the recognition model.

[0029] Figure 1This is a diagram illustrating a system environment for updating a recognition model of a robotic mobile device according to an embodiment of the present disclosure. (Refer to...) Figure 1 The system according to embodiments of this disclosure may include a server 100 and a robotic vacuum cleaner 200. As described above, one or more embodiments of this disclosure can be applied to various types of robotic mobile devices other than the robotic vacuum cleaner 200.

[0030] like Figure 1 As shown, assume that the robotic vacuum cleaner 200 operates in space A. For example, space A could be the home of a user using the robotic vacuum cleaner 200. Space A can be referred to as the target space, and the server 100 can update the identification model of the robotic vacuum cleaner 200 to be optimized for the target space.

[0031] Server 100 can be various types of electronic devices with computing capabilities. According to embodiments of this disclosure, server 100 can be an Internet of Things (IoT) server for controlling the robotic vacuum cleaner 200 or providing services associated with the robotic vacuum cleaner 200.

[0032] When the robotic vacuum cleaner 200 requests the server 100 to update its recognition model, the server 100 can update the recognition model and send the updated recognition model to the robotic vacuum cleaner 200. Specifically, the server 100 can store recognition models (e.g., the same recognition model as the one in the robotic vacuum cleaner 200 or a base recognition model), and the server 100 can update the recognition model by performing fine-tuning on the stored recognition model, then send the parameter information of the updated recognition model to the robotic vacuum cleaner 200. Optionally, according to embodiments of this disclosure, the server 100 can send training data to the robotic vacuum cleaner 200, and the robotic vacuum cleaner 200 can update its recognition model using the received training data.

[0033] When the robotic vacuum cleaner 200 is used in space A, server 100 can update the recognition model of the robotic vacuum cleaner 200 to achieve high recognition performance. To achieve this, server 100 can obtain training data by using a virtual space similar to space A, and update the recognition model by using the obtained training data. Server 100 can use a generative model to generate virtual objects that the recognition model is highly likely to fail to recognize, and use the generated virtual objects when obtaining training data. The purpose and reasons for server 100 to update the recognition model of robotic vacuum cleaner 200 are described in detail below.

[0034] A robotic vacuum cleaner 200 can identify objects in front of it by performing object recognition on images captured by a camera set in the robotic vacuum cleaner 200, and determine whether to remove the object or move forward while avoiding it based on the category of the identified object. The robotic vacuum cleaner 200 may include a neural network model (i.e., a recognition model) for object recognition to identify objects included in the captured images. However, when the recognition performance (recognition accuracy) of the recognition model set in the robotic vacuum cleaner 200 is low, problems may occur where the robotic vacuum cleaner incorrectly identifies objects and fails to avoid them while moving forward when it needs to. For example, the robotic vacuum cleaner 200 may fail to recognize cables on the floor and become stuck in them while moving forward, or it may fail to avoid pet feces while moving forward because it fails to recognize them. In this way, a situation where the recognition model of the robotic vacuum cleaner 200 fails to recognize objects, causing problems with the movement or operation of the robotic vacuum cleaner 200, can be referred to as a "failure event". Additionally, a "failure event" can refer to a situation where the robot vacuum cleaner 200's recognition model fails to recognize an object.

[0035] The recognition performance of the recognition model in the robotic vacuum cleaner 200 may be affected by the category of the object being recognized or the surrounding environment (e.g., the structure of the space, the brightness of the space, nearby objects, the background of the space, etc.). Therefore, the recognition performance of the recognition model may vary depending on the space in which the robotic vacuum cleaner 200 is used.

[0036] Typically, the recognition model loaded onto the robotic vacuum cleaner 200 in a factory is a generic model loaded onto all robotic vacuum cleaners, and is very likely not optimized for a specific space. Therefore, when the robotic vacuum cleaner 200 is used in a specific space, the recognition performance of the recognition model may not be as high as expected.

[0037] To address this issue, the recognition model needs further training to be specifically designed for spaces where the robotic vacuum cleaner 200 is used (e.g., a user's home), and it is difficult to obtain a sufficient amount of data for training when collecting training data by capturing images of the space with objects in it.

[0038] In one or more embodiments of this disclosure, server 100 can obtain a sufficient amount of training data by collecting training data using a virtual space. Specifically, server 100 can generate a virtual space (e.g., a 3D space) that reflects the characteristics of space A using the robotic vacuum cleaner 200, generate virtual objects that the recognition model is highly likely to fail to recognize, and then obtain training data by using the generated virtual space and virtual objects. The specific process by which server 100 obtains training data and updates the recognition model is described in more detail below with reference to other accompanying drawings.

[0039] Figure 2 This is a diagram illustrating hardware components included in server 100 according to an embodiment of the present disclosure. (Refer to...) Figure 2 According to embodiments of this disclosure, server 100 may include a communication interface 110, a processor 120, and a memory 130. However, server 100 may include more or fewer components than those described above. Some or all of the communication interface 110, processor 120, and memory 130 may be implemented as a single chip.

[0040] Communication interface 110 is a component used for sending signals (control commands, data, etc.) to and receiving signals (control commands, data, etc.) from external devices, either wired or wirelessly, and can be implemented as a communication chipset supporting various communication protocols. Communication interface 110 can receive signals from the outside and output signals to processor 120, or it can send signals output from processor 120 to the outside. Server 100 can communicate with robotic vacuum cleaner 200 via communication interface 110.

[0041] Server 100 can receive a request to update the recognition model from robotic vacuum cleaner 200 via communication interface 110, and send the parameter information of the updated recognition model to robotic vacuum cleaner 200.

[0042] Processor 120 is a component that controls a series of processes to enable server 100 to operate according to one or more embodiments of the present disclosure as described below, and can be configured as one or more processors. The one or more processors included in processor 120 can be circuits (such as system-on-chip (SoC), integrated circuit (IC), etc.). The one or more processors included in processor 120 can be general-purpose processors (such as central processing unit (CPU), microprocessor unit (MPU), application processor (AP), digital signal processor (DSP), etc.), dedicated graphics processors (such as graphics processing unit (GPU) and vision processing unit (VPU)), dedicated AI processors (such as neural processing unit (NPU)), or dedicated communication processors (such as communication processor (CP)). When the one or more processors included in processor 120 are dedicated AI processors, the corresponding dedicated AI processor can be designed with a hardware architecture specifically designed to process a particular AI model.

[0043] Processor 120 can write data to or read data stored in memory 130, and in particular, execute a program or at least one instruction stored in memory 130 to process data according to predefined operating rules or an AI model. Therefore, processor 120 can perform operations according to one or more embodiments of this disclosure as described below. Unless otherwise stated, the description of server 100 in one or more embodiments of this disclosure described below is provided by server 100 or detailed components ( Figure 4 The operations performed by the communication module 410 to the identification model 490 can be considered to be performed by the processor 120.

[0044] Memory 130 is a component for storing various programs or data and may include storage media such as read-only memory (ROM), RAM, hard disk, optical disc ROM (CD-ROM), and digital video disc (DVD), or combinations of storage media. Memory 130 may not exist alone but may be configured to be included in processor 120. Memory 130 may include volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory. Memory 130 may store programs or at least one instruction for performing operations according to one or more embodiments of the present disclosure as described below. Memory 130 may provide stored data to processor 120 upon request from processor 120.

[0045] Figure 3 This is a diagram illustrating components included in a robotic vacuum cleaner 200 according to an embodiment of the present disclosure. (Refer to...) Figure 3According to embodiments of this disclosure, the robotic vacuum cleaner 200 may include a communication interface 210, an input / output (I / O) interface 220, a memory 230, a processor 240, a camera 250, a LiDAR sensor 260, and a drive module 270. However, the components of the robotic vacuum cleaner 200 are not limited to the examples described above, and the robotic vacuum cleaner 200 may include more or fewer components than those described. In embodiments of this disclosure, some or all of the communication interface 210, I / O interface 220, memory 230, and processor 240 may be implemented as a single chip, and the processor 240 may include one or more processors.

[0046] Communication interface 210 is a component for transmitting and receiving signals (control commands, data, etc.) to and from external devices, either wired or wirelessly, and can be configured to include a communication chipset supporting various communication protocols. Communication interface 210 can receive signals from the outside and output signals to processor 240, or it can send signals output from processor 240 to the outside. According to embodiments of this disclosure, robotic vacuum cleaner 200 can transmit spatial scanning data obtained using camera 250 and LiDAR sensor 260, as described below, to server 100 via communication interface 210.

[0047] I / O interface 220 may include an input interface (e.g., touch screen, hard button, microphone, etc.) for receiving control commands or information from the user, and an output interface (e.g., display panel, speaker, etc.) for displaying the results of operations performed according to the user's control or the status of the robotic vacuum cleaner 200.

[0048] Memory 230 is a component for storing various programs or data and may include storage media (such as ROM, RAM, hard disk, CD-ROM, and DVD) or combinations of storage media. Memory 230 may not exist alone but may be configured to be included in processor 240. Memory 230 may include volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory. Memory 230 may store programs for performing operations according to one or more embodiments of this disclosure as described below. Memory 230 may provide stored data to processor 240 upon request from processor 240.

[0049] Processor 240 is a component that controls a series of processes to enable the robotic vacuum cleaner 200 to operate according to one or more embodiments of the present disclosure as described below, and can be configured as one or more processors. In this case, the one or more processors can be general-purpose processors (such as CPUs, APs, DSPs, etc.), dedicated graphics processors (such as GPUs and VPUs), or dedicated AI processors (such as NPUs). For example, when one or more processors are dedicated AI processors, the dedicated AI processors can be designed with hardware architectures specifically designed to process a particular AI model.

[0050] Processor 240 can write data to or read data stored in memory 230, and in particular, run programs stored in memory 230 to process data according to predefined operating rules or AI models. Therefore, processor 240 can perform operations according to one or more embodiments of this disclosure as described below, and unless otherwise stated, operations described below in one or more embodiments of this disclosure as performed by the robotic vacuum cleaner 200 can be considered as performed by processor 240.

[0051] when Figure 3 When processor 240 includes recognition model 241, recognition model 241 may refer to a neural network model implemented by processor 240 running a program stored in memory 230. Recognition model 241 can identify objects included in an image captured by camera 250. Processor 240 can update recognition model 241 based on parameter information received from server 100.

[0052] Camera 250 is a component used to capture images of the surrounding environment of the robotic vacuum cleaner 200. The robotic vacuum cleaner 200 can capture images of its front side using camera 250 and identify objects included in the captured images using recognition model 241.

[0053] The LiDAR sensor 260 is a component used to scan the distance (depth) to walls or objects in the surrounding space. The robotic vacuum cleaner 200 can measure depth values ​​for multiple areas included in the target space using the LiDAR sensor 260 and send the measured depth values ​​(LiDAR scan data) to the server 100.

[0054] The drive module 270 is a module for performing the movement or cleaning operations of the robotic vacuum cleaner 200, and may include a motor and a battery. The motor is a component for providing the power required for the robotic vacuum cleaner 200 to perform cleaning operations. According to embodiments of this disclosure, the robotic vacuum cleaner 200 can move within space and perform cleaning operations (e.g., suction, etc.) due to the driving force provided by the motor. The battery can provide power to the components included in the robotic vacuum cleaner 200.

[0055] Figure 4 This is a diagram illustrating detailed components of a server 100 according to an embodiment of the present disclosure, differentiated based on their function or role. Figure 4 The detailed components of server 100 shown (i.e., communication modules 410 to identification models 490) can be software components implemented by server 100's processor 120 running programs stored in memory 130, or they can be virtual components without actual matching hardware devices. In other words, the operations performed by server 100's processor 120 by running programs stored in memory 130 can be categorized into multiple groups according to function or purpose, and the entities performing the operations respectively included in the multiple groups can be... Figure 4 The detailed components (i.e., communication module 410 to recognition model 490) are shown. Therefore, Figure 4 The description shown herein, which describes the operations performed by the detailed components (i.e., communication modules 410 to identification model 490), can be considered as being actually performed by the processor 120 of server 100 by running a program stored in memory 130. (Refer to...) Figures 4 to 11 This disclosure describes in detail one or more embodiments of the server 100 updating the identification model 241 of the robotic vacuum cleaner 200.

[0056] 1. Collect and send failure log data When a failure event occurs while cleaning is being performed in space A (target space), the robotic vacuum cleaner 200 can store information related to the failure event as failure log data and send the failure log data to the server 100. In this case, as described above, the failure event can refer to the failure of the robotic vacuum cleaner 200's recognition model 241 to identify the object and the failure of the robotic vacuum cleaner 200's cleaning operation due to the recognition failure.

[0057] Failure log data can include the type and location of failure events, spatial information (e.g., spatial maps obtained via SLAM), and captured images of the failure situation (e.g., images captured a few seconds before the failure event occurred, images captured when the failure event occurred, etc.). In other words, failure log data can include information about where and what type of failure occurred, the spatial structure of the failure, and the circumstances surrounding the failure.

[0058] According to embodiments of this disclosure, the robotic vacuum cleaner 200 can store failure log data in a memory 230 and send the failure log data to a server 100 when certain conditions are met. For example, the robotic vacuum cleaner 200 may send the failure log data to the server 100 when a failure event occurs a preset number of times or when failure events occur at a preset ratio of the total number of recently performed cleaning operations. Optionally, for example, the robotic vacuum cleaner 200 may periodically send failure log data to the server 100.

[0059] Reference Figure 4 When the robotic vacuum cleaner 200 sends failure log data to the communication module 410 of the server 100, the communication module 410 can send the received failure log data to the log data analysis module 430. The log data analysis module 430 can analyze the failure log data and output multiple pieces of information included in the failure log data.

[0060] 2. Train a virtual object generation model using failure log data. Among the multiple pieces of information included in the failure log data, the type and location of the failure event, as well as spatial information, can be used to train the virtual object generation model 450. As described above, the virtual object generation model 450 is a generative AI model for generating virtual object data based on spatial information. The virtual object generation model 450 can learn the relationship between spatial characteristics (specifically, the categories and locations of objects arranged in the space) and objects (specifically, objects that may cause failure events). By training the virtual object generation model 450 based on information about where and what type of failure event actually occurred, the virtual object generation model 450 can generate data about virtual objects that are highly likely to cause failure events (e.g., the categories and locations of virtual objects) based on the characteristics of the space. In detail, when spatial information about the target space (e.g., the structure of the space and the categories and locations of objects arranged in the space) and information about failure events (the categories and locations of objects that caused failure events) are input into the virtual object generation model 450, the virtual object generation model 450 can add new objects that may be arranged in the target space (specifically, objects that may cause failure events), and the added new objects correspond to virtual objects.

[0061] For example, suppose a cable is laid on the floor in an environment near electronic devices, but the robot vacuum cleaner 200's recognition model 241 fails to recognize the cable, resulting in a failure event where the robot vacuum cleaner 200 gets stuck in the cable, and the virtual object generation model 450 is trained using information about this failure event. In this case, when spatial information representing a similar environment (the environment where the electronic devices are located) is input into the virtual object generation model 450, the virtual object generation model 450 can generate the cable laid around the electronic devices as a virtual object.

[0062] The virtual object data output by the virtual object generation model 450 may include information about the category and location of the virtual object. Furthermore, the virtual object data may also include contextual information about nearby objects (objects that are highly likely to be located near the virtual object).

[0063] The following reference Figure 6 This section describes in detail the specific method for training a virtual object generation model 450 using failure log data.

[0064] Figure 5 This is a diagram illustrating the operation of a robot mobile device sending failure log data to a server according to an embodiment of the present disclosure. (Refer to...) Figure 5 When the robotic vacuum cleaner 200 sends failure log data to the server 100, the server 100 can use the failure log data to train a virtual object generation model. Figure 4 (450).

[0065] Figure 5 A first screen 510 indicating a failure event and a composition 520 of failure log data are shown. According to an embodiment of this disclosure, when a failure event occurs in the robotic vacuum cleaner 200, the first screen 510 is displayed on a user device (e.g., a smartphone) for controlling the robotic vacuum cleaner 200, allowing the user to view information about the currently occurring failure event. A spatial map 511 for the target space (space A) may be displayed on the first screen 510, and the location where the failure event occurred may be displayed on the spatial map 511. Furthermore, the first screen 510 may display an image 512 capturing the situation where the failure event occurred and information 513 about the category of the object that caused the failure event.

[0066] Referring to the composition 520 of the failure log data, the failure log data may include an identifier (ID) for identifying the robotic vacuum cleaner 200 where the failure event occurred, the location of the failure event (location on a spatial map), the type of failure event (e.g., getting stuck in a cable, failing to avoid feces, etc.), spatial information (e.g., a spatial map obtained via SLAM), and sample images capturing the failure event (e.g., an image taken 5 seconds before the failure event, an image taken at the time of the failure event, etc.). According to embodiments of this disclosure, the server 100 can receive failure log data not only from the robotic vacuum cleaner 200 operating in space A, but also from various external sources (e.g., robotic vacuum cleaners) used in other spaces. Therefore, the server 100 can train the virtual object generation model 450 using information about failure events occurring in various spaces. Thus, the virtual object generation model 450 can fully learn the relationship between the characteristics of the space (e.g., the structure of the space, the arrangement of objects in the space, etc.) and the objects that caused the failure.

[0067] Furthermore, as described below, when generating training data for the recognition model, captured images of failure scenarios included in the failure log data can also be used. For example, the robotic vacuum cleaner 200 can accurately identify the category of the object that caused the failure event through additional recognition attempts after the failure event, and the server 100 can generate training data by labeling the captured images included in the failure log data with the object's category.

[0068] The following reference Figure 6 and Figure 7 This section describes in detail the specific method for training a virtual object generation model 450 using failure log data.

[0069] Figure 6 This is a diagram illustrating the process of performing training and inference in a virtual object generation model 450 included in a server 100 according to an embodiment of the present disclosure. (Refer to...) Figure 6 The virtual object generation model 450 may include a layout encoder 451, an object encoder 452, an obstacle encoder 453, a Transformer encoder 454, and an obstacle extractor 455.

[0070] The virtual object generation model 450 can be trained using spatial information about the target space and information about the type and location of failure events. According to embodiments of this disclosure, when spatial information about the target space and information about the type and location of failure events are input into the virtual object generation model 450, the virtual object generation model 450 can output information about virtual objects that may cause failure events (e.g., the category and location of the virtual object and information about objects near the virtual object).

[0071] The inputs to and outputs from the detailed components included in the virtual object generation model 450 (i.e., layout encoder 451 to obstacle extractor 455) are described in detail. According to embodiments of this disclosure, the virtual object generation model 450 can be implemented as a transformer encoder, and therefore unsupervised learning techniques can be used to train the virtual object generation model 450. That is, a loss function can be used to train the virtual object generation model 450, which compares the outputs of the layout encoder 451, object encoder 452, and obstacle encoder 453 with the output of the transformer encoder 454, and its detailed description is omitted here.

[0072] The processor 120 of server 100 can obtain an image 61 representing the structure of the target space from spatial information and input the obtained image 61 into layout encoder 451. According to embodiments of this disclosure, such as Figure 6 As shown, image 61, which is a two-dimensional (2D) top-view image of the target space, can be input into layout encoder 451. In this case, a convolutional neural network (CNN) suitable for extracting features from an image can be used as layout encoder 451. Layout encoder 451 can extract features from image 61 representing the structure of the target space and output a feature vector.

[0073] The processor 120 of server 100 can extract the category c representing the objects 62 arranged in the target space from the spatial information. j Position t j and size s j The vector is then input into the object encoder 452. In this case, the position t of object 62 is... j This could be coordinate information indicating the position on image 61, which is a 2D top-view image. Furthermore, in this case, the size s of object 62... j This can be information indicating the size measured based on a direction (e.g., a one-dimensional (1D) size). The object encoder 452 can output a feature vector for each of the objects 62.

[0074] The processor 120 of server 100 can extract the category c representing the object 63 that caused the failure from information about the type and location of the failure event. j and position t j The vector is then input into the obstacle encoder 453. In this case, the position t of the object 63 that caused the failure is determined. jAlternatively, it could be coordinate information indicating the position on image 61, which is a 2D top-view image. Obstacle encoder 453 can output feature vectors for each of the objects 63 that resulted in failure. According to embodiments of this disclosure, processor 120 can allow virtual object generation model 450 to be trained in an autoregressive manner by masking some of the feature vectors extracted from the objects 63 that resulted in failure.

[0075] Figure 7 The image shows an example of the input (training data) used to train the virtual object generation model 450. Figure 7 The first image 710 is an image obtained by visualizing the data (training data) input to the virtual object generation model 450, and the second image 720 shows the vectors input to the virtual object generation model 450.

[0076] The first image 710 shows the objects and the objects that caused the failure event on a 2D top-view image representing the structure of the target space. An image from the first image 710, with the objects and the objects that caused the failure event removed, can be input into the layout encoder 451.

[0077] The second image 720 shows vectors representing the categories and locations of the objects shown in the first image 710 and the objects that caused the failure event. The vectors shown in the second image 720 can be input to the object encoder 452 and the obstacle encoder 453.

[0078] When the feature vectors extracted from the layout encoder 451, object encoder 452, and obstacle encoder 453 are input to the transformer encoder 454, the transformer encoder 454 can output a feature vector similar to the input feature vectors. Specifically, for the object 63 that causes failure, as seen in the first region 610, the transformer encoder 454 can output the feature vectors of objects nearby as well as the feature vector of object 63, thereby learning the placement probability of the obstacle (the object 63 that causes failure) based on the placement of object 62. Therefore, the query vector 64, as the output of the transformer encoder 454, can include the feature vectors of nearby objects and the feature vectors of obstacles, similar to the vectors included in the first region 610. When the query vector 64 passes through the obstacle extractor 455, information about the category and location of the new virtual object (obstacle) 65 can be output, and as seen in the second region 620, the feature vector of the new virtual object 65 and the feature vectors of its nearby objects can be output. In other words, the output of the virtual object generation model 450 can include the category and location of the new virtual object 65, as well as contextual information about the new virtual object 65.

[0079] 3. Begin updating the recognition model for the robotic vacuum cleaner. This describes a method for starting to update the recognition model 241 of the robotic vacuum cleaner 200.

[0080] According to embodiments of this disclosure, the robotic vacuum cleaner 200 may prompt the user whether to update the recognition model 241 upon initial startup or upon reset, and may begin updating the recognition model 241 upon the user's consent. Optionally, the robotic vacuum cleaner 200 may automatically begin updating the recognition model 241 upon initial startup or upon reset.

[0081] According to embodiments of this disclosure, when a user requests an update to the recognition model 241 via the I / O interface 220, the robotic vacuum cleaner 200 may begin updating the recognition model 241.

[0082] According to embodiments of this disclosure, when a failure event occurs, the robotic vacuum cleaner 200 asks the user whether to update the identification model 241, and if the user agrees to update, the update of the identification model 241 begins. Optionally, when a failure event occurs, the robotic vacuum cleaner 200 may automatically begin updating the identification model 241. Specifically, when a failure event occurs a preset number of times or when failure events occur at a preset ratio of the total number of recently performed cleaning operations, the robotic vacuum cleaner 200 may send failure log data to the server 100, and then automatically or with the user's consent begin updating the identification model 241.

[0083] In this way, when preset conditions are met, the robotic vacuum cleaner 200 can automatically or with the user's consent begin to update the recognition model 241.

[0084] 4. Collect and transmit space scan data When updating the recognition model 241 of the robotic vacuum cleaner 200, the robotic vacuum cleaner 200 can scan the target space (space A) using sensors (camera 250 and LiDAR sensor 260) installed therein, collect spatial scan data, and send the spatial scan data to the server 100. The spatial scan data may include map data (e.g., LiDAR scan data) collected via the LiDAR sensor 260 of the robotic vacuum cleaner 200 and captured images (e.g., images of multiple areas within the target space under various environmental (illuminance levels) conditions) obtained via the camera 250.

[0085] Reference Figure 8 The method for obtaining spatial scanning data about a target space (space A) is described in detail by the robotic vacuum cleaner 200.

[0086] When a pre-generated spatial map 800 exists, the robotic vacuum cleaner 200 can set an optimal route on the pre-generated spatial map 800 (e.g., by setting the route using a one-stroke drawing technique) and collect spatial scanning data while traveling along the set route. For example, while traveling along the route, the robotic vacuum cleaner 200 can scan the depth of the space at specific distances or intervals using a LiDAR sensor 260 or capture images of the space using a camera 250. In this case, the captured images can then be used when the server 100 determines the category and location of objects arranged in the target space, analyzes the illumination characteristics of the target space, or analyzes the background of the target space (e.g., the pattern, texture, color perception, color distribution, and hue (white hue, wood hue, etc.) of the target space).

[0087] Even when a pre-generated spatial map 800 is not available, the robotic vacuum cleaner 200 can collect spatial scan data while traveling around the target space. For example, upon initial startup, the robotic vacuum cleaner 200 can travel around the target space and scan it to generate a spatial map 800, and the robotic vacuum cleaner 200 can acquire spatial scan data during this process. In this case, since the robotic vacuum cleaner 200 cannot pre-set its travel route, it can collect spatial scan data while simultaneously setting its travel route in real time based on information sensed during travel.

[0088] According to embodiments of this disclosure, the robotic vacuum cleaner 200 can automatically or upon user request collect additional spatial scanning data when the environment of the target space is changed (e.g., turning lights on and off, manually placing objects that may lead to failure, etc.).

[0089] When the robotic vacuum cleaner 200 sends the spatial scanning data collected according to the above method to the server 100, the server 100 can obtain spatial information and illuminance characteristic information by analyzing the spatial scanning data. When the communication module 410 of the server 100 sends the received spatial scanning data to the spatial information analysis module 420, the spatial information analysis module 420 can obtain spatial information and illuminance characteristic information about the target space (space A) by analyzing the spatial scanning data. Specifically, the spatial information analysis module 420 can obtain spatial information by analyzing the structure of the target space based on at least one of the map data or captured images, analyzing the categories and positions of objects arranged in the target space, and analyzing the background of the target space. In addition, the spatial information analysis module 420 can determine the illuminance characteristics of the target space based on the captured images.

[0090] The method for obtaining illuminance characteristics information about a target space, performed by the spatial information analysis module 420, is described in detail below.

[0091] The spatial information analysis module 420 can determine the color temperature and luminance distribution in a target space by analyzing captured images included in spatial scan data. According to embodiments of this disclosure, the spatial information analysis module 420 can analyze color based on the red, green, and blue (RGB) distribution in the captured image, convert the color space of the captured image (e.g., convert it from the RGB color space to the International Commission on Illumination (CIE) Luminance, Red / Green, Yellow / Blue (LAB) color space), and subsequently analyze the luminance of the captured image. Therefore, the spatial information analysis module 420 can obtain data on the color temperature and luminance of each region of the target space. Furthermore, the spatial information analysis module 420 can obtain illuminance characteristics information about the target space by analyzing captured images included in spatial scan data in various other ways.

[0092] 5. Generate training data (for training the recognition model) To train the recognition model 241 of the robotic vacuum cleaner 200 to be optimized for a target space (space A), images of objects need to be captured in multiple areas of space A, where various categories of objects (especially those highly likely to cause failure) are arranged. However, it is difficult to obtain a sufficient amount of real data when training the recognition model solely using real data (such as images of objects captured in real-world space, e.g., images of captured failure events). Finding a large amount of real data is time-consuming, and there are practical limitations on the amount of data available from the user's real-world space. Therefore, in one or more embodiments of this disclosure, server 100 may generate a virtual space similar to the target space (space A), generate virtual objects highly likely to cause failure events in the target space, and subsequently use the virtual space and virtual objects to generate synthetic data for training the recognition model 241.

[0093] In the following text, refer to Figure 9 and Figure 10 The process of server 100 generating virtual space and virtual objects and generating training data through simulation of the virtual space is described in detail.

[0094] (1) Generate virtual space As mentioned above, the recognition performance of recognition model 241 may be greatly affected by the structure of the space, the arrangement of objects in the space, the background of the space, the color of the space according to the light source, the illuminance characteristics of the space, etc. Therefore, in order to reflect the main factors affecting recognition performance in the training data in a manner similar to the user's real-world space (space A), server 100 may generate a virtual space based on spatial information about the user's space and illuminance characteristics information of the space.

[0095] Figure 9 This is a diagram illustrating a method for generating a virtual space performed by server 100 according to an embodiment of the present disclosure. (Refer to...) Figure 9 The virtual space generation module 440 of server 100 can generate a virtual space based on the spatial information and illuminance characteristic information received from the spatial information analysis module 420.

[0096] Figure 9 The first image 910 can be an image obtained by visualizing the spatial information received as input by the virtual space generation module 440. In other words, the first image 910 can be the result of visualizing spatial information obtained by analyzing map data included in the spatial scan data through the spatial information analysis module 420. The first image 910 shows the overall structure of the target space (space A) and the categories and locations of objects arranged in the target space. That is, the spatial information received as input by the virtual space generation module 440 can include information about the structure of the target space and information about the objects arranged in the target space. In addition, the spatial information received by the virtual space generation module 440 can also include information about the background of the space (e.g., the pattern, texture, color, color distribution, and tone (white tone, wood tone, etc.) of the floor or wallpaper).

[0097] The spatial information input to the virtual space generation module 440 may be in the form of an image (such as the first image 910) or in the form of text containing information about the structure and arrangement of objects in the target space. Furthermore, the spatial information input to the virtual space generation module 440 may include information about the background of the space represented as an image (e.g., a captured image of the floor or wallpaper).

[0098] The illuminance characteristic information input to the virtual space generation module 440 is the result obtained by the spatial information analysis module 420 analyzing the captured image included in the spatial scan data, and can be various types of data (e.g., images or text) representing the color temperature and brightness distribution in the target space. That is, the illuminance characteristic information input to the virtual space generation module 440 can be data representing the color temperature and brightness of each of the multiple regions included in the target space.

[0099] The virtual space generation module 440 can generate a virtual space based on the input spatial information and illuminance characteristic information, and output virtual space data representing the generated virtual space. Figure 9 The second image 920 and the third image 930 are images obtained by visualizing the process of generating virtual space by the virtual space generation module 440.

[0100] Referring to the second image 920, the virtual space generation module 440 can generate the structure of a virtual space based on the input spatial information. As described above, the spatial information may include information about the structure of the target space. Therefore, the virtual space generation module 440 can generate the structure of the virtual space by reflecting the structure of the target space. For example, the virtual space generation module 440 can generate the structure of a virtual space that is identical to the structure of the target space. Alternatively, the virtual space generation module 440 can generate the structure of the virtual space by maintaining the structural characteristics of the target space (e.g., the shape of the overall spatial structure, etc.) but changing the details.

[0101] According to embodiments of this disclosure, the virtual space generation module 440 can generate the structure of a virtual space using a realistic 3D model. In this way, when generating a virtual space using a realistic 3D model, there are advantages such as the ability to make various changes to the characteristics of the space, the time of image capture, etc., during the process of obtaining training data (synthetic data) through simulation.

[0102] Referring to the third image 930, the virtual space generation module 440 can arrange objects in a virtual space using a defined structure. As described above, the spatial information may include information about the category and location of the objects arranged in the target space. Therefore, the virtual space generation module 440 can arrange objects in the virtual space by reflecting the category and location of the objects arranged in the target space. For example, the virtual space generation module 440 can arrange objects in the virtual space in the same way as it arranges objects in the target space. Alternatively, the virtual space generation module 440 can arrange objects in the target space while maintaining the characteristics of how objects are arranged in the target space (e.g., at least one room has a table and a desktop computer is placed on the table), but changing the details.

[0103] Once the structure of the virtual space is generated and objects are arranged within it, the virtual space generation module 440 can adjust the illuminance of the virtual space based on illuminance characteristic information. As described below, illuminance adjustments for the virtual space can also be performed during the simulation process used to generate training data. According to embodiments of this disclosure, the virtual space generation module 440 can determine a corresponding region of the target space for each of a plurality of regions included in the virtual space, and can either fully apply the color temperature and brightness of the corresponding region of the target space to the virtual space, or adjust the color temperature and brightness of the region of the target space by a specific percentage (e.g., 20% to 50%) and apply the result of the adjustment to the virtual space.

[0104] The virtual space generation module 440 can adjust the background of the virtual space by reflecting information about the background of the target space included in the spatial information. For example, the virtual space generation module 440 can generate a pattern and texture of floor or wallpaper in the virtual space that is the same as or similar to the pattern and texture of the floor or wallpaper in the target space.

[0105] Therefore, the virtual space generation module 440 can generate a virtual space by applying a layout identical to that of the target space determined from the spatial information, or by generating a virtual space with slightly modified structures or object arrangements while maintaining the characteristics of the target space's layout (e.g., many pieces of furniture arranged in a living room, wood patterns on the floor, etc.). Furthermore, the virtual space generation module 440 can apply various illuminance changes to the virtual space based on illuminance characteristic information. Additionally, the virtual space generation module 440 can adjust the background of the virtual space based on information about the background of the target space.

[0106] As described above, server 100 generates virtual space based on spatial information of real-world target space, thus generating realistic synthetic data and using it as training data.

[0107] (2) Generate virtual objects As described above, the virtual object generation model 450 can learn the relationship between the characteristics of the space (e.g., the structure of the space, the arrangement of objects in the space, etc.) and the objects that cause failure. Therefore, when spatial information about the target space (e.g., the structure of the space and the categories and positions of objects arranged in the space) and information about failure events (the categories and positions of the objects that cause the failure events) are input into the virtual object generation model 450, the virtual object generation model 450 can generate new objects (virtual objects) that may cause failure in the target space. In other words, the virtual object generation model 450 can generate and output data about virtual objects that may cause failure based on context (such as the structure of the target space and the arrangement of objects in the target space). According to embodiments of this disclosure, the virtual object can be an object in the target space whose recognition model 241 of the robotic vacuum cleaner 200 has a probability of failure greater than or equal to a preset threshold.

[0108] Reference Figure 10 Describe in detail the process of generating virtual objects using Virtual Object Generation Model 450. Figure 10 The first image 1010 is an image obtained by visualizing the input of the virtual object generation model 450. The first image 1010 shows the structure of the target space (space A) and the categories and positions of objects arranged in the target space. Additionally, the first image 1010 shows the categories and positions of objects 11, 12, and 13 that caused the failure event. Figure 4 As shown, the virtual object generation model 450 can take the spatial information of the target space output by the spatial information analysis module 420 and the type and location of the failure event output by the log data analysis module 430 as input.

[0109] As shown above (refer to the reference) Figure 6 The virtual object generation model 450 can input multiple pieces of information (spatial information of the target space and the type and location of the failure event) included in the first image 1010 in image or vector form. More specifically, the following three types of data can be input into the virtual object generation model 450.

[0110] 1) A 2D top-view image of the target space (an image obtained by removing objects and objects 11, 12 and 13 that caused the failure event from the first image 1010).

[0111] 2) A vector representing the category, location (2D coordinate information), and size of the objects (household appliances and furniture) shown in the first image 1010.

[0112] 3) A vector representing the category and location (2D coordinate information) of the objects 11, 12 and 13 shown in the first image 1010 that caused the failure event.

[0113] Figure 10 The second image 1020 is an image obtained by visualizing the output of the virtual object generation model 450. Unlike the first image 1010, a new virtual object 14 is added to the second image 1020. The virtual object generation model 450 can learn the relationships between the objects 11, 12, and 13 that caused the failure and the structure of the space and the objects arranged in the space (e.g., household appliances, furniture, etc.), and can generate a new virtual object 14 based on the learning results and according to the context of the target space.

[0114] According to embodiments of this disclosure, the virtual object data actually output from the virtual object generation model 450 may include the category (feces) and location (2D coordinate information) of the new virtual object 14 as a newly generated virtual object. Furthermore, the virtual object data may also include contextual information (e.g., indicating that the feces are located around a TV), which is information about nearby objects associated with the new virtual object 14. Figure 10 The second image 1020 corresponds to the result of adding the new virtual object 14 to the target space according to the virtual object data described above.

[0115] According to embodiments of this disclosure, the location of a new virtual object 14 included in the virtual object data can be determined based on contextual information included in the virtual object data. For example, the virtual object 14 can be positioned based on contextual information indicating "feces are located around the TV". Figure 10 The location in the second image 1020 is determined to be around TV.

[0116] In summary, when given spatial information about the target space and information about failure events as input, the virtual object generation model 450 can generate virtual object data (the category and location of the virtual object, as well as contextual information).

[0117] In other words, the virtual object generation model 450 can understand which situations are most likely to lead to failure and create similar situations (the situation where virtual objects of a certain category are placed around a specific object in space).

[0118] In other words, based on the structure of the target space and the arrangement of objects (e.g., household appliances and furniture) within the target space, the virtual object generation model 450 can learn to understand: what category of objects is most likely to be placed in the target space and where in the target space the object is most likely to be placed, and further, where in the target space the object is placed when the recognition model 241 is most likely to fail to recognize a particular category of object, and generate virtual objects accordingly.

[0119] (3) Add virtual objects in virtual space The Add Module 460 can add virtual objects in a specific area within the virtual space based on virtual space data and virtual object data.

[0120] As described above, the virtual object data may include information about the location of the virtual object, and the location of the virtual object may be represented by coordinate information on a 2D top-view image of the target space. Furthermore, as described above, the virtual space generation module 440 can generate a virtual space structure that is the same as or similar to the structure of the target space. Therefore, according to embodiments of this disclosure, the adding module 460 can identify a location in the virtual space corresponding to a location on a 2D top-view image of the target space and place the virtual object at the identified location. For example, the adding module 460 can identify a location in the virtual space corresponding to a location on a 2D top-view image of the target space. Figure 10 The position on the second image 1020 corresponds to the position, and the virtual object (feces) is placed at the identified position.

[0121] As described above, the virtual object data may include contextual information, which is information about objects that may be located near the virtual object. Furthermore, as described above, the virtual space generation module 440 can arrange objects in the virtual space by reflecting the category and location of objects arranged in the target space. Therefore, according to embodiments of this disclosure, based on the objects arranged in the virtual space and the contextual information included in the virtual object data, the adding module 460 can determine the location where the virtual object will be placed in the virtual space. In other words, the adding module 460 can determine, based on the contextual information, the objects arranged in the virtual space that will be located near the virtual object, and place the virtual object near the determined object. For example, when the category of the generated virtual object is "cable" and the contextual information is "table and computer," the adding module 460 can place the virtual object (cable) near the location where the table and computer are arranged in the virtual space.

[0122] In summary, according to embodiments of this disclosure, based on the location information of the virtual object included in the virtual object data, the adding module 460 can determine the location where the virtual object will be placed in the virtual space and place the virtual object at the determined location. Furthermore, according to embodiments of this disclosure, based on the context information of the objects arranged in the virtual space and the virtual object, the adding module 460 can determine the location where the virtual object will be placed in the virtual space and place the virtual object at the determined location.

[0123] (4) Generating training data through simulation Once a virtual object is added to the virtual space, the training data generation module 470 can generate training data through simulation using the virtual robotic vacuum cleaner. According to embodiments of this disclosure, the training data generation module 470 can perform simulations such that the virtual robotic vacuum cleaner performs cleaning or scanning operations (e.g., capturing images with a camera while moving) in the virtual space where the virtual object is placed. The training data generation module 470 can generate training data for training the recognition model 241 by using images of the virtual object captured by the virtual robotic vacuum cleaner in the virtual space (hereinafter referred to as "virtual captured images" or "synthetic images").

[0124] Reference Figure 11 Describe in detail the process by which the training data generation module 470 generates training data.

[0125] Figure 11 The first image 1110 shows a virtual robot vacuum cleaner 1111 placed in a virtual space generated by the virtual space generation module 440 of the server 100. The training data generation module 470 can place the virtual robot vacuum cleaner 1111 in the virtual space for simulation. (See reference...) Figure 11 In the first image 1110, the virtual space generation module 440 can modify the lighting or texture rendering for the virtual space in various ways before performing the simulation. As described above, the virtual space generation module 440 can adjust the illumination of the virtual space based on illumination characteristic information about the target space. Furthermore, the virtual space generation module 440 can adjust the background of the virtual space (e.g., the pattern, texture, color, color distribution, hue (white hue, wood hue, etc.) of the floor or wallpaper) based on information about the background of the space included in the spatial information. By allowing the virtual space generation module 440 to modify the lighting or texture rendering for the virtual space, various environments that may appear in the real-world target space can be generated in the virtual space.

[0126] Figure 11 The second image 1120 is an image used to illustrate the process of adding the virtual object 14 by the adding module 460 in the virtual space. As described above, the virtual object data output by the virtual object generation model 450 may include information about the position of the virtual object 14 (this information is represented as coordinates on a 2D top-view image of the target space), and therefore, the adding module 460 can position and place the virtual object 14 in a virtual space that is implemented in the same or similar way as the target space.

[0127] Optionally, module 460 can add virtual objects in the virtual space according to any of the various methods described above.

[0128] Figure 11The third image 1130 shows a virtual captured image (synthetic data) obtained by the virtual robot vacuum cleaner 1111 when a virtual object 14 is added to a virtual space. The training data generation module 470 can perform a simulation such that the virtual robot vacuum cleaner 1111 performs a cleaning or scanning operation in the virtual space where the virtual object 14 is placed. When the simulation is performed, the virtual robot vacuum cleaner 1111 can obtain a virtual captured image 1132 by performing image capture while moving in the virtual space. The training data generation module 470 can generate training data by labeling the virtual captured image 1132 with the category (feces) of the virtual object 14.

[0129] In this way, the training data generation module 470 can effectively obtain training data for various situations by generating synthetic data (virtual captured images) via virtual space and generating training data by using the synthetic data.

[0130] According to embodiments of this disclosure, the training data generation module 470 can generate training data by selecting some synthetic images from synthetic images collected through simulation. For example, the training data generation module 470 can perform tests on the synthetic images using the recognition model 490 and determine whether to use the synthetic images as training data based on the test results. The recognition model 490 used in the test can be the same model as the recognition model 241 embedded in the robotic vacuum cleaner 200, or it can be a basic recognition model that is typically loaded on all devices (all robotic vacuum cleaners) in the factory.

[0131] In detail, the training data generation module 470 can attempt to perform object recognition on the synthetic images obtained through simulation using the recognition model 490, and determine whether the object recognition is successful. The training data generation module 470 can generate training data by adjusting the ratio between the number of synthetic images where the recognition model 490 successfully performs object recognition and the number of synthetic images where the recognition model 490 fails to perform object recognition to a preset value. For example, the training data generation module 470 can perform an object recognition test on the initially collected synthetic images and select synthetic images such that the ratio of the number of synthetic images with a "successful" test result to the number of synthetic images with a "failed" test result is a preset ratio (e.g., 1:1, 10:1, etc.). Subsequently, the training data generation module 470 can generate training data by labeling the selected synthetic images with the categories of virtual objects included in the selected synthetic images.

[0132] In other words, the training data generation module 470 can determine whether to use each synthetic image as training data by using the object recognition test of the recognition model 490, and when it is determined that synthetic images are to be used as training data, it generates training data by labeling the synthetic images with the categories of virtual objects.

[0133] Overfitting or underfitting may occur when the training data includes only synthetic images in which the recognition model 490 successfully identifies objects, or conversely, when the training data includes only synthetic images in which the recognition model 490 fails to identify objects. The training data generation module 470 can improve the training efficiency of the recognition model 490 by adjusting the number of synthetic images with "successful" test results and the number of synthetic images with "failed" test results according to a specific ratio to achieve data balance.

[0134] According to embodiments of this disclosure, when a failure event occurs while the virtual robot vacuum cleaner 1111 is performing a cleaning operation in a virtual space, the training data generation module 470 can generate training data by using a synthetic image capturing the failure event.

[0135] Furthermore, the training data generation module 470 can also generate training data by using captured images included in the failure log data. Such training data based on real data and training data based on synthetic data can be used together to train the recognition model 490.

[0136] 6. Fine-tune the recognition model using the generated training data.

[0137] Figure 12 This is a diagram illustrating the operation of a server 100 according to an embodiment of the present disclosure, which performs fine-tuning on a recognition model using training data and updates the recognition model of a robot mobile device based on the fine-tuning results.

[0138] Reference Figure 12 The fine-tuning module 480 of server 100 can fine-tune the recognition model 490 using the training data 1200 output from the training data generation module 470. In this case, the recognition model 490 can be the same recognition model as the recognition model 241 in the robotic vacuum cleaner 200, or it can be a recognition model applicable to all devices.

[0139] According to embodiments of this disclosure, the fine-tuning module 480 can update the parameters of the recognition model 490 by further training the recognition model 490 via supervised learning using training data labeled with real values ​​(object categories).

[0140] 7. Update the recognition model of the robotic vacuum cleaner based on the fine-tuned recognition model.

[0141] When the fine-tuning (additional training) of the recognition model 490 is completed, the server 100 can request an update to the recognition model 241 while sending updated parameter information of the recognition model 490 to the robotic vacuum cleaner 200 via the communication module 410. The robotic vacuum cleaner 200 can update the recognition model 241 based on the received parameter information.

[0142] Reference Figures 13 to 18 A method for updating the recognition model of a robotic mobile device according to one or more embodiments of the present disclosure is described. Because it includes... Figures 13 to 18 The operations in the flowchart are by Figure 1 and Figure 2 The server 100 is executing, therefore it has already referred to Figures 1 to 12 The provided description applies equally even if omitted below. Figures 13 to 18 .

[0143] Reference Figure 13 In operation 1301, server 100 can obtain spatial scan data about the target space from the robot mobile device. According to embodiments of this disclosure, upon initial startup or reset, the robot mobile device can ask the user whether to update the recognition model, and begin updating the recognition model when the user agrees. According to embodiments of this disclosure, when a failure event occurs, the robot mobile device can automatically or with the user's consent begin updating the recognition model. When updating the robot mobile device's recognition model begins, the robot mobile device can collect spatial scan data by scanning the target space via sensors (cameras, LiDAR sensors, etc.) disposed therein, and send the spatial scan data to server 100. The spatial scan data may include map data (e.g., LiDAR scan data) collected via the robot mobile device's LiDAR sensor and captured images obtained via the camera.

[0144] In operation 1302, server 100 can obtain spatial information based on spatial scan data, wherein the spatial information includes information about the structure of the target space and the objects arranged in the target space. Figure 14 The detailed operation included in operation 1302 is shown in the figure.

[0145] Reference Figure 14 In operation 1401, server 100 can analyze the structure of the target space based on map data included in the spatial scan data.

[0146] In operation 1402, server 100 can analyze the category and location of objects arranged in the target space based on map data and captured images included in the spatial scan data.

[0147] According to embodiments of this disclosure, server 100 can analyze the background of a target space based on captured images included in spatial scan data.

[0148] According to embodiments of this disclosure, server 100 can obtain illuminance characteristics information about a target space based on spatial scan data. For example, server 100 can determine the color temperature distribution and brightness distribution in the target space by analyzing captured images included in the spatial scan data.

[0149] Return to reference Figure 13 In operation 1303, server 100 can obtain virtual object data by inputting spatial information into a generative model (virtual object generation model), wherein the virtual object data includes information about the category and location of the virtual object. According to embodiments of this disclosure, the generative model can be a neural network model trained using the type of failure event occurring in the target space, the location of the failure event, and spatial information of the target space, and the failure event can be an event where the robot mobile device's recognition model fails to recognize an object. Furthermore, the virtual object can be an object in the target space whose probability of being failed to be recognized by the robot mobile device's recognition model is greater than or equal to a preset threshold.

[0150] According to embodiments of this disclosure, by learning the relationship between spatial characteristics (e.g., the structure of the space, the arrangement of objects in the space, etc.) and objects that may lead to failure, a generative model can generate and output data about virtual objects that may cause failure when spatial information is taken as input. (See also...) Figure 6 and Figure 10 A description of the training and inference of the generative model (generation of virtual objects) is provided. Furthermore, according to embodiments of this disclosure, the virtual object data may also include contextual information about nearby objects associated with the virtual object.

[0151] In operation 1304, server 100 can obtain training data by using spatial information and virtual object data. Figure 15 The detailed operation included in operation 1304 is shown in the figure.

[0152] Reference Figure 15 In operation 1501, server 100 can generate a virtual space based on spatial information. Figure 16 The detailed operation included in operation 1501 is shown in the figure.

[0153] Reference Figure 16In operation 1601, server 100 can determine the structure of a virtual space and the categories and positions of objects arranged in the virtual space based on spatial information. The spatial information may include information about the structure of the target space, and server 100 can generate the structure of the virtual space by reflecting the structure of the target space. For example, server 100 can generate a virtual space structure identical to that of the target space, or generate a virtual space structure by maintaining the structural characteristics of the target space but changing the details. The spatial information may include information about the categories and positions of objects arranged in the target space, and server 100 can arrange objects in the virtual space by reflecting the categories and positions of objects arranged in the target space. For example, server 100 can arrange objects in the same way as in the target space, or arrange objects in the target space by maintaining the characteristics of the object arrangement in the target space but changing the details.

[0154] In operation 1602, server 100 can determine the illuminance of each region in a plurality of regions within a virtual space based on illuminance characteristic information about the target space. According to embodiments of this disclosure, virtual space generation module 440 can determine a corresponding region of the target space for each of the plurality of regions included in the virtual space, and can apply the color temperature and brightness of the corresponding region of the target space to the virtual space as is, or adjust the color temperature and brightness of the region of the target space by a specific ratio and apply the adjustment result to the virtual space.

[0155] According to embodiments of this disclosure, server 100 adjusts the background of the virtual space by reflecting information about the background of the target space included in spatial information.

[0156] Return to reference Figure 15 In operation 1502, server 100 can add virtual objects in the virtual space. Figure 17 The detailed operation included in operation 1502 is shown in the figure.

[0157] Reference Figure 17 In operation 1701, server 100 can determine the location where the virtual object will be placed based on the objects arranged in the virtual space and context information, and then in operation 1702, server 100 can place the virtual object at the determined location.

[0158] The virtual object data may include information about the location of the virtual object, and the location of the virtual object may be represented by coordinate information on a 2D top-view image of the target space. Furthermore, the structure of the virtual space may be generated to be the same as or similar to the structure of the target space. Therefore, according to embodiments of this disclosure, server 100 can identify a location in the virtual space corresponding to a location on a 2D top-view image of the target space and place the virtual object at the identified location.

[0159] The virtual object data may also include contextual information, which is information about objects that may be located near the virtual object. Therefore, server 100 can determine the location in the virtual space where the virtual object will be placed based on the objects arranged in the virtual space and the contextual information included in the virtual object data.

[0160] Return to reference Figure 15 In operation 1503, server 100 may perform a simulation of a virtual space to obtain a composite image of a virtual object captured in the virtual space. According to embodiments of this disclosure, server 100 may perform a simulation to cause a virtual robot mobile device to perform a scanning operation in a virtual space (e.g., capture images with a camera while moving), thereby obtaining an image (composite image) of a virtual object captured by the virtual robot mobile device.

[0161] In operation 1504, server 100 can generate training data by using synthetic images. According to embodiments of this disclosure, server 100 can generate training data by labeling synthetic images with categories of virtual objects.

[0162] Furthermore, according to embodiments of this disclosure, server 100 can perform object recognition on the synthetic image using a recognition model embedded in server 100, determine whether the object recognition is successful, and then determine whether to use the synthetic image as training data based on the determined result. When it is determined that the synthetic image will be used as training data, server 100 can generate training data by labeling the synthetic image with the category of virtual object.

[0163] Return to reference Figure 13 In operation 1305, server 100 can update the recognition model of the robot's mobile device using training data. Figure 18 The detailed operation included in operation 1305 is shown in the figure.

[0164] Reference Figure 18 In operation 1801, server 100 can fine-tune the recognition model included in the electronic device (server 100) by using training data.

[0165] In operation 1802, server 100 may request an update to the recognition model while sending parameter information of the fine-tuned recognition model to the robot mobile device. The robot mobile device may update the recognition model embedded in the robot mobile device itself based on the received parameter information.

[0166] According to one or more embodiments disclosed above, high recognition performance in the target space can be achieved by updating the recognition model of the robot mobile device to be optimized for the target space where the robot mobile device is used. Furthermore, according to one or more embodiments of this disclosure, training data for various situations can be effectively obtained by using training data generated through virtual objects that reflect real failure events.

[0167] According to embodiments of this disclosure, a method for updating the recognition model of a robot mobile device may include: obtaining spatial scanning data about a target space from the robot mobile device by an electronic device; obtaining spatial information based on the spatial scanning data by the electronic device, wherein the spatial information includes information about the structure of the target space and objects arranged in the target space; obtaining virtual object data by the electronic device by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category and location of the virtual objects; obtaining training data by the electronic device using the spatial information and the virtual object data; and updating the recognition model of the robot mobile device by the electronic device using the training data.

[0168] According to embodiments of this disclosure, spatial scanning data may include map data obtained by scanning the target space using a LiDAR sensor and images of the target space captured using a camera.

[0169] According to embodiments of this disclosure, obtaining spatial information may include: analyzing the structure of a target space using electronic devices based on map data; and analyzing the categories and locations of objects arranged in the target space using electronic devices based on map data and captured images.

[0170] According to embodiments of this disclosure, the generative model can be a neural network model trained using the type of failure event occurring in the target space, the location of the failure event, and spatial information of the target space, and the failure event can be an event in which the robot's mobile device's recognition model fails to recognize an object.

[0171] According to embodiments of this disclosure, a virtual object can be an object in the target space whose recognition model of the robot mobile device fails to recognize it with a probability greater than or equal to a preset threshold.

[0172] According to embodiments of this disclosure, obtaining training data may include: generating a virtual space based on spatial information by an electronic device; adding virtual objects in the virtual space by an electronic device; obtaining a synthetic image of the virtual objects captured in the virtual space by an electronic device through performing a simulation of the virtual space; and generating training data by an electronic device using the synthetic image.

[0173] According to embodiments of this disclosure, obtaining spatial information may include: obtaining illuminance characteristic information about a target space by an electronic device based on spatial scanning data, and generating a virtual space may include: determining the structure of the virtual space and the categories and positions of objects arranged in the virtual space by the electronic device based on the spatial information; and determining the illuminance of each of the multiple regions in the virtual space by the electronic device based on the illuminance characteristic information.

[0174] According to embodiments of this disclosure, the virtual object data may further include contextual information about nearby objects associated with the virtual object, and adding a virtual object may include: determining, by an electronic device, the location where the virtual object will be placed based on objects arranged in the virtual space and the contextual information; and placing the virtual object at the determined location by the electronic device.

[0175] According to embodiments of this disclosure, generating training data using synthetic images may include: generating training data by labeling synthetic images with categories of virtual objects.

[0176] According to embodiments of this disclosure, generating training data using synthetic images may include: performing object recognition on the synthetic image using a recognition model; determining whether the object recognition was successful; determining whether to use the synthetic image as training data based on the determined result; and when it is determined that the synthetic image will be used as training data, generating training data by labeling the synthetic image with the category of a virtual object.

[0177] According to embodiments of this disclosure, updating the recognition model of a robot mobile device may include: fine-tuning the recognition model included in the electronic device by using training data; and requesting an update of the recognition model by sending parameter information of the fine-tuned recognition model to the robot mobile device.

[0178] According to embodiments of this disclosure, an electronic device for updating a recognition model of a robot mobile device may include: a memory storing a program or at least one instruction for updating the recognition model; and at least one processor configured to run the program or at least one instruction to cause the electronic device to perform the following operations: obtaining spatial scan data about a target space from the robot mobile device; obtaining spatial information based on the spatial scan data, wherein the spatial information includes information about the structure of the target space and objects arranged in the target space; obtaining virtual object data by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category and location of virtual objects; obtaining training data by using the spatial information and the virtual object data; and updating the recognition model of the robot mobile device by using the training data.

[0179] According to embodiments of this disclosure, spatial scanning data may include map data obtained by scanning the target space using a LiDAR sensor and images of the target space captured using a camera.

[0180] According to embodiments of this disclosure, when obtaining spatial information, the electronic device may be configured to: analyze the structure of a target space based on map data, and subsequently analyze the categories and locations of objects arranged in the target space based on the map data and captured images.

[0181] According to embodiments of this disclosure, the generative model can be a neural network model trained using the type of failure event occurring in the target space, the location of the failure event, and spatial information of the target space, and the failure event can be an event in which the robot's mobile device's recognition model fails to recognize an object.

[0182] According to embodiments of this disclosure, a virtual object can be an object in the target space whose recognition model of the robot mobile device fails to recognize it with a probability greater than or equal to a preset threshold.

[0183] According to embodiments of this disclosure, when obtaining training data, the electronic device may be configured to: generate a virtual space based on spatial information, add virtual objects in the virtual space, obtain a synthetic image of the virtual objects captured in the virtual space by performing a simulation of the virtual space, and generate training data by using the synthetic image.

[0184] According to embodiments of the present disclosure, at least one processor may also be configured to run a program or at least one instruction to enable an electronic device to obtain illuminance characteristics information about a target space based on spatial scan data, and when generating a virtual space, the electronic device may be configured to determine the structure of the virtual space and the categories and positions of objects arranged in the virtual space based on the spatial information, and subsequently determine the illuminance of each of the multiple regions in the virtual space based on the illuminance characteristics information.

[0185] According to embodiments of this disclosure, the virtual object data may further include contextual information about nearby objects associated with the virtual object, and when adding a virtual object, the electronic device may be configured to determine the location where the virtual object will be placed based on objects arranged in the virtual space and the contextual information, and then place the virtual object at the determined location.

[0186] One or more embodiments of this disclosure may be implemented or supported by one or more computer programs, which may be created from computer-readable program code and included on a computer-readable medium. As used herein, the terms "application" and "program" may refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media accessible by a computer, such as ROM, RAM, hard disk drive (HDD), CD, DVD, or various other types of storage.

[0187] Furthermore, machine-readable storage media may be provided in the form of non-transitory storage media. Here, "non-transitory storage media" refers to a tangible device and may not include wired communication links, wireless communication links, optical communication links, or other communication links that transmit transient electrical signals or other signals. Moreover, the term "non-transitory storage media" does not distinguish between data that is semi-permanently stored in the storage medium and data that is temporarily stored in the storage medium. For example, "non-transitory storage media" may include buffers for temporarily storing data. Computer-readable media can be any available medium accessible by a computer and includes both volatile and non-volatile media, as well as both removable and non-removable media. Computer-readable media include media on which data can be permanently stored and media on which data can be stored and later rewritten, such as rewritable optical discs or erasable memory devices.

[0188] According to embodiments of this disclosure, methods based on one or more embodiments of this disclosure set forth herein may be included in a computer program product when provided. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc ROM (CD-ROM)), or distributed online via an app store (e.g., downloaded or uploaded), or distributed directly between two user devices (e.g., smartphones). For online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be at least temporarily stored or temporarily generated in a machine-readable storage medium (such as the memory of a manufacturer's server, an app store's server, or a relay server).

[0189] The above description of this disclosure is provided for illustrative purposes, and those skilled in the art will understand that changes in form and detail can be readily made therein without departing from the technical concept or essential features of this disclosure. For example, suitable effects may be achieved even if the above-described techniques are performed in a different order than described above, and / or components of the above-described systems, structures, devices, circuits, etc., are combined or integrated in different forms and modes than described above, or replaced or supplemented by other components or their equivalents. Therefore, the above embodiments and all aspects of this disclosure are merely examples and not limitations. For example, each component defined as an integrated component can be implemented in a distributed manner, and similarly, components defined as individual components can be implemented in an integrated form.

[0190] The scope of this disclosure is not limited by its detailed description but by the appended claims, and all changes or modifications to the meaning and scope of the appended claims and their equivalents shall be construed as being included within the scope of this disclosure.

Claims

1. A method for updating a recognition model of a robot mobile device, the method comprising: Spatial scan data about the target space is obtained from the robot mobile device (200) by the electronic device (100); The electronic device (100) obtains spatial information based on the spatial scanning data, wherein the spatial information includes information about the structure of the target space and objects in the target space; The electronic device (100) obtains virtual object data by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category of the virtual object and the location of the virtual object; The electronic device (100) obtains training data by using the spatial information and the virtual object data; and The electronic device (100) updates the recognition model of the robot mobile device (200) by using the training data.

2. The method of claim 1, wherein, The spatial scanning data includes images of the target space captured by a camera (250) and map data obtained by scanning the target space using a light detection and ranging LiDAR sensor (260).

3. The method of claim 1 or claim 2, wherein, Obtaining the spatial information includes: The electronic device (100) analyzes the structure of the target space based on the map data; and The electronic device (100) analyzes the category and location of the objects in the target space based on the map data and the image of the target space.

4. The method according to any one of claims 1 to 3, wherein, The generative model includes a neural network model, wherein the neural network model is trained using the type of failure event occurring in the target space, the location of the failure event, and the spatial information. The failure events include events where the recognition model fails to recognize an object.

5. The method of any one of claims 1 to 4, wherein, The virtual object is an object in the target space whose probability of being failed to be recognized by the recognition model is greater than or equal to a preset threshold.

6. The method according to any one of claims 1 to 5, wherein, Obtaining the training data includes: The electronic device (100) generates a virtual space based on the spatial information; The electronic device (100) adds the virtual object in the virtual space; The electronic device (100) obtains a composite image of the virtual object captured in the virtual space by performing a simulation of the virtual space; and The training data is generated by the electronic device (100) using the synthesized image.

7. The method according to any one of claims 1 to 6, in, Obtaining the spatial information includes: the electronic device (100) obtaining illuminance characteristic information about the target space based on the spatial scanning data, and Generating the virtual space includes: The electronic device (100) determines the structure of the virtual space, the category of the object, and the position of the object in the virtual space based on the spatial information; and The electronic device (100) determines the illuminance of each of the multiple regions in the virtual space based on the illuminance characteristic information.

8. The method as claimed in any one of claims 1 to 7, in, The virtual object data also includes contextual information about nearby objects related to the virtual object, and Adding the virtual object includes: The electronic device (100) determines the location in the virtual space where the virtual object will be placed, based on the object in the virtual space and the context information; and The electronic device (100) places the virtual object at a specific location in the virtual space.

9. The method according to any one of claims 1 to 8, wherein, Generating the training data using the synthetic images includes labeling the synthetic images with the categories of the virtual objects.

10. The method according to any one of claims 1 to 9, wherein, Generating the training data using the synthesized image includes: The recognition model is used to perform object recognition on the synthetic image; Determine whether the object recognition was successful; Based on the result of determining whether the object recognition was successful, determine whether to use the synthetic image as the training data; and Based on the determination to use the synthetic image as the training data, the training data is generated by labeling the synthetic image with the category of the virtual object.

11. An electronic device (100) for updating a recognition model of a robot mobile device (200), said electronic device comprising: Memory (130), storing a program or at least one instruction; as well as At least one processor (120) is operatively coupled to the memory (130). Wherein, when the program or the at least one instruction is executed by the at least one processor (120), the electronic device (100) causes the electronic device (100) to perform the following operations: Spatial scan data about the target space is obtained from the robot's mobile device (200); Spatial information is obtained based on the spatial scanning data, wherein the spatial information includes information about the structure of the target space and objects in the target space; Virtual object data is obtained by inputting the spatial information into a generative model, wherein the virtual object data includes information about the category of the virtual object and the location of the virtual object; Training data is obtained by using the spatial information and the virtual object data; and The recognition model of the robot mobile device (200) is updated by using the training data.

12. The electronic device of claim 11, wherein, The spatial scanning data includes images of the target space captured by a camera and map data obtained by scanning the target space using a light detection and ranging LiDAR sensor (260).

13. The electronic device as claimed in claim 11 or claim 12, wherein, When the program or the at least one instruction is executed by the at least one processor (120), the electronic device (100) acquires the spatial information by performing the following operations: The structure of the target space is analyzed based on the map data; as well as The category and location of the objects in the target space are analyzed based on the map data and the image of the target space.

14. The electronic device as claimed in any one of claims 11 to 13, in, The generative model includes a neural network model, wherein the neural network model is trained using the type of failure event occurring in the target space, the location of the failure event, and the spatial information. The failure events include events where the recognition model fails to recognize an object.

15. The electronic device according to any one of claims 11 to 14, wherein, When the program or the at least one instruction is executed by the at least one processor (120), the electronic device (100) acquires the training data by: A virtual space is generated based on the aforementioned spatial information. Add the virtual object to the virtual space. A composite image of the virtual object captured in the virtual space is obtained by simulating the virtual space. The training data is generated using the synthesized images.